Ana Freire

dblp:78/1624 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
1since 2021 · last 2021
0000-0003-4698-2129ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 3 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Information retrieval · 97% Query processing and optimization · 3%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Energy-efficient computing · 50% Distributed systems · 50%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
web search
0.522017
Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web Search · ACM Trans. Inf. Syst. 2017
A self-adapting latency/power tradeoff model for replicated search engines · WSDM 2014
Information retrieval › user behavior › search behavior
click behavior
0.312017
Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web Search · ACM Trans. Inf. Syst. 2017
Information retrieval
user behavior
0.312017
Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web Search · ACM Trans. Inf. Syst. 2017
Information retrieval › distributed information retrieval
distributed search
0.212014
A self-adapting latency/power tradeoff model for replicated search engines · WSDM 2014
Energy-efficient computing
power management
0.212014
A self-adapting latency/power tradeoff model for replicated search engines · WSDM 2014
Distributed systems
replication
0.212014
A self-adapting latency/power tradeoff model for replicated search engines · WSDM 2014
Information retrieval
distributed information retrieval
0.112012
Scheduling queries across replicas · SIGIR 2012
Information retrieval
query processing
0.112017
Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web Search · ACM Trans. Inf. Syst. 2017
Information retrieval › evaluation › query performance prediction
query efficiency prediction
0.012012
Scheduling queries across replicas · SIGIR 2012
Query processing and optimization
query scheduling
0.012012
Scheduling queries across replicas · SIGIR 2012

Methods — techniques the papers use, named apart from their topics

dynamic programming · 0.4machine learning · 0.3response-time prediction · 0.1dynamic pruning · 0.1
YearPublicationVenuePosition
2021 How diverse is the ACII community? Analysing gender, geographical and business diversity of Affective Computing research
abstract
ACII is the premier international forum for presenting the latest research on affective computing. In this work, we monitor, quantify and reflect on the diversity in ACII conference across time by computing a set of indexes. We measure diversity in terms of gender, geographic location and academia vs research centres vs industry, and consider three different actors: authors, keynote speakers and organizers. Results raise awareness on the limited diversity in the field, in all studied facets, and compared to other AI conferences. While gender diversity is relatively high, equality is far from being reached. The community is dominated by European, Asian and North American researchers, leading the rest of continents under-represented. There is also a strong absence of companies and research centres focusing on applied research and products. This study fosters discussion in the community on the need for diversity and related challenges in terms of minimizing potential biases of the developed systems to the represented groups. We intend our paper to contribute with a first analysis to consider as a monitoring tool when implementing diversity initiatives. The data collected for this study are publicly released through the European divinAI initiative.
Isabelle Hupont, Songül Tolan, Ana Freire, Lorenzo Porcaro, Sara Estevez, Emilia Gómez
ACII3
2020 Enhanced Word Embeddings for Anorexia Nervosa Detection on Social Media
abstract
Anorexia Nervosa (AN) is a serious mental disorder that has been proved to be traceable on social media through the analysis of users’ written posts. Here we present an approach to generate word embeddings enhanced for a classification task dedicated to the detection of Reddit users with AN. Our method extends Word2vec ’s objective function in order to put closer domain-specific and semantically related words. The approach is evaluated through the calculation of an average similarity measure, and via the usage of the embeddings generated as features for the AN screening task. The results show that our method outperforms the usage of fine-tuned pre-learned word embeddings, related methods dedicated to generate domain adapted embeddings, as well as representations learned on the training set using Word2vec . This method can potentially be applied and evaluated on similar tasks that can be formalized as document categorization problems. Regarding our use case, we believe that this approach can contribute to the development of proper automated detection tools to alert and assist clinicians.
Diana Ramírez-Cifuentes, Christine Largeron, Julien Tissier, Ana Freire, Ricardo Baeza-Yates
IDA4
2017 Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web Search
abstract
The interplay between the response latency of web search systems and users’ search experience has only recently started to attract research attention, despite the important implications of response latency on monetisation of such systems. In this work, we carry out two complementary studies to investigate the impact of response latency on users’ searching behaviour in web search engines. We first conduct a controlled user study to investigate the sensitivity of users to increasing delays in response latency. This study shows that the users of a fast search system are more sensitive to delays than the users of a slow search system. Moreover, the study finds that users are more likely to notice the response latency delays beyond a certain latency threshold, their search experience potentially being affected. We then analyse a large number of search queries obtained from Yahoo Web Search to investigate the impact of response latency on users’ click behaviour. This analysis demonstrates the significant change in click behaviour as the response latency increases. We also find that certain user, context, and query attributes play a role in the way increasing response latency affects the click behaviour. To demonstrate a possible use case for our findings, we devise a machine-learning framework that leverages the latency impact, together with other features, to predict whether a user will issue any clicks on web search results. As a further extension of this use case, we investigate whether this machine-learning framework can be exploited to help search engines reduce their energy consumption during query processing.
Xiao Bai 0002, Ioannis Arapakis, Berkant Barla Cambazoglu, Ana Freire
ACM Trans. Inf. Syst.4
2014 A self-adapting latency/power tradeoff model for replicated search engines
abstract
For many search settings, distributed/replicated search engines deploy a large number of machines to ensure efficient retrieval. This paper investigates how the power consumption of a replicated search engine can be automatically reduced when the system has low contention, without compromising its efficiency. We propose a novel self-adapting model to analyse the trade-off between latency and power consumption for distributed search engines. When query volumes are high and there is contention for the resources, the model automatically increases the necessary number of active machines in the system to maintain acceptable query response times. On the other hand, when the load of the system is low and the queries can be served easily, the model is able to reduce the number of active machines, leading to power savings. The model bases its decisions on examining the current and historical query loads of the search engine. Our proposal is formulated as a general dynamic decision problem, which can be quickly solved by dynamic programming in response to changing query loads. Thorough experiments are conducted to validate the usefulness of the proposed adaptive model using historical Web search traffic submitted to a commercial search engine. Our results show that our proposed self-adapting model can achieve an energy saving of 33% while only degrading mean query completion time by 10 ms compared to a baseline that provisions replicas based on a previous day's traffic.
Ana Freire, Craig Macdonald, Nicola Tonellotto, Iadh Ounis, Fidel Cacheda
WSDM1
2013 Hybrid Query Scheduling for a Replicated Search Engine
Ana Freire, Craig Macdonald, Nicola Tonellotto, Iadh Ounis, Fidel Cacheda
ECIR1
2012 Scheduling queries across replicas
abstract
For increased efficiency, an information retrieval system can split its index into multiple shards, and then replicate these shards across many query servers. For each new query, an appropriate replica for each shard must be selected, such that the query is answered as quickly as possible. Typically, the replica with the lowest number of queued queries is selected. However, not every query takes the same time to execute, particularly if a dynamic pruning strategy is applied by each query server. Hence, the replica's queue length is an inaccurate indicator of the workload of a replica, and can result in inefficient usage of the replicas. In this work, we propose that improved replica selection can be obtained by using query efficiency prediction to measure the expected workload of a replica. Experiments are conducted using 2.2k queries, over various numbers of shards and replicas for the large GOV2 collection. Our results show that query waiting and completion times can be markedly reduced, showing that accurate response time predictions can improve scheduling accuracy and attesting the benefit of the proposed scheduling algorithm.
Ana Freire, Craig Macdonald, Nicola Tonellotto, Iadh Ounis, Fidel Cacheda
SIGIR1
2009 Revisiting N-Gram Based Models for Retrieval in Degraded Large Collections
Javier Parapar, Ana Freire, Álvaro Barreiro
ECIR2