Ida Mele

dblp:20/9056 · DBLP profile ↗
← Back
27ranked-venue papers in the field
10as first author
6since 2021 · last 2024
0000-0002-3730-6383ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 22 (9 first)Data Mining & Knowledge Discovery · 3 (1 first)Database Systems & Data Management · 2
YearPublicationVenuePosition
2024 A template-based approach for question answering over knowledge bases
abstract
Abstract In this paper, we address the problem of answering complex questions formulated by users in natural language. Since traditional information retrieval systems are not suitable for complex questions, these questions are usually run over knowledge bases, such as Wikidata or DBpedia. We propose a semi-automatic approach for transforming a natural language question into a SPARQL query that can be easily processed over a knowledge base. The approach applies classification techniques to associate a natural language question with a proper query template from a set of predefined templates. The nature of our approach is semi-automatic as the query templates are manually written by human assessors, who are the experts of the knowledge bases, whereas the classification and query processing steps are completely automatic. Our experiments on the large-scale CSQA dataset for question-answering corroborate the effectiveness of our approach.
Anna Formica, Ida Mele, Francesco Taglino
Knowl. Inf. Syst.2
2024 Caching Historical Embeddings in Conversational Search
abstract
Rapid response, namely, low latency, is fundamental in search applications; it is particularly so in interactive search sessions, such as those encountered in conversational settings. An observation with a potential to reduce latency asserts that conversational queries exhibit a temporal locality in the lists of documents retrieved. Motivated by this observation, we propose and evaluate a client-side document embedding cache, improving the responsiveness of conversational search systems. By leveraging state-of-the-art dense retrieval models to abstract document and query semantics, we cache the embeddings of documents retrieved for a topic introduced in the conversation, as they are likely relevant to successive queries. Our document embedding cache implements an efficient metric index, answering nearest-neighbor similarity queries by estimating the approximate result sets returned. We demonstrate the efficiency achieved using our cache via reproducible experiments based on Text Retrieval Conference Conversational Assistant Track datasets, achieving a hit rate of up to 75% without degrading answer quality. Our achieved high cache hit rates significantly improve the responsiveness of conversational systems while likewise reducing the number of queries managed on the search back-end.
Ophir Frieder, Ida Mele, Cristina Ioana Muntean, Franco Maria Nardini, Raffaele Perego 0001, Nicola Tonellotto
ACM Trans. Web2
2023 A Knowledge-Based Approach to Business Process Analysis: From Informal to Formal
Antonio De Nicola 0001, Anna Formica, Ida Mele, Michele Missikoff, Francesco Taglino
DEXA (1)3
2022 The 2nd Workshop on Mixed-Initiative ConveRsatiOnal Systems (MICROS)
abstract
The Mixed-Initiative ConveRsatiOnal Systems workshop (MICROS) aims at bringing novel ideas and investigating new solutions on conversational assistant systems. The increasing popularity of personal assistant systems, as well as smartphones, has changed the way users access online information, posing new challenges for information seeking and filtering. MICROS has a particular focus on mixed-initiative conversational systems, namely, systems that can provide answers in a proactive way (e.g., asking for clarification or proposing possible interpretations for ambiguous and vague requests). We invite people working on conversational systems or interested in the workshop topics to send us their position and research manuscripts.
Ida Mele, Cristina Ioana Muntean, Mohammad Aliannejadi, Nikos Voskarides
CIKM1
2021 MICROS: Mixed-Initiative ConveRsatiOnal Systems Workshop
Ida Mele, Cristina Ioana Muntean, Mohammad Aliannejadi, Nikos Voskarides
ECIR (2)1
2021 Adaptive utterance rewriting for conversational search
Ida Mele, Cristina Ioana Muntean, Franco Maria Nardini, Raffaele Perego 0001, Nicola Tonellotto, Ophir Frieder
Inf. Process. Manag.1
2020 Topic Propagation in Conversational Search
abstract
In a conversational context, a user expresses her multi-faceted information need as a sequence of natural-language questions, i.e., utterances. Starting from a given topic, the conversation evolves through user utterances and system replies. The retrieval of documents relevant to a given utterance in a conversation is challenging due to ambiguity of natural language and to the difficulty of detecting possible topic shifts and semantic relationships among utterances. We adopt the 2019 TREC Conversational Assistant Track (CAsT) framework to experiment with a modular architecture performing: (i) topic-aware utterance rewriting, (ii) retrieval of candidate passages for the rewritten utterances, and (iii) neural-based re-ranking of candidate passages. We present a comprehensive experimental evaluation of the architecture assessed in terms of traditional IR metrics at small cutoffs. Experimental results show the effectiveness of our techniques that achieve an improvement of up to $0.28$ (+93%) for [email protected] and $0.19$ (+89.9%) for [email protected] w.r.t. the CAsT baseline.
Ida Mele, Cristina Ioana Muntean, Franco Maria Nardini, Raffaele Perego 0001, Nicola Tonellotto, Ophir Frieder
SIGIR1
2020 Topical result caching in web search engines
Ida Mele, Nicola Tonellotto, Ophir Frieder, Raffaele Perego 0001
Inf. Process. Manag.1
2019 Predicting the Topic of Your Next Query for Just-In-Time IR
Seyed Ali Bahrainian, Fattane Zarrinkalam, Ida Mele, Fabio Crestani
ECIR (1)3
2019 Event mining and timeliness analysis from heterogeneous news streams
Ida Mele, Seyed Ali Bahrainian, Fabio Crestani
Inf. Process. Manag.1
2018 Predicting Topics in Scholarly Papers
Seyed Ali Bahrainian, Ida Mele, Fabio Crestani
ECIR2
2018 Emotional Influence Prediction of News Posts
Anastasia Giahanou, Paolo Rosso, Ida Mele, Fabio Crestani
ICWSM3
2018 Early Commenting Features for Emotional Reactions Prediction
Anastasia Giahanou, Paolo Rosso, Ida Mele, Fabio Crestani
SPIRE3
2017 Linking News across Multiple Streams for Timeliness Analysis
abstract
Linking multiple news streams based on the reported events and analyzing the streams' temporal publishing patterns are two very important tasks for information analysis, discovering newsworthy stories, studying the event evolution, and detecting untrustworthy sources of information. In this paper, we propose techniques for cross-linking news streams based on the reported events with the purpose of analyzing the temporal dependencies among streams. Our research tackles two main issues: (1) how news streams are connected as reporting an event or the evolution of the same event and (2) how timely the newswires report related events using different publishing platforms. Our approach is based on dynamic topic modeling for detecting and tracking events over the timeline and on clustering news according to the events. We leverage the event-based clustering to link news across different streams and present two scoring functions for ranking the streams based on their timeliness in publishing news about a specific event.
Ida Mele, Seyed Ali Bahrainian, Fabio Crestani
CIKM1
2017 Sentiment Propagation for Predicting Reputation Polarity
Anastasia Giahanou, Julio Gonzalo 0001, Ida Mele, Fabio Crestani
ECIR3
2017 Event Detection for Heterogeneous News Streams
Ida Mele, Fabio Crestani
NLDB1
2017 A Cross-Platform Collection for Contextual Suggestion
abstract
Suggesting personalized venues helps users to find interesting places on location-based social networks (LBSNs). Although there are many LBSNs online, none of them is known to have thorough information about all venues. The Contextual Suggestion track at TREC aimed at providing a collection consisting of places as well as user context to enable researchers to examine and compare different approaches, under the same evaluation setting. However, the officially released collection of the track did not meet many participants' needs related to venue content, online reviews, and user context. That is why almost all successful systems chose to crawl information from different LBSNs. For example, one of the best proposed systems in the TREC 2016 Contextual Suggestion track crawled data from multiple LBSNs and enriched it with venue-context appropriateness ratings, collected using a crowdsourcing platform. Such collection enabled the system to better predict a venue's appropriateness to a given user's context. In this paper, we release both collections that were used by the system above. We believe that these datasets give other researchers the opportunity to compare their approaches with the top systems in the track. Also, it provides the opportunity to explore different methods to predicting contextually appropriate venues.
Mohammad Aliannejadi, Ida Mele, Fabio Crestani
SIGIR2
2017 A Collection for Detecting Triggers of Sentiment Spikes
abstract
The advent of social media has given the opportunity to users to publicly express and share their opinion about any topic. Public opinion is very important for the interested entities that can leverage such information in the process of making decisions. In addition, identifying sentiment changes and the likely causes that have triggered them allows interested parties to adjust their strategies and attract more positive sentiment. With the aim to facilitate research on this problem, we describe a collection of tweets that can be used for detecting and ranking the likely triggers of sentiment spikes towards different entities. To build the collection, we first group tweets by topic which are then manually annotated according to sentiment polarity and strength. We believe that this collection can be useful for further research on detecting sentiment change triggers, sentiment analysis and sentiment prediction.
Anastasia Giahanou, Ida Mele, Fabio Crestani
SIGIR2
2016 Explaining Sentiment Spikes in Twitter
abstract
Tracking public opinion in social media provides important information to enterprises or governments during a decision making process. In addition, identifying and extracting the causes of sentiment spikes allows interested parties to redesign and adjust strategies with the aim to attract more positive sentiments. In this paper, we focus on the problem of tracking sentiment towards different entities, detecting sentiment spikes and on the problem of extracting and ranking the causes of a sentiment spike. Our approach combines LDA topic model with Relative Entropy. The former is used for extracting the topics discussed in the time window before the sentiment spike. The latter allows to rank the detected topics based on their contribution to the sentiment spike.
Anastasia Giahanou, Ida Mele, Fabio Crestani
CIKM2
2016 Network-Aware Recommendations of Novel Tweets
abstract
With the rapid proliferation of microblogging services such as Twitter, a large number of tweets is published everyday often making users feel overwhelmed with information. Helping these users to discover potentially interesting tweets is an important task for such services. In this paper, we present a novel tweet-recommendation approach, which exploits network, content, and retweet analyses for making recommendations of tweets. The idea is to recommend tweets that are not visible to the user (i.e., they do not appear in the user timeline) because nobody in her social circles published or retweeted them. To do that, we create the user's ego-network up to depth two and apply the transitivity property of the friends-of-friends relationship to determine interesting recommendations, which are then ranked to best match the user's interests. Experimental results demonstrate that our approach improves the state-of-the-art technique.
Noor Aldeen Alawad, Aris Anagnostopoulos, Stefano Leonardi 0001, Ida Mele, Fabrizio Silvestri
SIGIR4
2016 R-Susceptibility: An IR-Centric Approach to Assessing Privacy Risks for Users in Online Communities
abstract
Privacy of Internet users is at stake because they expose personal information in posts created in online communities, in search queries, and other activities. An adversary that monitors a community may identify the users with the most sensitive properties and utilize this knowledge against them (e.g., by adjusting the pricing of goods or targeting ads of sensitive nature). Existing privacy models for structured data are inadequate to capture privacy risks from user posts.
Asia J. Biega, Krishna P. Gummadi, Ida Mele, Dragan Milchevski, Christos Tryfonopoulos, Gerhard Weikum
SIGIR3
2015 The Importance of Being Expert: Efficient Max-Finding in Crowdsourcing
abstract
Crowdsourcing is a computational paradigm whose distinctive feature is the involvement of human workers in key steps of the computation. It is used successfully to address problems that would be hard or impossible to solve for machines. As we highlight in this work, the exclusive use of nonexpert individuals may prove ineffective in some cases, especially when the task at hand or the need for accurate solutions demand some degree of specialization to avoid excessive uncertainty and inconsistency in the answers. We address this limitation by proposing an approach that combines the wisdom of the crowd with the educated opinion of experts. We present a computational model for crowdsourcing that envisions two classes of workers with different expertise levels. One of its distinctive features is the adoption of the threshold error model, whose roots are in psychometrics and which we extend from previous theoretical work. Our computational model allows to evaluate the performance of crowdsourcing algorithms with respect to accuracy and cost. We use our model to develop and analyze an algorithm for approximating the best, in a broad sense, of a set of elements. The algorithm uses naïve and expert workers to find an element that is a constant-factor approximation to the best. We prove upper and lower bounds on the number of comparisons needed to solve this problem, showing that our algorithm uses expert and naïve workers optimally up to a constant factor. Finally, we evaluate our algorithm on real and synthetic datasets using the CrowdFlower crowdsourcing platform, showing that our approach is also effective in practice.
Aris Anagnostopoulos, Luca Becchetti, Adriano Fazzone, Ida Mele, Matteo Riondato
SIGMOD Conference4
2015 Stochastic Query Covering for Fast Approximate Document Retrieval
abstract
We design algorithms that, given a collection of documents and a distribution over user queries, return a small subset of the document collection in such a way that we can efficiently provide high-quality answers to user queries using only the selected subset. This approach has applications when space is a constraint or when the query-processing time increases significantly with the size of the collection. We study our algorithms through the lens of stochastic analysis and prove that even though they use only a small fraction of the entire collection, they can provide answers to most user queries, achieving a performance close to the optimal. To complement our theoretical findings, we experimentally show the versatility of our approach by considering two important cases in the context of Web search. In the first case, we favor the retrieval of documents that are relevant to the query, whereas in the second case we aim for document diversification. Both the theoretical and the experimental analysis provide strong evidence of the potential value of query covering in diverse application scenarios.
Aris Anagnostopoulos, Luca Becchetti, Ilaria Bordino, Stefano Leonardi 0001, Ida Mele, Piotr Sankowski
ACM Trans. Inf. Syst.5
2014 Phrase Query Optimization on Inverted Indexes
abstract
Phrase queries are a key functionality of modern search engines. Beyond that, they increasingly serve as an important building block for applications such as entity-oriented search, text analytics, and plagiarism detection. Processing phrase queries is costly, though, since positional information has to be kept in the index and all words, including stopwords, need to be considered.
Avishek Anand, Ida Mele, Srikanta J. Bedathur, Klaus Berberich
CIKM2
2013 Web usage mining for enhancing search-result delivery and helping users to find interesting web content
abstract
Web usage mining is the application of data mining techniques to the data generated by the interactions of users with web servers. This kind of data, stored in server logs, represents a valuable source of information, which can be exploited to optimize the document-retrieval task, or to better understand, and thus, satisfy user needs.
Ida Mele
WSDM1
2012 The early-adopter graph and its application to web-page recommendation
abstract
In this paper we present a novel graph-based data abstraction for modeling the browsing behavior of web users. The objective is to identify users who discover interesting pages before others. We call these users early adopters. By tracking the browsing activity of early adopters we can identify new interesting pages early, and recommend these pages to similar users. We focus on news and blog pages, which are more dynamic in nature and more appropriate for recommendation.
Ida Mele, Francesco Bonchi, Aristides Gionis
CIKM1
2011 Stochastic query covering
abstract
In this paper we introduce the problem of query covering as a means to efficiently cache query results. The general idea is to populate the cache with documents that contribute to the result pages of a large number of queries, as opposed to caching the top documents for each query. It turns out that the problem is hard and solving it requires knowledge of the structure of the queries and the results space, as well as knowledge of the input query distribution. We formulate the problem under the framework of stochastic optimization; theoretically it can be seen as a stochastic universal version of set multicover. While the problem is NP-hard to be solved exactly, we show that for any distribution it can be approximated using a simple greedy approach. Our theoretical findings are complemented by experimental activity on real datasets, showing the feasibility and potential interest of query-covering approaches in practice.
Aris Anagnostopoulos, Luca Becchetti, Stefano Leonardi 0001, Ida Mele, Piotr Sankowski
WSDM4