EDBT 2026 Demo / reviewers in the wild / expert
Fabio Crestani
dblp:c/FabioCrestani · also Fabio A. Crestani
· DBLP profile ↗
167ranked-venue papers in the field
23as first author
26since 2021 · last 2026
0000-0001-8672-0700ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 139 (17 first)Database Systems & Data Management · 14 (4 first)Data Mining & Knowledge Discovery · 9Knowledge Engineering, Semantic Web & Information Systems · 4 (2 first)Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring User Simulators in Conversational Search: A Comparison Between LLMs and Humans
Lili Lu, Fabio Crestani |
ECIR (2) | 2 |
| 2026 | eRisk 2026: Tasks on Symptoms Ranking, Contextual and Conversational Approaches for Early Mental Health Detection
Anxo Pérez, Javier Parapar, Xi Wang 0012, Fabio Crestani |
ECIR (4) | 4 |
| 2026 | Toward Exploring Mixed-Initiative Conversation Generation Based on Community Question AnsweringabstractConversational search addresses users’ information needs through multi-turn and context-aware interactions. Given that user queries are often ambiguous, the use of clarifying questions can effectively reduce uncertainty and enable a mixed-initiative conversational system. However, current datasets for clarifying questions remain limited in the following three aspects: (1) underrepresented multi-turn conversational data, (2) limited diversity, and (3) heavily reliance on crowdsourcing, thereby suffering from limitations such as high annotation cost. To address these issues, we propose a large language model (LLM)-based three-stage framework that relies on an existing community question answering dataset. It encompasses: (1) extracting essential information from the initial user query with the relevant contextual information, (2) generating clarifying questions paired with corresponding answers, and (3) refining conversations to ensure coherence and a natural conversational flow. We assess our multi-stage method against a baseline that directly prompts LLMs to generate conversations in a single-step process, evaluating on an answer retrieval task using recall, precision, normalized discounted cumulative gain and mean average precision. Results show that our three-stage generation approach consistently outperforms the baseline particularly in recall, while also achieving competitive results across other metrics. Human and automatic evaluations further indicate the high quality of generated conversations and fine-tuning on them improves retrieval performance, highlighting the pipeline’s potential. Lili Lu, Pranav Kasela, Federico Ravenda, Chuan Meng, Gabriella Pasi, Fabio Crestani |
ACM Trans. Inf. Syst. | 6 |
| 2025 | TalkDep: Clinically Grounded LLM Personas for Conversation-Centric Depression ScreeningabstractThe increasing demand for mental health services has outpaced the availability of real training data to develop clinical professionals, leading to limited support for the diagnosis of depression. This shortage has motivated the development of simulated or virtual patients to assist in training and evaluation, but existing approaches often fail to generate clinically valid, natural, and diverse symptom presentations. In this work, we embrace the recent advanced language models as the backbone and propose a novel clinician-in-the-loop patient simulation pipeline, TalkDep, with access to diversified patient profiles to develop simulated patients. By conditioning the model on psychiatric diagnostic criteria, symptom severity scales, and contextual factors, our goal is to create authentic patient responses that can better support diagnostic model training and evaluation. We verify the reliability of these simulated patients with thorough assessments conducted by clinical professionals. The availability of validated simulated patients offers a scalable and adaptable resource for improving the robustness and generalisability of automatic depression diagnosis systems. Xi Wang 0012, Anxo Pérez, Javier Parapar, Fabio Crestani |
CIKM | 4 |
| 2025 | Zero-Shot and Efficient Clarification Need Prediction in Conversational Search
Lili Lu, Chuan Meng, Federico Ravenda, Mohammad Aliannejadi, Fabio Crestani |
ECIR (1) | 5 |
| 2025 | eRisk 2025: Contextual and Conversational Approaches for Depression Challenges
Javier Parapar, Anxo Pérez, Xi Wang 0012, Fabio Crestani |
ECIR (5) | 4 |
| 2025 | A self-supervised seed-driven approach to topic modelling and clusteringabstractAbstract Topic models are useful tools for extracting the most salient themes within a collection of documents, grouping them to construct clusters representative of each specific topic. These clusters summarize and represent the semantic contents of the documents for better document interpretation. In this work, we present a light approach able to learn topic representations in a Self-Supervised fashion. More specifically, we propose a lightweight and scalable architecture using a seed-word driven approach to simultaneously co-learn a representation from a document and its corresponding word embeddings. The results obtained on a variety of datasets of different sizes and natures show that our model is capable of extracting meaningful topics. Furthermore, our experiments on five benchmark datasets illustrate that our model outperforms both traditional and neural topic modelling baseline models in terms of different coherence and clustering accuracy measures. Federico Ravenda, Seyed Ali Bahrainian, Andrea Raballo, Antonietta Mira, Fabio Crestani |
J. Intell. Inf. Syst. | 5 |
| 2024 | Towards Self-Contained Answers: Entity-Based Answer Rewriting in Conversational SearchabstractConversational Information Seeking (CIS) is an emerging paradigm for knowledge acquisition and exploratory search. Traditional web search interfaces enable easy exploration of entities, but this is limited in conversational settings due to the limited-bandwidth interface. This paper explore ways to rewrite answers in CIS, so that users can understand them without having to resort to external services or sources. Specifically, we focus on salient entities—entities that are central to understanding the answer. As our first contribution, we create a dataset of conversations annotated with entities for saliency. Our analysis of the collected data reveals that the majority of answers contain salient entities. As our second contribution, we propose two answer rewriting strategies aimed at improving the overall user experience in CIS. One approach expands answers with inline definitions of salient entities, making the answer self-contained. The other approach complements answers with follow-up questions, offering users the possibility to learn more about specific entities. Results of a crowdsourcing-based study indicate that rewritten answers are clearly preferred over the original ones. We also find that inline definitions tend to be favored over follow-up questions, but this choice is highly subjective, thereby providing a promising future direction for personalization. Ivan Sekulic, Krisztian Balog, Fabio Crestani |
CHIIR | 3 |
| 2024 | eRisk 2024: Depression, Anorexia, and Eating Disorder Challenges
Javier Parapar, Patricia Martín-Rodilla, David E. Losada, Fabio Crestani |
ECIR (5) | 4 |
| 2024 | Estimating the Usefulness of Clarifying Questions and Answers for Conversational Search
Ivan Sekulic, Weronika Lajewska, Krisztian Balog, Fabio Crestani |
ECIR (3) | 4 |
| 2024 | Advances in information retrieval collection on the European conference on information retrieval 2023abstractAbstract This paper introduces the Collection on ECIR 2023. The 45th European Conference on Information Retrieval (ECIR 2023) was held in Dublin, Ireland, during April 2–6, 2023. The conference was the largest ECIR ever, and brought together hundreds of researchers from Europe and abroad. A selection of papers shortlisted for the best paper awards was asked to submit expanded versions appearing in this Discover Computing (formerly the Information Retrieval Journal) Collection on ECIR 2023. First, an analytic paper on incorporating first stage retrieval status values as input in neural cross-encoder re-rankers. Second, new models and new data for a new task of temporal natural language inference. Third, a weak supervision approach to video retrieval overcoming the need for large-scale human labeled training data. Together, these papers showcase the breadth and diversity of current research on information retrieval. Jaap Kamps, Lorraine Goeuriot, Fabio Crestani |
Discov. Comput. | 3 |
| 2024 | Analysing Utterances in LLM-Based User Simulation for Conversational SearchabstractClarifying underlying user information needs by asking clarifying questions is an important feature of modern conversational search systems. However, evaluation of such systems through answering prompted clarifying questions requires significant human effort, which can be time-consuming and expensive. In our recent work, we proposed an approach to tackle these issues with a user simulator,USi. Given a description of an information need,USiis capable of automatically answering clarifying questions about the topic throughout the search session. However, while the answers generated byUSiare both in line with the underlying information need and in natural language, a deeper understanding of such utterances is lacking. Thus, in this work, we explore utterance formulation of large language model (LLM)–based user simulators. To this end, we first analyze the differences betweenUSi, based on GPT-2, and the next generation of generative LLMs, such as GPT-3. Then, to gain a deeper understanding of LLM-based utterance generation, we compare the generated answers to the recently proposed set of patterns of human-based query reformulations. Finally, we discuss potential applications as well as limitations of LLM-based user simulators and outline promising directions for future work on the topic. Ivan Sekulic, Mohammad Aliannejadi, Fabio Crestani |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2023 | eRisk 2023: Depression, Pathological Gambling, and Eating Disorder Challenges
Javier Parapar, Patricia Martín-Rodilla, David E. Losada, Fabio Crestani |
ECIR (3) | 4 |
| 2023 | Exploiting Simulated User Feedback for Conversational Search: Ranking, Rewriting, and BeyondabstractThis research aims to explore various methods for assessing user feedback in mixed-initiative conversational search (CS ) systems. While CS systems enjoy profuse advancements across multiple aspects, recent research fails to successfully incorporate feedback from the users. One of the main reasons for that is the lack of system-user conversational interaction data. To this end, we propose a user simulator-based framework for multi-turn interactions with a variety of mixed-initiative CS systems. Specifically, we develop a user simulator, dubbed ConvSim, that, once initialized with an information need description, is capable of providing feedback to system's responses, as well as answering potential clarifying questions. Our experiments on a wide variety of state-of-the-art passage retrieval and neural re-ranking models show that effective utilization of user feedback can lead to 16% retrieval performance increase in terms of nDCG@3. Moreover, we observe consistent improvements as the number of feedback rounds increases (35% relative improvement in terms of nDCG@3 after three rounds). This points to a research gap in the development of specific feedback processing modules and opens a potential for significant advancements in CS. To support further research in the topic, we release over 30 000 transcripts of system-simulator interactions based on well-established CS datasets. Paul Owoicho, Ivan Sekulic, Mohammad Aliannejadi, Jeff Dalton 0001, Fabio Crestani |
SIGIR | 5 |
| 2022 | eRisk 2022: Pathological Gambling, Depression, and Eating Disorder Challenges
Javier Parapar, Patricia Martín-Rodilla, David E. Losada, Fabio Crestani |
ECIR (2) | 4 |
| 2022 | Exploiting Document-Based Features for Clarification in Conversational Search
Ivan Sekulic, Mohammad Aliannejadi, Fabio Crestani |
ECIR (1) | 3 |
| 2022 | Evaluating Mixed-initiative Conversational Search Systems via User SimulationabstractClarifying the underlying user information need by asking clarifying questions is an important feature of modern conversational search system. However, evaluation of such systems through answering prompted clarifying questions requires significant human effort, which can be time-consuming and expensive. In this paper, we propose a conversational User Simulator, called USi, for automatic evaluation of such conversational search systems. Given a description of an information need, USi is capable of automatically answering clarifying questions about the topic throughout the search session. Through a set of experiments, including automated natural language generation metrics and crowdsourcing studies, we show that responses generated by USi are both inline with the underlying information need and comparable to human-generated answers. Moreover, we make the first steps towards multi-turn interactions, where conversational search systems asks multiple questions to the (simulated) user with a goal of clarifying the user need. To this end, we expand on currently available datasets for studying clarifying questions, i.e., Qulac and ClariQ, by performing a crowdsourcing-based multi-turn data acquisition. We show that our generative, GPT2-based model, is capable of providing accurate and natural answers to unseen clarifying questions in the single-turn setting and discuss capabilities of our model in the multi-turn setting. We provide the code, data, and the pre-trained model to be used for further research on the topic. Ivan Sekulic, Mohammad Aliannejadi, Fabio Crestani |
WSDM | 3 |
| 2022 | The impact of psycholinguistic patterns in discriminating between fake news spreaders and fact checkersabstractFake news is a threat to society. A huge amount of fake news is posted every day on social networks which is read, believed and sometimes shared by a number of users. On the other hand, with the aim to raise awareness, some users share posts that debunk fake news by using information from fact-checking websites. In this paper, we are interested in exploring the role of various psycholinguistic characteristics in differentiating between users that tend to share fake news and users that tend to debunk them. Psycholinguistic characteristics represent the different linguistic information that can be used to profile users and can be extracted or inferred from users’ posts. We present the CheckerOrSpreader model that uses a Convolution Neural Network (CNN) to differentiate between spreaders and checkers of fake news. The experimental results showed that CheckerOrSpreader is effective in classifying a user as a potential spreader or checker. Our analysis showed that checkers tend to use more positive language and a higher number of terms that show causality compared to spreaders who tend to use a higher amount of informal language, including slang and swear words. Anastasia Giahanou, Bilal Ghanem, Esteban A. Ríssola, Paolo Rosso, Fabio Crestani, Daniel L. Oberski |
Data Knowl. Eng. | 5 |
| 2022 | Mental disorders on online social media through the lens of language and behaviour: Analysis and visualisationabstractDue to the worldwide accessibility to the Internet along with the continuous advances in mobile technologies, physical and digital worlds have become completely blended, and the proliferation of social media platforms has taken a leading role over this evolution. In this paper, we undertake a thorough analysis towards better visualising and understanding the factors that characterise and differentiate social media users affected by mental disorders. We perform different experiments studying multiple dimensions of language, including vocabulary uniqueness, word usage, linguistic style, psychometric attributes, emotions’ co-occurrence patterns, and online behavioural traits, including social engagement and posting trends. Our findings reveal significant differences on the use of function words, such as adverbs and verb tense, and topic-specific vocabulary, such as biological processes. As for emotional expression, we observe that affected users tend to share emotions more regularly than control individuals on average. Overall, the monthly posting variance of the affected groups is higher than the control groups. Moreover, we found evidence suggesting that language use on micro-blogging platforms is less distinguishable for users who have a mental disorder than other less restrictive platforms. In particular, we observe on Twitter less quantifiable differences between affected and control groups compared to Reddit. Esteban A. Ríssola, Mohammad Aliannejadi, Fabio Crestani |
Inf. Process. Manag. | 3 |
| 2022 | CATS: Customizable Abstractive Topic-based SummarizationabstractNeural sequence-to-sequence models are the state-of-the-art approach used in abstractive summarization of textual documents, useful for producing condensed versions of source text narratives without being restricted to using only words from the original text. Despite the advances in abstractive summarization, custom generation of summaries (e.g., towards a user’s preference) remains unexplored. In this article, we present CATS, an abstractive neural summarization model that summarizes content in a sequence-to-sequence fashion while also introducing a new mechanism to control the underlying latent topic distribution of the produced summaries. We empirically illustrate the efficacy of our model in producing customized summaries and present findings that facilitate the design of such systems. We use the well-known CNN/DailyMail dataset to evaluate our model. Furthermore, we present a transfer-learning method and demonstrate the effectiveness of our approach in a low resource setting, i.e., abstractive summarization of meetings minutes, where combining the main available meetings’ transcripts datasets, AMI and International Computer Science Institute(ICSI) , results in merely a few hundred training documents. Seyed Ali Bahrainian, George Zerveas, Fabio Crestani, Carsten Eickhoff |
ACM Trans. Inf. Syst. | 3 |
| 2022 | A Systematic Analysis on the Impact of Contextual Information on Point-of-Interest RecommendationabstractAs the popularity of Location-based Social Networks increases, designing accurate models for Point-of-Interest (POI) recommendation receives more attention. POI recommendation is often performed by incorporating contextual information into previously designed recommendation algorithms. Some of the major contextual information that has been considered in POI recommendation are the location attributes (i.e., exact coordinates of a location, category, and check-in time), the user attributes (i.e., comments, reviews, tips, and check-in made to the locations), and other information, such as the distance of the POI from user’s main activity location and the social tie between users. The right selection of such factors can significantly impact the performance of the POI recommendation. However, previous research does not consider the impact of the combination of these different factors. In this article, we propose different contextual models and analyze the fusion of different major contextual information in POI recommendation. The major contributions of this article are as follows: (i) providing an extensive survey of context-aware location recommendation; (ii) quantifying and analyzing the impact of different contextual information (e.g., social, temporal, spatial, and categorical) in the POI recommendation on available baselines and two new linear and non-linear models, which can incorporate all the major contextual information into a single recommendation model; and (iii) evaluating the considered models using two well-known real-world datasets. Our results indicate that while modeling geographical and temporal influences can improve recommendation quality, fusing all other contextual information into a recommendation model is not always the best strategy. Hossein A. Rahmani, Mohammad Aliannejadi, Mitra Baratchi, Fabio Crestani |
ACM Trans. Inf. Syst. | 4 |
| 2021 | eRisk 2021: Pathological Gambling, Self-harm and Depression Challenges
Javier Parapar, Patricia Martín-Rodilla, David E. Losada, Fabio Crestani |
ECIR (2) | 4 |
| 2021 | User Engagement Prediction for Clarification in Search
Ivan Sekulic, Mohammad Aliannejadi, Fabio Crestani |
ECIR (1) | 3 |
| 2021 | The Impact of User Demographics and Task Types on Cross-App Mobile Search
Mohammad Aliannejadi, Fabio Crestani, Theo Huibers, Monica Landoni, Emiliana Murgia, Maria Soledad Pera |
FQAS | 2 |
| 2021 | The impact of emotional signals on credibility assessmentabstractFake news is considered one of the main threats of our society. The aim of fake news is usually to confuse readers and trigger intense emotions to them in an attempt to be spread through social networks. Even though recent studies have explored the effectiveness of different linguistic patterns for fake news detection, the role of emotional signals has not yet been explored. In this paper, we focus on extracting emotional signals from claims and evaluating their effectiveness on credibility assessment. First, we explore different methodologies for extracting the emotional signals that can be triggered to the users when they read a claim. Then, we present emoCred, a model that is based on a long-short term memory model that incorporates emotional signals extracted from the text of the claims to differentiate between credible and non-credible ones. In addition, we perform an analysis to understand which emotional signals and which terms are the most useful for the different credibility classes. We conduct extensive experiments and a thorough analysis on real-world datasets. Our results indicate the importance of incorporating emotional signals in the credibility assessment problem. Anastasia Giahanou, Paolo Rosso, Fabio Crestani |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2021 | Context-aware Target Apps Selection and Recommendation for Enhancing Personal Mobile AssistantsabstractUsers install many apps on their smartphones, raising issues related to information overload for users and resource management for devices. Moreover, the recent increase in the use of personal assistants has made mobile devices even more pervasive in users’ lives. This article addresses two research problems that are vital for developing effective personal mobile assistants: target apps selection and recommendation . The former is the key component of a unified mobile search system: a system that addresses the users’ information needs for all the apps installed on their devices with a unified mode of access. The latter, instead, predicts the next apps that the users would want to launch. Here we focus on context-aware models to leverage the rich contextual information available to mobile devices. We design an in situ study to collect thousands of mobile queries enriched with mobile sensor data (now publicly available for research purposes). With the aid of this dataset, we study the user behavior in the context of these tasks and propose a family of context-aware neural models that take into account the sequential, temporal, and personal behavior of users. We study several state-of-the-art models and show that the proposed models significantly outperform the baselines. Mohammad Aliannejadi, Hamed Zamani, Fabio Crestani, W. Bruce Croft |
ACM Trans. Inf. Syst. | 3 |
| 2020 | Harnessing Evolution of Multi-Turn Conversations for Effective Answer RetrievalabstractWith the improvements in speech recognition and voice generation technologies over the last years, a lot of companies have sought to develop conversation understanding systems that run on mobile phones or smart home devices through natural language interfaces. Conversational assistants, such as Google Assistant and Microsoft Cortana, can help users to complete various types of tasks. This requires an accurate understanding of the user's information need as the conversation evolves into multiple turns. Finding relevant context in a conversation's history is challenging because of the complexity of natural language and the evolution of a user's information need. In this work, we present an extensive analysis of language, relevance, dependency of user utterances in a multi-turn information-seeking conversation. To this aim, we have annotated relevant utterances in the conversations released by the TREC CaST 2019 track. The annotation labels determine which of the previous utterances in a conversation can be used to improve the current one. Furthermore, we propose a neural utterance relevance model based on BERT fine-tuning, outperforming competitive baselines. We study and compare the performance of multiple retrieval models, utilizing different strategies to incorporate the user's context. The experimental results on both classification and retrieval tasks show that our proposed approach can effectively identify and incorporate the conversation context. We show that processing the current utterance using the predicted relevant utterance leads to a 38% relative improvement in terms of [email protected] Finally, to foster research in this area, we have released the dataset of the annotations. Mohammad Aliannejadi, Manajit Chakraborty, Esteban A. Ríssola, Fabio Crestani |
CHIIR | 4 |
| 2020 | A Tool for Conducting User Studies on Mobile DevicesabstractWith the ever-growing interest in the area of mobile information retrieval and the ongoing fast development of mobile devices and, as a consequence, mobile apps, an active research area lies in studying users' behavior and search queries users submit on mobile devices. However, many researchers require to develop an app that collects useful information from users while they search on their phones or participate in a user study. In this paper, we aim to address this need by providing a comprehensive Android app, called Omicron, which can be used to collect mobile query logs and perform user studies on mobile devices. Omicron, at its current version, can collect users' mobile queries, relevant documents, sensor data as well as user activity and interaction data in various study settings. Furthermore, we designed Omicron in such a way that it is conveniently extendable to conduct more specific studies and collect other types of sensor data. Finally, we provide a tool to monitor the participants and their data both during and after the collection process. Luca Costa, Mohammad Aliannejadi, Fabio Crestani |
CHIIR | 3 |
| 2020 | eRisk 2020: Self-harm and Depression Challenges
David E. Losada, Fabio Crestani, Javier Parapar |
ECIR (2) | 2 |
| 2020 | Joint Geographical and Temporal Modeling Based on Matrix Factorization for Point-of-Interest Recommendation
Hossein A. Rahmani, Mohammad Aliannejadi, Mitra Baratchi, Fabio Crestani |
ECIR (1) | 4 |
| 2020 | Beyond Modelling: Understanding Mental Disorders in Online Social Media
Esteban A. Ríssola, Mohammad Aliannejadi, Fabio Crestani |
ECIR (1) | 3 |
| 2020 | The Role of Personality and Linguistic Patterns in Discriminating Between Fake News Spreaders and Fact CheckersabstractUsers play a critical role in the creation and propagation of fake news online by consuming and sharing articles with inaccurate information either intentionally or unintentionally. Fake news are written in a way to confuse readers and therefore understanding which articles contain fabricated information is very challenging for non-experts. Given the difficulty of the task, several fact checking websites have been developed to raise awareness about which articles contain fabricated information. As a result of those platforms, several users are interested to share posts that cite evidence with the aim to refute fake news and warn other users. These users are known as fact checkers . However, there are users who tend to share false information, who can be characterised as potential fake news spreaders . In this paper, we propose the CheckerOrSpreader model that can classify a user as a potential fact checker or a potential fake news spreader. Our model is based on a Convolutional Neural Network (CNN) and combines word embeddings with features that represent users’ personality traits and linguistic patterns used in their tweets. Experimental results show that leveraging linguistic patterns and personality traits can improve the performance in differentiating between checkers and spreaders. Anastasia Giahanou, Esteban A. Ríssola, Bilal Ghanem, Fabio Crestani, Paolo Rosso |
NLDB | 4 |
| 2020 | A Joint Two-Phase Time-Sensitive Regularized Collaborative Ranking Model for Point of Interest RecommendationabstractThe popularity of location-based social networks (LBSNs) has led to a tremendous amount of user check-in data. Recommending points of interest (POIs) plays a key role in satisfying users needs in LBSNs. While recent work has explored the idea of adopting collaborative ranking (CR) for recommendation, there have been few attempts to incorporate temporal information for POI recommendation using CR. In this article, we propose a two-phase CR algorithm that incorporates the geographical influence of POIs and is regularized based on the variance of POIs popularity and users activities over time. The time-sensitive regularizer penalizes user and POIs that have been more time-sensitive in the past, helping the model to account for their long-term behavioral patterns while learning from user-POI interactions. Moreover, in the first phase, it attempts to rank visited POIs higher than the unvisited ones, and at the same time, apply the geographical influence. In the second phase, our algorithm tries to rank users favorite POIs higher on the recommendation list. Both phases employ a collaborative learning strategy that enables the model to capture complex latent associations from two different perspectives. Experiments on real-world datasets show that our proposed time-sensitive collaborative ranking model beats state-of-the-art POI recommendation methods. Mohammad Aliannejadi, Dimitrios Rafailidis, Fabio Crestani |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2019 | Understanding Mobile Search Task Relevance and User Behaviour in ContextabstractImprovements in mobile technologies have led to a dramatic change in how and when people access and use information, and is having a profound impact on how users address their daily information needs. Smart phones are rapidly becoming our main method of accessing information and are frequently used to perform "on-the-go'' search tasks. As research into information retrieval continues to evolve, evaluating search behaviour in context is relatively new. Previous research has studied the effects of context through either self-reported diary studies or quantitative log analysis; however, neither approach is able to accurately capture context of use at the time of searching. Mohammad Aliannejadi, Morgan Harvey, Luca Costa, Matthew Pointon, Fabio Crestani |
CHIIR | 5 |
| 2019 | Predicting the Topic of Your Next Query for Just-In-Time IR
Seyed Ali Bahrainian, Fattane Zarrinkalam, Ida Mele, Fabio Crestani |
ECIR (1) | 4 |
| 2019 | Early Detection of Risks on the Internet: An Exploratory Campaign
David E. Losada, Fabio Crestani, Javier Parapar |
ECIR (2) | 2 |
| 2019 | Anticipating Depression Based on Online Social Media Behaviour
Esteban A. Ríssola, Seyed Ali Bahrainian, Fabio Crestani |
FQAS | 3 |
| 2019 | Asking Clarifying Questions in Open-Domain Information-Seeking ConversationsabstractUsers often fail to formulate their complex information needs in a single query. As a consequence, they may need to scan multiple result pages or reformulate their queries, which may be a frustrating experience. Alternatively, systems can improve user satisfaction by proactively asking questions of the users to clarify their information needs. Asking clarifying questions is especially important in conversational systems since they can only return a limited number of (often only one) result(s). Mohammad Aliannejadi, Hamed Zamani, Fabio Crestani, W. Bruce Croft |
SIGIR | 3 |
| 2019 | Leveraging Emotional Signals for Credibility DetectionabstractThe spread of false information on the Web is one of the main problems of our society. Automatic detection of fake news posts is a hard task since they are intentionally written to mislead the readers and to trigger intense emotions to them in an attempt to be disseminated in the social networks. Even though recent studies have explored different linguistic patterns of false claims, the role of emotional signals has not yet been explored. In this paper, we study the role of emotional signals in fake news detection. In particular, we propose an LSTM model that incorporates emotional signals extracted from the text of the claims to differentiate between credible and non-credible ones. Experiments on real world datasets show the importance of emotional signals for credibility assessment. Anastasia Giahanou, Paolo Rosso, Fabio Crestani |
SIGIR | 3 |
| 2019 | Adversarial Training for Review-Based RecommendationsabstractRecent studies have shown that incorporating users' reviews into the collaborative filtering strategy can significantly boost the recommendation accuracy. A pressing challenge resides on learning how reviews influence users' rating behaviors. In this paper, we propose an Adversarial Training approach for Review-based recommendations, namely ATR. We design a neural architecture of sequence-to-sequence learning to calculate the deep representations of users' reviews on items following an adversarial training strategy. At the same time we jointly learn to factorize the rating matrix, by regularizing the deep representations of reviews with the user and item latent features. In doing so, our model captures the non-linear associations among reviews and ratings while producing a review for each user-item pair. Our experiments on publicly available datasets demonstrate the effectiveness of the proposed model, outperforming other state-of-the-art methods. Dimitrios Rafailidis, Fabio Crestani |
SIGIR | 2 |
| 2019 | Personality Recognition in Conversations using Capsule Neural NetworksabstractAutomatic identification of personality in conversations has many applications in natural language processing, such as community role identification (e.g., group leader) in online social media conversations as well as meeting transcripts. Conversation utterances provide a lot of information about the parties involved in a conversation such as cues to the participants’ personality traits, one of human’s most distinguishable attributes. However, traditional computational personality assessment models rely on limited domain-knowledge and various psychometric indicators. In this paper, we propose a novel model based on capsule neural networks to extract meaningful hidden patterns from conversations and use them to assess the personality of individuals. Our experimental results on a real-world dataset reveals evidence that personality can be captured from conversation utterances outperforming traditional approaches. Esteban A. Ríssola, Seyed Ali Bahrainian, Fabio Crestani |
WI | 3 |
| 2019 | Propagating sentiment signals for estimating reputation polarity
Anastasia Giahanou, Julio Gonzalo 0001, Fabio Crestani |
Inf. Process. Manag. | 3 |
| 2019 | Event mining and timeliness analysis from heterogeneous news streams
Ida Mele, Seyed Ali Bahrainian, Fabio Crestani |
Inf. Process. Manag. | 3 |
| 2018 | Friend Recommendation in Location-Based Social Networks via Deep Pairwise LearningabstractGenerating friend recommendations in location-based social networks is a challenging task, as we have to learn how different contextual factors influence users' behavior to form social relationships. For example, the contextual information of users' check-in behavior at common locations and users' activities at close regions may impact users' relationships. In this paper we propose a deep pairwise learning model, namely FDPL. Our model first learns the low dimensional latent embeddings of users' social relationships by jointly factorizing them with the available contextual information based on a multi-view learning strategy. In addition, to account for the fact that the contextual information is non-linearly correlated with users' social relationships we design a deep pairwise learning architecture based on a Bayesian personalized ranking strategy. We learn the non-linear deep representations of the computed low dimensional latent embeddings by formulating the top- k friend recommendation task at location-based social networks as a ranking task in our deep pairwise learning strategy. Our experiments on three real world location-based social networks from Brightkite, Gowalla and Foursquare show that the proposed FDPL model significantly outperforms other state-of-the-art methods. Finally, we evaluate the impact of contextual information on our model and we experimentally show that it is a key factor to boost the friend recommendation accuracy at location-based social networks. Dimitrios Rafailidis, Fabio Crestani |
ASONAM | 2 |
| 2018 | Augmentation of Human Memory: Anticipating Topics that Continue in the Next MeetingabstractMemory augmentation is the process of providing human memory with information that facilitates and complements the recall of an event in a person»s past. Recently, there has been a lot of attention on processing the content of meetings for later reuse, such as reviewing a meeting for supporting failing memories, keeping in mind key issues, verification, etc. That is due to the fact that meetings are essential for sharing knowledge in organizations. In this paper, we propose four novel time-series methods for predicting the topics that one should review in preparation for a next meeting. The predicted/recommended topics can be reviewed by a user as a memory augmentation process to facilitate recall of key points of a previous meeting. With the growing number of meetings at an organization that one may attend weekly and with the growing number of topics discussed, forgetting past meetings becomes eminent, hence recommending certain topics to the user in order to prepare the user for a future meeting is beneficial and important. Our experimental results on real-world data, demonstrate that our methods significantly outperform a state-of-the-art Hidden Markov Model baseline. This indicates the efficacy of our proposed methods for modeling semantics in temporal data. Seyed Ali Bahrainian, Fabio Crestani |
CHIIR | 2 |
| 2018 | In Situ and Context-Aware Target Apps Selection for Unified Mobile SearchabstractWith the recent growth in the use of conversational systems and intelligent assistants such as Google Assistant and Microsoft Cortana, mobile devices are becoming even more pervasive in our lives. As a consequence, users are getting engaged with mobile apps and frequently search for an information need using different apps. Recent work has stated the need for a unified mobile search system that would act as meta search on users' mobile devices: it would identify the target apps for the user's query, submit the query to the apps, and present the results to the user. Moreover, mobile devices provide rich contextual information about users and their whereabouts. In this paper, we introduce the task of context-aware target apps selection as part of a unified mobile search framework. To this aim, we designed an in situ study to collect thousands of mobile queries enriched with mobile sensor data from 255 users during a three month period. With the aid of this dataset, we were able to study user behavior as they performed cross-app search. We finally study the performance of state-of-the-art retrieval models for this task and propose a simple yet effective neural model that significantly outperforms the baselines. Our neural approach is based on learning high-dimensional representations for mobile apps and contextual information. Furthermore, we show that incorporating context improves the performance by 20% in terms of [email protected], enabling the model to perform better for 57% of users. Our data is publicly available for research purposes. Mohammad Aliannejadi, Hamed Zamani, Fabio Crestani, W. Bruce Croft |
CIKM | 3 |
| 2018 | Predicting Topics in Scholarly Papers
Seyed Ali Bahrainian, Ida Mele, Fabio Crestani |
ECIR | 3 |
| 2018 | A Micromodule Approach for Building Real-Time Systems with Python-Based Models: Application to Early Risk Detection of Depression on Social Media
Rodrigo Martínez-Castaño, Juan Carlos Pichel, David E. Losada, Fabio Crestani |
ECIR | 4 |
| 2018 | Emotional Influence Prediction of News Posts
Anastasia Giahanou, Paolo Rosso, Ida Mele, Fabio Crestani |
ICWSM | 4 |
| 2018 | GeoDCF: Deep Collaborative Filtering with Multifaceted Contextual Information in Location-Based Social Networks
Dimitrios Rafailidis, Fabio Crestani |
ECML/PKDD (2) | 2 |
| 2018 | Target Apps Selection: Towards a Unified Search Framework for Mobile DevicesabstractWith the recent growth of conversational systems and intelligent assistants such as Apple Siri and Google Assistant, mobile devices are becoming even more pervasive in our lives. As a consequence, users are getting engaged with the mobile apps and frequently search for an information need in their apps. However, users cannot search within their apps through their intelligent assistants. This requires a unified mobile search framework that identifies the target app(s) for the user's query, submits the query to the app(s), and presents the results to the user. In this paper, we take the first step forward towards developing unified mobile search. In more detail, we introduce and study the task of target apps selection, which has various potential real-world applications. To this aim, we analyze attributes of search queries as well as user behaviors, while searching with different mobile apps. The analyses are done based on thousands of queries that we collected through crowdsourcing. We finally study the performance of state-of-the-art retrieval models for this task and propose two simple yet effective neural models that significantly outperform the baselines. Our neural approaches are based on learning high-dimensional representations for mobile apps. Our analyses and experiments suggest specific future directions in this research area. Mohammad Aliannejadi, Hamed Zamani, Fabio Crestani, W. Bruce Croft |
SIGIR | 3 |
| 2018 | Early Commenting Features for Emotional Reactions Prediction
Anastasia Giahanou, Paolo Rosso, Ida Mele, Fabio Crestani |
SPIRE | 4 |
| 2018 | Personalized Context-Aware Point of Interest RecommendationabstractPersonalized recommendation of Points of Interest (POIs) plays a key role in satisfying users on Location-Based Social Networks (LBSNs). In this article, we propose a probabilistic model to find the mapping between user-annotated tags and locations’ taste keywords. Furthermore, we introduce a dataset on locations’ contextual appropriateness and demonstrate its usefulness in predicting the contextual relevance of locations. We investigate four approaches to use our proposed mapping for addressing the data sparsity problem: one model to reduce the dimensionality of location taste keywords and three models to predict user tags for a new location. Moreover, we present different scores calculated from multiple LBSNs and show how we incorporate new information from the mapping into a POI recommendation approach. Then, the computed scores are integrated using learning to rank techniques. The experiments on two TREC datasets show the effectiveness of our approach, beating state-of-the-art methods. Mohammad Aliannejadi, Fabio Crestani |
ACM Trans. Inf. Syst. | 2 |
| 2017 | Linking News across Multiple Streams for Timeliness AnalysisabstractLinking multiple news streams based on the reported events and analyzing the streams' temporal publishing patterns are two very important tasks for information analysis, discovering newsworthy stories, studying the event evolution, and detecting untrustworthy sources of information. In this paper, we propose techniques for cross-linking news streams based on the reported events with the purpose of analyzing the temporal dependencies among streams. Our research tackles two main issues: (1) how news streams are connected as reporting an event or the evolution of the same event and (2) how timely the newswires report related events using different publishing platforms. Our approach is based on dynamic topic modeling for detecting and tracking events over the timeline and on clustering news according to the events. We leverage the event-based clustering to link news across different streams and present two scoring functions for ranking the streams based on their timeliness in publishing news about a specific event. Ida Mele, Seyed Ali Bahrainian, Fabio Crestani |
CIKM | 3 |
| 2017 | A Collaborative Ranking Model for Cross-Domain RecommendationsabstractWith the advent of social media, generating high quality cross-domain recommendations has become more and more important for users of heterogeneous domains. In this study, we propose a collaborative ranking model to generate cross-domain recommendations. Given a target domain, we design an objective function aimed at performing push of relevant items at the top of a recommendation list. Also, as users may have different behaviours in multiple domains in our collaborative ranking model we propose a weighting strategy to control the influence of user preferences from auxiliary domains when producing the recommendation lists. Our experiments on ten cross-domain recommendation tasks show that the proposed approach achieves higher recommendation accuracy than other state-of-the-art methods. Dimitrios Rafailidis, Fabio Crestani |
CIKM | 2 |
| 2017 | Personalized Keyword Boosting for Venue Suggestion Based on Multiple LBSNs
Mohammad Aliannejadi, Dimitrios Rafailidis, Fabio Crestani |
ECIR | 3 |
| 2017 | Sentiment Propagation for Predicting Reputation Polarity
Anastasia Giahanou, Julio Gonzalo 0001, Ida Mele, Fabio Crestani |
ECIR | 4 |
| 2017 | Multiple Random Walks for Personalized Ranking with Trust and Distrust
Dimitrios Rafailidis, Fabio Crestani |
TPDL | 2 |
| 2017 | Event Detection for Heterogeneous News Streams
Ida Mele, Fabio Crestani |
NLDB | 2 |
| 2017 | A Regularization Method with Inference of Trust and Distrust in Recommender Systems
Dimitrios Rafailidis, Fabio Crestani |
ECML/PKDD (2) | 2 |
| 2017 | Learning to Rank with Trust and Distrust in Recommender SystemsabstractThe sparsity of users' preferences can significantly degrade the quality of recommendations in the collaborative filtering strategy. To account for the fact that the selections of social friends and foes may improve the recommendation accuracy, we propose a learning to rank model that exploits users' trust and distrust relationships. Our learning to rank model focusses on the performance at the top of the list, with the recommended items that end-users will actually see. In our model, we try to push the relevant items of users and their friends at the top of the list, while ranking low those of their foes. Furthermore, we propose a weighting strategy to capture the correlations of users' preferences with friends' trust and foes' distrust degrees in two intermediate trust- and distrust-preference user latent spaces, respectively. Our experiments on the Epinions dataset show that the proposed learning to rank model significantly outperforms other state-of-the-art methods in the presence of sparsity in users' preferences and when a part of trust and distrust relationships is not available. Furthermore, we demonstrate the crucial role of our weighting strategy in our model, to balance well the influences of friends and foes on users' preferences. Dimitrios Rafailidis, Fabio Crestani |
RecSys | 2 |
| 2017 | Venue Appropriateness Prediction for Personalized Context-Aware Venue SuggestionabstractPersonalized context-aware venue suggestion plays a critical role in satisfying the users' needs on location-based social networks (LBSNs). In this paper, we present a set of novel scores to measure the similarity between a user and a candidate venue in a new city. The scores are based on user's history of preferences in other cities as well as user's context. We address the data sparsity problem in venue recommendation with the aid of a proposed approach to predict contextually appropriate places. Furthermore, we show how to incorporate different scores to improve the performance of recommendation. The experimental results of our participation in the TREC 2016 Contextual Suggestion track show that our approach beats state-of-the-art strategies. Mohammad Aliannejadi, Fabio Crestani |
SIGIR | 2 |
| 2017 | A Cross-Platform Collection for Contextual SuggestionabstractSuggesting personalized venues helps users to find interesting places on location-based social networks (LBSNs). Although there are many LBSNs online, none of them is known to have thorough information about all venues. The Contextual Suggestion track at TREC aimed at providing a collection consisting of places as well as user context to enable researchers to examine and compare different approaches, under the same evaluation setting. However, the officially released collection of the track did not meet many participants' needs related to venue content, online reviews, and user context. That is why almost all successful systems chose to crawl information from different LBSNs. For example, one of the best proposed systems in the TREC 2016 Contextual Suggestion track crawled data from multiple LBSNs and enriched it with venue-context appropriateness ratings, collected using a crowdsourcing platform. Such collection enabled the system to better predict a venue's appropriateness to a given user's context. In this paper, we release both collections that were used by the system above. We believe that these datasets give other researchers the opportunity to compare their approaches with the top systems in the track. Also, it provides the opportunity to explore different methods to predicting contextually appropriate venues. Mohammad Aliannejadi, Ida Mele, Fabio Crestani |
SIGIR | 3 |
| 2017 | A Collection for Detecting Triggers of Sentiment SpikesabstractThe advent of social media has given the opportunity to users to publicly express and share their opinion about any topic. Public opinion is very important for the interested entities that can leverage such information in the process of making decisions. In addition, identifying sentiment changes and the likely causes that have triggered them allows interested parties to adjust their strategies and attract more positive sentiment. With the aim to facilitate research on this problem, we describe a collection of tweets that can be used for detecting and ranking the likely triggers of sentiment spikes towards different entities. To build the collection, we first group tweets by topic which are then manually annotated according to sentiment polarity and strength. We believe that this collection can be useful for further research on detecting sentiment change triggers, sentiment analysis and sentiment prediction. Anastasia Giahanou, Ida Mele, Fabio Crestani |
SIGIR | 3 |
| 2017 | Comparative opinion mining: A reviewabstractOpinion mining refers to the use of natural language processing, text analysis, and computational linguistics to identify and extract subjective information in textual material. Opinion mining, also known as sentiment analysis, has received a lot of attention in recent times, as it provides a number of tools to analyze public opinion on a number of different topics. Comparative opinion mining is a subfield of opinion mining which deals with identifying and extracting information that is expressed in a comparative form (e.g., “paper X is better than the Y”). Comparative opinion mining plays a very important role when one tries to evaluate something because it provides a reference point for the comparison. This paper provides a review of the area of comparative opinion mining. It is the first review that cover specifically this topic as all previous reviews dealt mostly with general opinion mining. This survey covers comparative opinion mining from two different angles. One from the perspective of techniques and the other from the perspective of comparative opinion elements. It also incorporates preprocessing tools as well as data set that were used by past researchers that can be useful to future researchers in the field of comparative opinion mining. Kasturi Dewi Varathan, Anastasia Giahanou, Fabio Crestani |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2016 | Network completion via joint node clustering and similarity learningabstractIn this study, we investigate the problem of network completion by considering the similarities between the node attributes. Given a sample of observed nodes with their incident edges, how can we efficiently reconstruct the network by completing the missing edges of unobserved nodes? Apart from the missing edges, in real settings the node attributes may be partially missing, as well as they may introduce noise when completing the network. We propose a network completion method based on joint clustering and similarity learning. The proposed approach differs from competitive strategies, which consider attribute-based similarities at the node-level. First we generate clusters based on the node attributes, thus reducing the noise and the sparsity in the case that the attributes may be missing. We design a joint objective function to jointly factorize the adjacency matrix of the observed edges with the cluster-based similarities of the node attributes, while at the same time the clusters are adapted, accordingly. In addition, we propose an optimization algorithm to solve the network completion problem via alternating minimization. Our experiments on two real world social networks from Facebook and Google+ show that the proposed approach achieves high completion accuracy, compared to other state-of-the-art methods. Dimitrios Rafailidis, Fabio Crestani |
ASONAM | 2 |
| 2016 | Explaining Sentiment Spikes in TwitterabstractTracking public opinion in social media provides important information to enterprises or governments during a decision making process. In addition, identifying and extracting the causes of sentiment spikes allows interested parties to redesign and adjust strategies with the aim to attract more positive sentiments. In this paper, we focus on the problem of tracking sentiment towards different entities, detecting sentiment spikes and on the problem of extracting and ranking the causes of a sentiment spike. Our approach combines LDA topic model with Relative Entropy. The former is used for extracting the topics discussed in the time window before the sentiment spike. The latter allows to rank the detected topics based on their contribution to the sentiment spike. Anastasia Giahanou, Ida Mele, Fabio Crestani |
CIKM | 3 |
| 2016 | Joint Collaborative Ranking with Social Relationships in Top-N RecommendationabstractWith the advent of learning to rank methods, relevant studies showed that Collaborative Ranking (CR) models can produce accurate ranked lists in the top-N recommendation problem. However, in practice several real-world problems decrease their ranking performance, such as the sparsity and cold-start problems, which often occur in recommendation systems for inactive or new users. In this study, to account for the fact that the selections of social friends can improve the recommendation accuracy, we propose a joint CR model based on the users' social relationships. We propose two different CR strategies based on the notions of Social Reverse Height and Social Height, which consider how well the relevant and irrelevant items of users and their social friends have been ranked at the top of the list, respectively. We focus on the top of the list mainly because users see the top-N recommendations in real-world applications, and not the whole ranked list. Furthermore, we formulate a joint objective function to consider both CR strategies, and propose an alternating minimization algorithm to learn our joint CR model. Our experiments on benchmark datasets show that our proposed joint CR model outperforms other state-of-the-art models that either consider social relationships or focus on the ranking performance at the top of the list. Dimitrios Rafailidis, Fabio Crestani |
CIKM | 2 |
| 2016 | Topic-Specific Stylistic Variations for Opinion Retrieval on Twitter
Anastasia Giahanou, Morgan Harvey, Fabio Crestani |
ECIR | 3 |
| 2016 | Top-N Recommendation via Joint Cross-Domain User Clustering and Similarity Learning
Dimitrios Rafailidis, Fabio Crestani |
ECML/PKDD (2) | 2 |
| 2016 | Tracking Sentiment by Time Series AnalysisabstractIn recent years social media have emerged as popular platforms for people to share their thoughts and opinions on all kind of topics. Tracking opinion over time is a powerful tool that can be used for sentiment prediction or to detect the possible reasons of a sentiment change. Understanding topic and sentiment evolution allows enterprises or government to capture negative sentiment and act promptly. In this study, we explore conventional time series analysis methods and their applicability on topic and sentiment trend analysis. We use data collected from Twitter that span over nine months. Finally, we study the usability of outliers detection and different measures such as sentiment velocity and acceleration on the task of sentiment tracking. Anastasia Giahanou, Fabio Crestani |
SIGIR | 2 |
| 2016 | Cluster-based Joint Matrix Factorization Hashing for Cross-Modal RetrievalabstractCross-modal retrieval has been an emerging topic over the last years, as modern applications have to efficiently search for multimedia documents with different modalities. In this study, we propose a cross-modal hashing method by following a cluster-based joint matrix factorization strategy. Our method first builds clusters for each modality separately and then generates a cross-modal cluster representation for each document. We formulate a joint matrix factorization process with the constraint that pushes the documents' representations of the different modalities and the cross-modal cluster representations into a common consensus matrix. In doing so, we capture the inter-modality, intra-modality and cluster-based similarities in a unified latent space. Finally, we present an efficient way to generate the hash codes using the maximum entropy principle and compute the binary codes for external queries. In our experiments with two publicly available data sets, we show that the proposed method outperforms state-of-the-art hashing methods for different cross-modal retrieval tasks. Dimitrios Rafailidis, Fabio Crestani |
SIGIR | 2 |
| 2016 | Collaborative Ranking with Social Relationships for Top-N RecommendationsabstractRecommendation systems have gained a lot of attention because of their importance for handling the unprecedentedly large amount of available content on the Web, such as movies, music, books, etc. Although Collaborative Ranking (CR) models can produce accurate recommendation lists, in practice several real-world problems decrease their ranking performance, such as the sparsity and cold start problems. Here, to account for the fact that the selections of social friends can leverage the recommendation accuracy, we propose SCR, a Social CR model. Our model learns personalized ranking functions collaboratively, using the notion of Social Reverse Height, that is, considering how well the relevant items of users and their social friends have been ranked at the top of the list. The reason that we focus on the top of the list is that users mainly see the top-N recommendations, and not the whole ranked list. In our experiments with a benchmark data set from Epinions, we show that our SCR model performs better than state-of-the-art CR models that either consider social relationships or focus on the ranking performance at the top of the list. Dimitrios Rafailidis, Fabio Crestani |
SIGIR | 2 |
| 2016 | A farewell message from the editor-in-chief
Fabio Crestani |
Inf. Process. Manag. | 1 |
| 2015 | Long Time, No Tweets! Time-aware Personalised Hashtag Suggestion
Morgan Harvey, Fabio Crestani |
ECIR | 2 |
| 2015 | 5th Workshop on Context-Awareness in Retrieval and Recommendation
Ernesto William De Luca, Alan Said, Fabio Crestani, David Elsweiler |
ECIR | 3 |
| 2015 | A geometric framework for data fusion in information retrieval
Shengli Wu 0001, Fabio Crestani |
Inf. Syst. | 2 |
| 2014 | Query-Driven Mining of Citation Networks for Patent Citation Retrieval and RecommendationabstractPrior art search or recommending citations for a patent application is a challenging task. Many approaches have been proposed and shown to be useful for prior art search. However, most of these methods do not consider the network structure for integrating and diffusion of different kinds of information present among tied patents in the citation network. In this paper, we propose a method based on a time-aware random walk on a weighted network of patent citations, the weights of which are characterized by contextual similarity relations between two nodes on the network. The goal of the random walker is to find influential documents in the citation network of a query patent, which can serve as candidates for drawing query terms and bigrams for query refinement. The experimental results on CLEF-IP datasets (CLEF-IP 2010 and CLEF-IP 2011) show the effectiveness of encoding contextual similarities (common classification codes, common inventor, and common applicant) between nodes in the citation network. Our proposed approach can achieve significantly better results in terms of recall and Mean Average Precision rates compared to strong baselines of prior art search. Parvaz Mahdabi, Fabio Crestani |
CIKM | 2 |
| 2014 | Vertical-Aware Click Model-Based Effectiveness MetricsabstractToday's web search systems present users with heterogeneous information coming from sources of different types, also known as verticals. Evaluating such systems is an important but complex task, which is still far from being solved. In this paper we examine the hypothesis that the use of models that capture user search behavior on heterogeneous result pages helps to improve the quality of offline metrics. We propose two vertical-aware metrics based on user click models for federated search and evaluate them using query logs of the Yandex search engine. We show that depending on the type of vertical, the proposed metrics have higher correlation with online user behavior than other state-of-the-art techniques. Ilya Markov, Eugene Kharitonov, Vadim Nikulin, Pavel Serdyukov, Maarten de Rijke, Fabio Crestani |
CIKM | 6 |
| 2014 | A Personalised Recommendation System for Context-Aware Suggestions
Andrei Rikitianskii, Morgan Harvey, Fabio Crestani |
ECIR | 3 |
| 2014 | The effect of citation analysis on query expansion for patent retrieval
Parvaz Mahdabi, Fabio Crestani |
Inf. Retr. | 2 |
| 2014 | Patent Query Formulation by Synthesizing Multiple Sources of Relevance EvidenceabstractPatent prior art search is a task in patent retrieval with the goal of finding documents which describe prior art work related to a query patent. A query patent is a full patent application composed of hundreds of terms which does not represent a single focused information need. Fortunately, other relevance evidence sources (i.e., classification tags and bibliographical data) provide additional details about the underlying information need. In this article, we propose a unified framework that integrates multiple relevance evidence components for query formulation. We first build a query model from the textual fields of a query patent. To overcome the term mismatch, we expand this initial query model with the term distribution of documents in the citation graph, modeling old and recent domain terminology. We build an IPC lexicon and perform query expansion using this lexicon incorporating proximity information. We performed an empirical evaluation on two patent datasets. Our results show that employing the temporal features of documents has a precision enhancing effect, while query expansion using IPC lexicon improves the recall of the final rank list. Parvaz Mahdabi, Fabio Crestani |
ACM Trans. Inf. Syst. | 2 |
| 2014 | Theoretical, Qualitative, and Quantitative Analyses of Small-Document Approaches to Resource SelectionabstractIn a distributed retrieval setup, resource selection is the problem of identifying and ranking relevant sources of information for a given user’s query. For better usage of existing resource-selection techniques, it is desirable to know what the fundamental differences between them are and in what settings one is superior to others. However, little is understood still about the actual behavior of resource-selection methods. In this work, we focus on small-document approaches to resource selection that rank and select sources based on the ranking of their documents. We pose a number of research questions and approach them by three types of analyses. First, we present existing small-document techniques in a unified framework and analyze them theoretically. Second, we propose using a qualitative analysis to study the behavior of different small-document approaches. Third, we present a novel experimental methodology to evaluate small-document techniques and to validate the results of the qualitative analysis. This way, we answer the posed research questions and provide insights about small-document methods in general and about each technique in particular. Ilya Markov, Fabio Crestani |
ACM Trans. Inf. Syst. | 2 |
| 2013 | Building user profiles from topic models for personalised searchabstractPersonalisation is an important area in the field of IR that attempts to adapt ranking algorithms so that the results returned are tuned towards the searcher's interests. In this work we use query logs to build personalised ranking models in which user profiles are constructed based on the representation of clicked documents over a topic space. Instead of employing a human-generated ontology, we use novel latent topic models to determine these topics. Our experiments show that by subtly introducing user profiles as part of the ranking algorithm, rather than by re-ranking an existing list, we can provide personalised ranked lists of documents which improve significantly over a non-personalised baseline. Further examination shows that the performance of the personalised system is particularly good in cases where prior knowledge of the search query is limited. Morgan Harvey, Fabio Crestani, Mark J. Carman |
CIKM | 2 |
| 2013 | Generalizing diversity detection in blog feed retrievalabstractThe goal of a blog retrieval system is to retrieve and rank blogs, as collections of documents, in response to a given query. Previous studies have shown that diversity among the top retrieved posts from a blog is a positive feature for indicating relevance of the blog to the query. However, existing methods capture the diversity of a blog using post-level properties that limits their application to a specific category of retrieval methods. In this paper, we propose a blog-level diversity measure where there is no assumption made about the underlying blog-ranking technique. The proposed measure enables us to integrate diversity in any existing blog retrieval method. Our experimental results show that the proposed method, while being more general, produces comparable results to the post-level diversity detection methods. Mostafa Keikha, Fabio Crestani, W. Bruce Croft |
CIKM | 2 |
| 2013 | Distributed Information Retrieval and Applications
Fabio Crestani, Ilya Markov |
ECIR | 1 |
| 2013 | Reducing the Uncertainty in Resource Selection
Ilya Markov, Leif Azzopardi, Fabio Crestani |
ECIR | 3 |
| 2013 | On CORI Results Merging
Ilya Markov, Avi Arampatzis, Fabio Crestani |
ECIR | 3 |
| 2013 | Leveraging conceptual lexicon: query disambiguation using proximity information for patent retrievalabstractPatent prior art search is a task in patent retrieval where the goal is to rank documents which describe prior art work related to a patent application. One of the main properties of patent retrieval is that the query topic is a full patent application and does not represent a focused information need. This query by document nature of patent retrieval introduces new challenges and requires new investigations specific to this problem. Researchers have addressed this problem by considering different information resources for query reduction and query disambiguation. However, previous work has not fully studied the effect of using proximity information and exploiting domain specific resources for performing query disambiguation. Parvaz Mahdabi, Shima Gerani, Jimmy Huang 0001, Fabio Crestani |
SIGIR | 4 |
| 2013 | The likelihood property in general retrieval operations
Richard Bache, Mark Baillie, Fabio Crestani |
Inf. Sci. | 3 |
| 2012 | Diversity in blog feed retrievalabstractBlog distillation (blog feed retrieval) is a task in blog retrieval where the goal is to rank blogs according to their recurrent relevance to a query topic. One of the main properties of blog feed retrieval is that the unit of retrieval is a collection of documents as opposed to a single document as in other IR tasks. This collection retrieval nature of blog distillation introduces new challenges and requires new investigations specific to this problem. Mostafa Keikha, Fabio Crestani, W. Bruce Croft |
CIKM | 2 |
| 2012 | Score Transformation in Linear Combination for Multi-criteria Relevance Ranking
Shima Gerani, ChengXiang Zhai, Fabio Crestani |
ECIR | 3 |
| 2012 | Automatic refinement of patent queries using concept importance predictorsabstractPatent prior art queries are full patent applications which are much longer than standard web search topics. Such queries are composed of hundreds of terms and do not represent a focused information need. One way to make the queries more focused is to select a group of key terms as representatives. Existing works show that such a selection to reduce patent queries is a challenging task mainly because of the presence of ambiguous terms. Given this setup, we present a query modeling approach where we utilize patent-specific characteristics to generate more precise queries. We propose to automatically disambiguate query terms by employing noun phrases that are extracted using the global analysis of the patent collection. We further introduce a method for predicting whether expansion using noun phrases would improve the retrieval effectiveness. Parvaz Mahdabi, Linda Andersson, Mostafa Keikha, Fabio Crestani |
SIGIR | 4 |
| 2012 | Unsupervised linear score normalization revisitedabstractWe give a fresh look into score normalization for merging result-lists, isolating the problem from other components. We focus on three of the simplest, practical, and widely-used linear methods which do not require any training data, i.e. MinMax, Sum, and Z-Score. We provide theoretical arguments on why and when the methods work, and evaluate them experimentally. We find that MinMax is the most robust under many circumstances, and that Sum is - in contrast to previous literature - the worst. Based on the insights gained, we propose another three simple methods which work as good or better than the baselines. Ilya Markov, Avi Arampatzis, Fabio Crestani |
SIGIR | 3 |
| 2012 | Linguistic aggregation methods in blog retrieval
Mostafa Keikha, Fabio Crestani |
Inf. Process. Manag. | 2 |
| 2012 | Employing document dependency in blog searchabstractAbstract The goal in blog search is to rank blogs according to their recurrent relevance to the topic of the query. State‐of‐the‐art approaches view it as an expert search or resource selection problem. We investigate the effect of content‐based similarity between posts on the performance of the retrieval system. We test two different approaches for smoothing (regularizing) relevance scores of posts based on their dependencies. In the first approach, we smooth term distributions describing posts by performing a random walk over a document‐term graph in which similar posts are highly connected. In the second, we directly smooth scores for posts using a regularization framework that aims to minimize the discrepancy between scores for similar documents. We then extend these approaches to consider the time interval between the posts in smoothing the scores. The idea is that if two posts are temporally close, then they are good sources for smoothing each other's relevance scores. We compare these methods with the state‐of‐the‐art approaches in blog search that employ Language Modeling‐based resource selection algorithms and fusion‐based methods for aggregating post relevance scores. We show performance gains over the baseline techniques which do not take advantage of the relation between posts for smoothing relevance estimates. Mostafa Keikha, Fabio Crestani, Mark J. Carman |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2012 | Aggregation Methods for Proximity-Based Opinion RetrievalabstractThe enormous amount of user-generated data available on the Web provides a great opportunity to understand, analyze, and exploit people’s opinions on different topics. Traditional Information Retrieval methods consider the relevance of documents to a topic but are unable to differentiate between subjective and objective documents. Opinion retrieval is a retrieval task in which not only the relevance of a document to the topic is important but also the amount of opinion expressed in the document about the topic. In this article, we address the blog post opinion retrieval task and propose methods that rank blog posts according to their relevance and opinionatedness toward a topic. We propose estimating the opinion density at each position in a document using a general opinion lexicon and kernel density functions. We propose and investigate different models for aggregating the opinion density at query terms positions to estimate the opinion score of every document. We then combine the opinion score with the relevance score based on a probabilistic justification. Experimental results on the BLOG06 dataset show that the proposed method provides significant improvement over the standard TREC baselines. The proposed models also achieve much higher performance compared to all state of the art methods. Shima Gerani, Mark J. Carman, Fabio Crestani |
ACM Trans. Inf. Syst. | 3 |
| 2011 | Bayesian latent variable models for collaborative item rating predictionabstractCollaborative filtering systems based on ratings make it easier for users to find content of interest on the Web and as such they constitute an area of much research. In this paper we first present a Bayesian latent variable model for rating prediction that models ratings over each user's latent interests and also each item's latent topics. We describe a Gibbs sampling procedure that can be used to estimate its parameters and show by experiment that it is competitive with the gradient descent SVD methods commonly used in state-of-the-art systems. We then proceed to make an important and novel extension to this model, enhancing it with user-dependent and item-dependant biases to significantly improve rating estimation. We show by experiment on a large set of real ratings data that these models are able to outperform 3 common baselines, including a very competitive and modern SVD-based model. Furthermore we illustrate other advantages of our approach beyond simply its ability to provide more accurate ratings and show that it is able to perform better on the common and important case where the user profile is short. Morgan Harvey, Mark J. Carman, Ian Ruthven, Fabio Crestani |
CIKM | 4 |
| 2011 | Predicting document effectiveness in pseudo relevance feedbackabstractPseudo relevance feedback (PRF) is one of effective practices in Information Retrieval. In particular, PRF via the relevance model (RM) has been widely used due to the theoretical soundness and effectiveness. In a PRF scenario, an underlying relevance model is inferred by combining language models of the top retrieved documents where the contribution of each document is assumed to be proportional to its score for the initial query. However, it is not clear that selecting the top retrieved documents only by the initial retrieval scores is actually the optimal way for query expansion. Mostafa Keikha, Jangwon Seo, W. Bruce Croft, Fabio Crestani |
CIKM | 4 |
| 2011 | Personal Blog Retrieval Using Opinion Features
Shima Gerani, Mostafa Keikha, Mark J. Carman, Fabio Crestani |
ECIR | 4 |
| 2011 | TEMPER: A Temporal Relevance Feedback Method
Mostafa Keikha, Shima Gerani, Fabio Crestani |
ECIR | 3 |
| 2011 | Investigating the Statistical Properties of User-Generated Documents
Giacomo Inches, Mark J. Carman, Fabio Crestani |
FQAS | 3 |
| 2011 | Aggregating multiple opinion evidence in proximity-based opinion retrievalabstractBlog post opinion retrieval is the problem of ranking blog posts according to the likelihood that the post is relevant to the query and that the author was expressing an opinion about the topic (of the query). A recent study has proposed a method for finding the opinion density at query term positions in a document which uses the proximity of query term and opinion term as an indicator of their relatedness. The maximum opinion density between different query positions was used as an opinion score of the whole document. In this paper we investigate the effect of exploiting multiple opinion evidence of a document. We propose using the ordered weighted averaging (OWA) operator in order to combine the opinion score of different query positions for a final score of a document, in the proximity-based opinion retrieval system. Shima Gerani, Mostafa Keikha, Fabio Crestani |
SIGIR | 3 |
| 2011 | Time-based relevance modelsabstractThis paper addresses blog feed retrieval where the goal is to retrieve the most relevant blog feeds for a given user query. Since the retrieval unit is a blog, as a collection of posts, performing relevance feedback techniques and selecting the most appropriate documents for query expansion becomes challenging. By assuming time as an effective parameter on the blog posts content, we propose a time-based query expansion method. In this method, we select terms for expansion using most relevant days for the query, as opposed to most relevant documents. This provide us with more trustable terms for expansion. Our preliminary experiments on Blog08 collection shows that this method can outperform state of the art relevance feedback methods in blog retrieval. Mostafa Keikha, Shima Gerani, Fabio Crestani |
SIGIR | 3 |
| 2011 | A multi-collection latent topic model for federated search
Mark Baillie, Mark J. Carman, Fabio Crestani |
Inf. Retr. | 3 |
| 2010 | Towards query log based personalization using topic modelsabstractWe investigate the utility of topic models for the task of personalizing search results based on information present in a large query log. We define generative models that take both the user and the clicked document into account when estimating the probability of query terms. These models can then be used to rank documents by their likelihood given a particular query and user pair. Mark J. Carman, Fabio Crestani, Morgan Harvey, Mark Baillie |
CIKM | 2 |
| 2010 | Statistics of Online User-Generated Short Documents
Giacomo Inches, Mark J. Carman, Fabio Crestani |
ECIR | 3 |
| 2010 | An Approach to Indexing and Clustering News Stories Using Continuous Language Models
Richard Bache, Fabio Crestani |
NLDB | 2 |
| 2010 | Ranking Sequential Patterns with Respect to Significance
Robert Gwadera, Fabio Crestani |
PAKDD (1) | 2 |
| 2010 | Proximity-based opinion retrievalabstractBlog post opinion retrieval aims at finding blog posts that are relevant and opinionated about a user's query. In this paper we propose a simple probabilistic model for assigning relevant opinion scores to documents. The key problem is how to capture opinion expressions in the document, that are related to the query topic. Current solutions enrich general opinion lexicons by finding query-specific opinion lexicons using pseudo-relevance feedback on external corpora or the collection itself. In this paper we use a general opinion lexicon and propose using proximity information in order to capture opinion term relatedness to the query. We propose a proximity-based opinion propagation method to calculate the opinion density at each point in a document. The opinion density at the position of a query term in the document can then be considered as the probability of opinion about the query term at that position. The effect of different kernels for capturing the proximity is also discussed. Experimental results on the BLOG06 dataset show that the proposed method provides significant improvement over standard TREC baselines and achieves a 2.5% increase in MAP over the best performing run in the TREC 2008 blog track. Shima Gerani, Mark J. Carman, Fabio Crestani |
SIGIR | 3 |
| 2010 | A Language Modelling approach to linking criminal styles with offender characteristics
Richard Bache, Fabio Crestani, David Canter, Donna Youngs |
Data Knowl. Eng. | 2 |
| 2009 | Mining and ranking streams of news stories using cross-stream sequential patternsabstractWe present a new method for mining and ranking streams of news stories using cross-stream sequential patterns and content similarity. In particular, we focus on stories reporting the same event across the streams within a given time window, where an event is defined as a specific thing that happens at a specific time and place. For every discovered cluster of stories reporting the same event we create an itemset-sequence consisting of stream identifiers of the stories in the cluster, where the sequence is ordered according to the timestamps of the stories. Furthermore, we record exact timestamps and content similarities between the respective stories. Given such a collection of itemset-sequences we use it for two tasks: (I) to discover recurrent temporal publishing patterns between the news streams in terms of frequent sequential patterns and content similarity and (II) to rank the streams of news stories with respect to timeliness of reporting important events and content authority. We demonstrate the applicability of the presented method on a multi-stream of news stories was gathered from RSS feeds of major world news agencies. Robert Gwadera, Fabio Crestani |
CIKM | 2 |
| 2009 | A Topic-Based Measure of Resource Description Quality for Distributed Information Retrieval
Mark Baillie, Mark J. Carman, Fabio Crestani |
ECIR | 3 |
| 2009 | Investigating Learning Approaches for Blog Post Opinion Retrieval
Shima Gerani, Mark J. Carman, Fabio Crestani |
ECIR | 3 |
| 2009 | Effectiveness of Aggregation Methods in Blog Distillation
Mostafa Keikha, Fabio Crestani |
FQAS | 2 |
| 2009 | Design of an Interface for Interactive Topic Detection and Tracking
Masnizah Mohd, Fabio Crestani, Ian Ruthven |
FQAS | 2 |
| 2009 | A statistical comparison of tag and query logsabstractWe investigate tag and query logs to see if the terms people use to annotate websites are similar to the ones they use to query for them. Over a set of URLs, we compare the distribution of tags used to annotate each URL with the distribution of query terms for clicks on the same URL. Understanding the relationship between the distributions is important to determine how useful tag data may be for improving search results and conversely, query data for improving tag prediction. In our study, we compare both term frequency distributions using vocabulary overlap and relative entropy. We also test statistically whether the term counts come from the same underlying distribution. Our results indicate that the vocabulary used for tagging and searching for content are similar but not identical. We further investigate the content of the websites to see which of the two distributions (tag or query) is most similar to the content of the annotated/searched URL. Finally, we analyze the similarity for different categories of URLs in our sample to see if the similarity between distributions is dependent on the topic of the website or the popularity of the URL. Mark J. Carman, Mark Baillie, Robert Gwadera, Fabio Crestani |
SIGIR | 4 |
| 2009 | Blog distillation using random walksabstractThis paper addresses the blog distillation problem. That is, given a user query find the blogs most related to the query topic. We model the blogosphere as a single graph that includes extra information besides the content of the posts. By performing a random walk on this graph we extract most relevant blogs for each query. Our experiments on the TREC'07 data set show 15% improvement in MAP and 8% improvement in [email protected] over the Language Modeling baseline. Mostafa Keikha, Mark J. Carman, Fabio Crestani |
SIGIR | 3 |
| 2009 | Measuring the likelihood property of scoring functions in general retrieval modelsabstractAbstract Although retrieval systems based on probabilistic models will rank the objects (e.g., documents) being retrieved according to the probability of some matching criterion (e.g., relevance), they rarely yield an actual probability, and the scoring function is interpreted to be purely ordinal within a given retrieval task. In this brief communication, it is shown that some scoring functions possess the likelihood property, which means that the scoring function indicates the likelihood of matching when compared to other retrieval tasks, which is potentially more useful than pure ranking although it cannot be interpreted as an actual probability. This property can be detected by using two modified effectiveness measures: entire precision and entire recall. Richard Bache, Mark Baillie, Fabio Crestani |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2008 | Estimating real-valued characteristics of criminals from their recorded crimesabstractOffender profiling concerns making inferences about a criminal from the crime(s) he has committed. Where descriptionsof the crimes are recorded electronically, text mining techniques provide a means by which recorded characteristics of the offenders can be linked with features of his crimes as revealed in the text. Past studies have used Language Modelling to identify characteristics that can be described by a categorical variable e.g. gender. Here we adapt the Language Modelling approach to allow estimation of numerical quantities such as age and distance travelled. Richard Bache, Fabio Crestani |
CIKM | 2 |
| 2008 | A Comparison of Named Entity Patterns from a User Analysis and a System Analysis
Masnizah Mohd, Fabio Crestani, Ian Ruthven |
ECIR | 2 |
| 2008 | Discovering Significant Patterns in Multi-stream SequencesabstractDiscovering significant patterns in synchronized multi-stream sequences also known as multi-attribute event sequences (multi-sequences), is an important problem in many domains, including monitoring systems and information retrieval. In this paper we propose a new approach for assessing significance of multi-stream patterns in multi-attribute event sequences. In experiments on physiological multi-stream data we show applicability of our method. Robert Gwadera, Fabio Crestani |
ICDM | 2 |
| 2008 | A Language Modelling Approach to Linking Criminal Styles with Offender Characteristics
Richard Bache, Fabio Crestani, David Canter, Donna Youngs |
NLDB | 2 |
| 2008 | Towards personalized distributed information retrievalabstractOur aim is to investigate if and how the performance of Distributed Information Retrieval (DIR) systems can be improved through personalization. Toward this aim we are building a testbed of document collections and corresponding personalized relevance judgments. In this paper we discuss our intended approach for personalizing the three different phases of the DIR process. We also describe the test collection we are building and discuss our methodology for evaluating personalized DIR using relevance information taken from social bookmarking data. Mark J. Carman, Fabio Crestani |
SIGIR | 2 |
| 2008 | 'Show me more': Incremental length summarisation using novelty detection
Simon O. Sweeney, Fabio Crestani, David E. Losada |
Inf. Process. Manag. | 2 |
| 2008 | Preface
Fabio Crestani, Paolo Ferragina, Mark Sanderson |
Inf. Retr. | 1 |
| 2008 | Metadata harvesting for content-based distributed information retrievalabstractAbstract We propose an approach to content‐based Distributed Information Retrieval based on the periodic and incremental centralization of full‐content indices of widely dispersed and autonomously managed document sources. Inspired by the success of the Open Archive Initiative's (OAI) Protocol for metadata harvesting, the approach occupies middle ground between content crawling and distributed retrieval. As in crawling, some data move toward the retrieval process, but it is statistics about the content rather than content itself; this grants more efficient use of network resources and wider scope of application. As in distributed retrieval, some processing is distributed along with the data, but it is indexing rather than retrieval; this reduces the costs of content provision while promoting the simplicity, effectiveness, and responsiveness of retrieval. Overall, we argue that the approach retains the good properties of centralized retrieval without renouncing to cost‐effective, large‐scale resource pooling. We discuss the requirements associated with the approach and identify two strategies to deploy it on top of the OAI infrastructure. In particular, we define a minimal extension of the OAI protocol which supports the coordinated harvesting of full‐content indices and descriptive metadata for content resources. Finally, we report on the implementation of a proof‐of‐concept prototype service for multimodel content‐based retrieval of distributed file collections. Fabio Simeoni, Murat Yakici, Steve Neely, Fabio Crestani |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2007 | Language models, probability of relevance and relevance likelihoodabstractThis paper proposes a measure of relevance likelihood derived specifically for language models. Such a measure may be used to guide a user on how far to browse through the list of retrieved items or for pseudo-relevance feedback. To derive this measure, it is necessary to make the assumption that a user is seeking an ideal (usually non-existent) document and the actual relevant documents in the collection will contain fragments of this ideal document. Thus, in deriving this measure we propose a novel way of capturing relevance in Language Modelling. Richard Bache, Mark Baillie, Fabio Crestani |
CIKM | 3 |
| 2007 | Summarisation and Novelty: An Experimental Investigation
Simon O. Sweeney, Fabio Crestani, David E. Losada |
ECIR | 2 |
| 2007 | Investigation of the Effectiveness of Cross-Media Indexing
Murat Yakici, Fabio Crestani |
ECIR | 2 |
| 2007 | The DILIGENT framework for distributed information retrievalabstractNo abstract available. Fabio Simeoni, Fabio Crestani, Ralf Bierig |
SIGIR | 2 |
| 2007 | Modelling epistemic uncertainty in ir evaluationabstractModern information retrieval (IR) test collections violate the completeness assumption of the Cranfield paradigm. In order to maximise the available resources, only a sample of documents (i.e. the pool) are judged for relevance by a human assessor(s). The subsequent evaluation protocol does not make any distinctions between assessed or unassesseddocuments, as documents that are not in the pool are assumedto be not relevant for the topic. This is beneficial from a practical point of view, as the relative performance can be compared with confidence if the experimental conditions are fair for all systems. However, given the incompleteness of relevance assessments, two forms of uncertainty emerge during evaluation. The first is Aleatory uncertainty, which refers to variation in system performance across the topic set, which is often addressed through the use of statistical significance tests. The second form of uncertainty is Epistemic, which refers to the amount of knowledge (or ignorance) we have about the estimate of a system's performance. Epistemic uncertainty is a consequence of incompleteness and is not addressed by the current evaluation protocol. In this study, we present a first attempt at modelling both aleatory and epistemic uncertainty associatedwith IR evaluation. We aim to account for both the variability associated with system performance and the amount of knowledge known about the performance estimate. Murat Yakici, Mark Baillie, Ian Ruthven, Fabio Crestani |
SIGIR | 4 |
| 2007 | Introduction to special issue on contextual information retrieval systems
Fabio Crestani, Ian Ruthven |
Inf. Retr. | 1 |
| 2006 | Adaptive query-based sampling for distributed IRabstractNo abstract available. Leif Azzopardi, Mark Baillie, Fabio Crestani |
SIGIR | 3 |
| 2006 | PENG: integrated search of distributed news archivesabstractNo abstract available. Mark Baillie, Fabio Crestani, Monica Landoni |
SIGIR | 2 |
| 2006 | Adaptive Query-Based Sampling of Distributed Collections
Mark Baillie, Leif Azzopardi, Fabio Crestani |
SPIRE | 3 |
| 2006 | Testing the cluster hypothesis in distributed information retrieval
Fabio Crestani, Shengli Wu 0001 |
Inf. Process. Manag. | 1 |
| 2006 | Effective search results summary size and device screen size: Is there a relationship?
Simon O. Sweeney, Fabio Crestani |
Inf. Process. Manag. | 2 |
| 2006 | Written versus spoken queries: A qualitative and quantitative comparative analysisabstractAbstract The authors report on an experimental study on the differences between spoken and written queries. A set of written and spontaneous spoken queries are generated by users from written topics. These two sets of queries are compared in qualitative terms and in terms of their retrieval effectiveness. Written and spoken queries are compared in terms of length, duration, and part of speech. In addition, assuming perfect transcription of the spoken queries, written and spoken queries are compared in terms of their aptitude to describe relevant documents. The retrieval effectiveness of spoken and written queries is compared using three different information retrieval models. The results show that using speech to formulate one's information need provides a way to express it more naturally and encourages the formulation of longer queries. Despite that, longer spoken queries do not seem to significantly improve retrieval effectiveness compared with written queries. Fabio Crestani, Heather Du |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2004 | Retrieval Effectiveness of Written and Spoken Queries: An Experimental Evaluation
Heather Du, Fabio Crestani |
FQAS | 2 |
| 2004 | A graphical user interface for the retrieval of hierarchically structured documents
Fabio Crestani, Jesús Vegas, Pablo de la Fuente |
Inf. Process. Manag. | 1 |
| 2003 | WebDocBall: A Graphical Visualization Tool for Web Search Results
Jesús Vegas, Pablo de la Fuente, Fabio Crestani |
ECIR | 3 |
| 2003 | Experiments with Document Archive Size Detection
Shengli Wu 0001, Forbes Gibb, Fabio Crestani |
ECIR | 3 |
| 2003 | Resource selection and data fusion in multimedia distributed digital librariesabstractCallan, Jamie (ed.): SIGIR '03: Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval. New York: ACM, 2003, 363–364. - ISBN: 978-1-58113-646-3 Online also available at: https://doi.org/10.1145/860435.860502 Jamie Callan, Fabio Crestani, Henrik Nottelmann, Pietro Pala, Xiao Mang Shou |
SIGIR | 2 |
| 2003 | Ranking Structured Documents Using Utility Theory in the Bayesian Network Retrieval Model
Fabio Crestani, Luis M. de Campos, Juan M. Fernández-Luna, Juan F. Huete |
SPIRE | 1 |
| 2003 | Handling vagueness, subjectivity, and imprecision in information access: an introduction to the special issue
Fabio Crestani, Gabriella Pasi |
Inf. Process. Manag. | 1 |
| 2003 | Automatic construction of hypertexts for self-referencing: the Hyper-TextBook project
Fabio Crestani, Massimo Melucci |
Inf. Syst. | 1 |
| 2003 | Mathematical, Logical and Formal Methods in Information Retrieval: An Introduction to the Specia IssueabstractAbstract Research on the use of mathematical, logical, and formal methods, has been central to Information Retrieval research for long time. Research in this area is important not only because it helps enhancing retrieval effectiveness, but also because it helps clarifying the underlying concepts of Information Retrieval. In this article we outline some of the major aspects of the subject, and summarize the papers of this special issue with respect to how they relate to these aspects. We conclude by highlighting some directions of future research, which are needed to better understand the formal characteristics of Information Retrieval. Fabio Crestani, Sándor Dominich, Mounia Lalmas-Roelleke, C. J. van Rijsbergen |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2003 | Vocal Access to a Newspaper Archive: Assessing the Limitations of Current Voice Information Access Technology
Fabio Crestani |
J. Intell. Inf. Syst. | 1 |
| 2002 | Data fusion with estimated weightsabstractThis paper proposes an adptive approach for data fusion of information retrieval systems, which exploits estimated performances of all component input systems without relevance judgement or training. The estimation is conducted prior to the fusion but uses the same data as fusion applies. The experiment shows that our algorithms are competitive with, and often outperform CombMNZ, one of the most effective algorithms in use. Shengli Wu 0001, Fabio Crestani |
CIKM | 2 |
| 2002 | A Graphical User Interface for Structured Document Retrieval
Jesús Vegas, Pablo de la Fuente, Fabio Crestani |
ECIR | 3 |
| 2002 | Experimenting with graphical user interfaces for structured document retrievalabstractNo abstract available. Fabio Crestani, Pablo de la Fuente, Jesús Vegas |
SIGIR | 1 |
| 2002 | Spoken query processing for interactive information retrieval
Fabio Crestani |
Data Knowl. Eng. | 1 |
| 2001 | Towards the use of Prosodic Information for Spoken Document RetrievalabstractNo abstract available. Fabio Crestani |
SIGIR | 1 |
| 2001 | Design of a Graphical User Interface for Structured Documents RetrievalabstractMany document collections contain documents that have significant structure. Structured document retrieval requires different models and interfaces from standard Information Retrieval. An Information Retrieval system dealing with structured documents has to enable a user to query, browse retrieved documents, and provide query refinement and relevance feedback based not only on full documents but also on specific parts of them, according to their structure. Currently, very few IR systems enable such level of flexibility and interaction, because of limitations in indexing and retrieval models and in interfaces. In this paper, we present the design of a new graphical user interface for structured document retrieval. This interface provides the user with an intuitive and yet powerful set of tools for structured document searching, retrieved list navigation, and search refinement. 1 Fabio Crestani, Pablo de la Fuente, Jesús Vegas |
SPIRE | 1 |
| 2000 | Word Recognition Errors and Relevance Feedback in Spoken Query ProcessingabstractGiven the typical length of queries submitted to Information Retrieval systems, it is easy to imagine that the effects of word recognition errors (WRE) in spoken queries must be severely destructive on the system’s effectiveness. The effects of word recognition errors in Spoken Document Retrieval have been well studied and well reported in recent Information Retrieval literature, but much less experimental work has been devoted to studying the effects of word recognition errors in Spoken Query Processing (SQP). The experimental work reported in this paper shows that the use of classical IR techniques for SQP is quite robust to considerably high levels of WRE, in particular for long queries. Moreover, in the case of short queries, both standard relevance feedback and pseudo relevance feedback can be effectively employed to improve the effectiveness of SQP. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Fabio Crestani |
FQAS | 1 |
| 2000 | Prosodic Stress and Topic Detection in Spoken SentencesabstractThe relationship between acoustic stress and the information content of words is investigated. On the one hand, the average acoustic stress is measured for each word throughout each utterance. On the other hand, an information retrieval (IR) index is calculated, based on the word frequency throughout the particular spoken sentence and throughout the collection of analysed spoken sentences. The scatter plots of the two measures (average acoustic stress on the y-axis and IR index on the x-axis) show higher values of average acoustic stress with increasing IR index of the word in the majority of the analysed utterances. A statistically more valid proof of such a relationship is derived from a histogram of the words with high average acoustic stress vs. the IR index. This confirms that a word with high average acoustic stress also has a high value of the IR index and, if we trust IR indexes, also a high information content. Rosaria Silipo, Fabio Crestani |
SPIRE | 2 |
| 2000 | Searching the web by constrained spreading activation
Fabio Crestani, Puay Leng Lee |
Inf. Process. Manag. | 1 |
| 2000 | Exploiting the Similarity of Non-Matching Terms at Retrieval Time
Fabio Crestani |
Inf. Retr. | 1 |
| 2000 | Users' perception of relevance of spoken documentsabstractWe present the results of a study of user's perception of relevance of documents. The aim is to study experimentally how users' perception varies depending on the form that retrieved documents are presented. Documents retrieved in response to a query are presented to users in a variety of ways, from full text to a machine spoken query-biased automatically-generated summary, and the difference in users' perception of relevance is studied. The experimental results suggest that the effectiveness of advanced multimedia Information Retrieval applications may be affected by the low level of users' perception of relevance of retrieved documents. Anastasios Tombros, Fabio Crestani |
J. Am. Soc. Inf. Sci. | 2 |
| 1999 | Probabilistic learning for selective dissemination of information
Gianni Amati, Fabio Crestani |
Inf. Process. Manag. | 2 |
| 1998 | A Case study of Automatic Authoring: From a Textbook to a Hyper-Textbook
Fabio Crestani, Massimo Melucci |
Data Knowl. Eng. | 1 |
| 1998 | A Study of Probability Kinematics in Information RetrievalabstractWe analyze the kinematics of probabilistic term weights at retrieval time for different Information Retrieval models. We present four models based on different notions of probabilistic retrieval. Two of these models are based on classical probability theory and can be considered as prototypes of models long in use in Information Retrieval, like the Vector Space Model and the Probabilistic Model. The two other models are based on a logical technique of evaluating the probability of a conditional called imaging; one is a generalization of the other. We analyze the transfer of probabilities occurring in the term space at retrieval time for these four models, compare their retrieval performance using classical test collections, and discuss the results. We believe that our results provide useful suggestions on how to improve existing probabilistic models of Information Retrieval by taking into consideration term-term similarity. Fabio Crestani, C. J. van Rijsbergen |
ACM Trans. Inf. Syst. | 1 |
| 1997 | On the Use of Information Retrieval Techniques for the Automatic Construction of Hypertext
Maristella Agosti, Fabio Crestani, Massimo Melucci |
Inf. Process. Manag. | 2 |
| 1997 | A Model for Adaptive Information Retrieval
Fabio Crestani, C. J. van Rijsbergen |
J. Intell. Inf. Syst. | 1 |
| 1996 | Design and Implementation of a Tool for the Automatic Construction of Hypertexts for Information Retrieval
Maristella Agosti, Fabio Crestani, Massimo Melucci |
Inf. Process. Manag. | 2 |
| 1995 | Probability Kinematics in Information RetrievalabstractIn this paper we discuss the dynamics of probabilistic term Fabio Crestani, C. J. van Rijsbergen |
SIGIR | 1 |