Ricardo Campos 0001

dblp:38/2963-1 · also Ricardo Nuno Taborda Campos · DBLP profile ↗
← Back
58ranked-venue papers in the field
22as first author
34since 2021 · last 2026
0000-0002-8767-8126ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 53 (19 first)Other / Interdisciplinary · 3 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 2 (1 first)
YearPublicationVenuePosition
2026 MiNER: A Two-Stage Pipeline for Metadata Extraction from Municipal Meeting Minutes
Rodrigo Batista, Luís Filipe Cunha, Purificação Silvano, Nuno Guimarães, Alípio Mário Jorge, Evelin Amorim, Ricardo Campos 0001
ECIR (2)7
2026 The 9th International Workshop on Narrative Extraction from Text: Text2Story 2026
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak
ECIR (3)1
2026 CitiLink-Minutes: A Multilayer Annotated Dataset of Municipal Meeting Minutes
Ricardo Campos 0001, Ana Filipa Pacheco, Ana Luísa Fernandes, Inês Cantante, Rute Rebouças, Luís Filipe Cunha, José Isidro, José Pedro Evans, Miguel Marques, Rodrigo Batista, Evelin Amorim, Alípio Mário Jorge, Nuno Guimarães, Sérgio Nunes 0001, Antonio Leal-Millán, Purificação Silvano
ECIR (4)1
2026 ClaimPT: A Portuguese Dataset of Annotated Claims in News Articles
Ricardo Campos 0001, Raquel Sequeira, Sara Nerea, Inês Cantante, Diogo Folques, Luís Filipe Cunha, João Canavilhas, António Branco, Alípio Mário Jorge, Sérgio Nunes 0001, Nuno Guimarães, Purificação Silvano
ECIR (4)1
2026 pt-image-ir-dataset: An Image Retrieval Dataset in European Portuguese
Rodrigo Duarte, António Branco, Hugo Proença 0001, Ricardo Campos 0001
ECIR (4)4
2026 ImageSeek: A Hybrid Text-to-Image Image Retrieval System for Domain-Specific Collections
Rodrigo Duarte, António Branco, Hugo Proença 0001, Ricardo Campos 0001
ECIR (4)5
2026 CitiLink: Enhancing Municipal Transparency and Citizen Engagement Through Searchable Meeting Minutes
José Pedro Evans, José Isidro, Miguel Marques, Afonso Fonseca, Ricardo Morais, João Canavilhas, Arian Pasquali, Purificação Silvano, Alípio Mário Jorge, Nuno Guimarães, Sérgio Nunes 0001, Ricardo Campos 0001
ECIR (4)13
2026 Looking for the Bottleneck in Fine-grained Temporal Relation Classification
abstract
Temporal relation classification is the task of determining the temporal relation between pairs of temporal entities in a text. Despite recent advancements in natural language processing, temporal relation classification remains a considerable challenge. Early attempts framed this task using a comprehensive set of temporal relations between events and temporal expressions. However, due to the task complexity, datasets have been progressively simplified, leading recent approaches to focus on the relations between event pairs and to use only a subset of relations. In this work, we revisit the broader goal of classifying interval relations between temporal entities by considering the full set of relations that can hold between two time intervals. The proposed approach, Interval from Point, involves first classifying the point relations between the endpoints of the temporal entities and then decoding these point relations into an interval relation. Evaluation on the TempEval-3 dataset shows that this approach can yield effective results, achieving a temporal awareness score of 70.1 percent, a new state-of-the-art on this benchmark.
Hugo Sousa, Ricardo Campos 0001, Alipio Jorge
SIGIR2
2026 CitiLink-Summ: A Dataset of Discussion Subjects Summaries in European Portuguese Municipal Meeting Minutes
abstract
Municipal meeting minutes are formal records documenting the discussions and decisions of local government, yet their content is often lengthy, dense, and difficult for citizens to navigate. Automatic summarization can help address this challenge by producing concise summaries for each discussion subject. Despite its potential, research on summarizing discussion subjects in municipal meeting minutes remains largely unexplored, especially in low-resource languages, where the inherent complexity of these documents adds further challenges. A major bottleneck is the scarcity of datasets containing high-quality, manually crafted summaries, which limits the development and evaluation of effective summarization models for this domain. In this paper, we present CitiLink-Summ, a new corpus of European Portuguese municipal meeting minutes, comprising 120 documents and 2,880 manually hand-written summaries, each corresponding to a distinct discussion subject. Leveraging this dataset, we establish baseline results for automatic summarization in this domain, employing state-of-the-art generative models (e.g., BART, PRIMERA) as well as large language models (LLMs), evaluated with both lexical and semantic metrics such as ROUGE, BLEU, METEOR, and BERTScore. CitiLink-Summ provides the first benchmark for municipal-domain summarization in European Portuguese, offering a valuable resource for advancing NLP research on complex administrative texts.
Miguel Marques, Ana Luísa Fernandes, Ana Filipa Pacheco, Rute Rebouças, Inês Cantante, José Isidro, Luís Filipe Cunha, Alípio Mário Jorge, Nuno Guimarães, Sérgio Nunes 0001, António Leal, Purificação Silvano, Ricardo Campos 0001
WWW13
2025 The Temporal Game: A New Perspective on Temporal Relation Extraction
abstract
In this paper we demo the Temporal Game, a novel approach to temporal relation extraction that casts the task as an interactive game. Instead of directly annotating interval-level relations, our approach decomposes them into point-wise comparisons between the start and end points of temporal entities. At each step, players classify a single point relation, and the system applies temporal closure to infer additional relations and enforce consistency. This point-based strategy naturally supports both interval and instant entities, enabling more fine-grained and flexible annotation than any previous approach. The Temporal Game also lays the groundwork for training reinforcement learning agents, by treating temporal annotation as a sequential decision-making task. To showcase this potential, the demo presented in this paper includes a Game mode, in which users annotate texts from the TempEval-3 dataset and receive feedback based on a scoring system, and an Annotation mode, that allows custom documents to be annotated and resulting timeline to be exported. Therefore, this demo serves both as a research tool and an annotation interface. The demo is publicly available at https://temporal-game.inesctec.pt, and the source code is open-sourced to foster further research and community-driven development in temporal reasoning and annotation.
Hugo O. Sousa, Ricardo Campos 0001, Alípio Mário Jorge
CIKM2
2025 The 8th International Workshop on Narrative Extraction from Texts: Text2Story 2025
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak
ECIR (5)1
2025 MedLink: Retrieval and Ranking of Case Reports to Assist Clinical Decision Making
Luís Filipe Cunha, Nuno Guimarães, Alexandra Mendes, Ricardo Campos 0001, Alípio Mário Jorge
ECIR (5)4
2025 Leveraging LLMs to Improve Human Annotation Efficiency with INCEpTION
Luís Filipe Cunha, Nana Yu, Purificação Silvano, Ricardo Campos 0001, Alípio Mário Jorge
ECIR (5)4
2025 CLEF 2025 JOKER Lab: Humour in the Machine
Liana Ermakova, Anne-Gwenn Bosser, Tristan Miller, Ricardo Campos 0001
ECIR (5)4
2025 Rebuilding the Past: Reconstructing Portuguese News Outlets with Web Archives
Ricardo Campos 0001
ECIR (5)2
2025 ICDAR 2025 Competition on Automatic Classification of Literary Epochs
Irina Rabaev, Marina Litvak, Roza Bass, Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt
ICDAR (5)4
2024 Physio: An LLM-Based Physiotherapy Advisor
Rúben Almeida, Hugo O. Sousa, Luís Filipe Cunha, Nuno Guimarães, Ricardo Campos 0001, Alípio Mário Jorge
ECIR (5)5
2024 The 7th International Workshop on Narrative Extraction from Texts: Text2Story 2024
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak
ECIR (5)1
2024 ACE-2005-PT: Corpus for Event Extraction in Portuguese
abstract
Event extraction is an NLP task that commonly involves identifying the central word (trigger) for an event and its associated arguments in text. ACE-2005 is widely recognised as the standard corpus in this field. While other corpora, like PropBank, primarily focus on annotating predicate-argument structure, ACE-2005 provides comprehensive information about the overall event structure and semantics. However, its limited language coverage restricts its usability. This paper introduces ACE-2005-PT, a corpus created by translating ACE-2005 into Portuguese, with European and Brazilian variants. To speed up the process of obtaining ACE-2005-PT, we rely on automatic translators. This, however, poses some challenges related to automatically identifying the correct alignments between multi-word annotations in the original text and in the corresponding translated sentence. To achieve this, we developed an alignment pipeline that incorporates several alignment techniques: lemmatization, fuzzy matching, synonym matching, multiple translations and a BERT-based word aligner. To measure the alignment effectiveness, a subset of annotations from the ACE-2005-PT corpus was manually aligned by a linguist expert. This subset was then compared against our pipeline results which achieved exact and relaxed match scores of 70.55% and 87.55% respectively. As a result, we successfully generated a Portuguese version of the ACE-2005 corpus, which has been accepted for publication by LDC.
Luís Filipe Cunha, Purificação Silvano, Ricardo Campos 0001, Alípio Mário Jorge
SIGIR3
2024 Keywords attention for fake news detection using few positive labels
Mariana Caravanti de Souza, Marcos P. S. Gôlo, Alípio Mário Jorge, Evelin Amorim, Ricardo Campos 0001, Ricardo M. Marcacini, Solange Oliveira Rezende
Inf. Sci.5
2023 Contrastive Keyword Extraction from Versioned Documents
abstract
Versioned documents are common in many situations and play a vital part in numerous applications enabling an overview of the revisions made to a document or document collection. However, as documents increase in size, it gets difficult to summarize and comprehend all the changes made to versioned documents. In this paper, we propose a novel research problem of contrastive keyword extraction from versioned documents, and introduce an unsupervised approach that extracts keywords to reflect the key changes made to an earlier document version. In order to provide an easy-to-use comparison and summarization tool, an open-source demonstration is made available which can be found at https://contrastive-keyword-extraction.streamlit.app/
Lukas Eder, Ricardo Campos 0001, Adam Jatowt
CIKM2
2023 TEI2GO: A Multilingual Approach for Fast Temporal Expression Identification
abstract
Temporal expression identification is crucial for understanding texts written in natural language. Although highly effective systems such as HeidelTime exist, their limited runtime performance hampers adoption in large-scale applications and production environments. In this paper, we introduce the TEI2GO models, matching HeidelTime's effectiveness but with significantly improved runtime, supporting six languages, and achieving state-of-the-art results in four of them. To train the TEI2GO models, we used a combination of manually annotated reference corpus and developed ``Professor HeidelTime'', a comprehensive weakly labeled corpus of news texts annotated with HeidelTime. This corpus comprises a total of $138,069$ documents (over six languages) with $1,050,921$ temporal expressions, the largest open-source annotated dataset for temporal expression identification to date. By describing how the models were produced, we aim to encourage the research community to further explore, refine, and extend the set of models to additional languages and domains. Code, annotations, and models are openly available for community exploration and use. The models are conveniently on HuggingFace for seamless integration and application.
Hugo O. Sousa, Ricardo Campos 0001, Alípio Mário Jorge
CIKM2
2023 Public News Archive: A Searchable Sub-archive to Portuguese Past News Articles
Ricardo Campos 0001, Diogo Correia, Adam Jatowt
ECIR (3)1
2023 The 6th International Workshop on Narrative Extraction from Texts: Text2Story 2023
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak
ECIR (3)1
2023 TweetStream2Story: Narrative Extraction from Tweets in Real Time
Mafalda Castro, Alípio Mário Jorge, Ricardo Campos 0001
ECIR (3)3
2023 Text2Storyline: Generating Enriched Storylines from Text
Francisco Gonçalves, Ricardo Campos 0001, Alípio Mário Jorge
ECIR (3)2
2023 The 1st International Workshop on Implicit Author Characterization from Texts for Search and Retrieval (IACT'23)
abstract
The first edition of the Implicit Author Characterization from Texts for Search and Retrieval (IACT'23) aims at bringing to the forefront the challenges involved in identifying and extracting from texts implicit information about authors (e.g., human or AI) and using it in IR tasks. The IACT workshop provides a common forum to consolidate multi-disciplinary efforts and foster discussions to identify the wide-ranging issues related to the task of extracting implicit author-related information from the textual content, including novel tasks and datasets. We will also discuss the ethical implications of implicit information extraction. In addition, we announce a shared task focused on automatically determining the literary epochs of written books.
Marina Litvak, Irina Rabaev, Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt
SIGIR3
2023 tieval: An Evaluation Framework for Temporal Information Extraction Systems
abstract
Temporal information extraction (TIE) has attracted a great deal of interest over the last two decades. Such endeavors have led to the development of a significant number of datasets. Despite its benefits, having access to a large volume of corpora makes it difficult to benchmark TIE systems. On the one hand, different datasets have different annotation schemes, which hinders the comparison between competitors across different corpora. On the other hand, the fact that each corpus is disseminated in a different format requires a considerable engineering effort for a researcher/practitioner to develop parsers for all of them. These constraints force researchers to select a limited amount of datasets to evaluate their systems which consequently limits the comparability of the systems. Yet another obstacle to the comparability of TIE systems is the evaluation metric employed. While most research works adopt traditional metrics such as precision, recall, and F1, a few others prefer temporal awareness -- a metric tailored to be more comprehensive on the evaluation of temporal systems. Although the reason for the absence of temporal awareness in the evaluation of most systems is not clear, one of the factors that certainly weighs on this decision is the need to implement the temporal closure algorithm, which is neither straightforward to implement nor easily available. All in all, these problems have limited the fair comparison between approaches and consequently, the development of TIE systems. To mitigate these problems, we have developed tieval, a Python library that provides a concise interface for importing different corpora and is equipped with domain-specific operations that facilitate system evaluation. In this paper, we present the first public release of tieval and highlight its most relevant features. The library is available as open source, under MIT License, at PyPI and GitHub.
Hugo O. Sousa, Ricardo Campos 0001, Alípio Mário Jorge
SIGIR2
2022 Tweet2Story: A Web App to Extract Narratives from Twitter
Vasco Campos, Ricardo Campos 0001, Pedro Mota, Alípio Mário Jorge
ECIR (2)2
2022 The 5th International Workshop on Narrative Extraction from Texts: Text2Story 2022
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak
ECIR (2)1
2021 Time-Matters: Temporal Unfolding of Texts
Ricardo Campos 0001, Jorge Duque, Tiago Cândido, Gaël Dias, Alípio Mário Jorge, Celia Nunes
ECIR (2)1
2021 The 4th International Workshop on Narrative Extraction from Texts: Text2Story 2021
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Mark A. Finlayson
ECIR (2)1
2021 Exploding TV Sets and Disappointing Laptops: Suggesting Interesting Content in News Archives Based on Surprise Estimation
Adam Jatowt, I-Chen Hung, Michael Färber 0001, Ricardo Campos 0001, Masatoshi Yoshikawa
ECIR (1)4
2021 TLS-Covid19: A New Annotated Corpus for Timeline Summarization
Arian Pasquali, Ricardo Campos 0001, Alexandre Ribeiro, Brenda Salenave Santana, Alípio Mário Jorge, Adam Jatowt
ECIR (1)2
2020 The 3rd International Workshop on Narrative Extraction from Texts: Text2Story 2020
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia
ECIR (2)1
2020 YAKE! Keyword extraction from single documents using multiple local features
abstract
As the amount of generated information grows, reading and summarizing texts of large collections turns into a challenging task. Many documents do not come with descriptive terms, thus requiring humans to generate keywords on-the-fly. The need to automate this kind of task demands the development of keyword extraction systems with the ability to automatically identify keywords within the text. One approach is to resort to machine-learning algorithms. These, however, depend on large annotated text corpora, which are not always available. An alternative solution is to consider an unsupervised approach. In this article, we describe YAKE!, a light-weight unsupervised automatic keyword extraction method which rests on statistical text features extracted from single documents to select the most relevant keywords of a text. Our system does not need to be trained on a particular set of documents, nor does it depend on dictionaries, external corpora, text size, language, or domain. To demonstrate the merits and significance of YAKE!, we compare it against ten state-of-the-art unsupervised approaches and one supervised method. Experimental results carried out on top of twenty datasets show that YAKE! significantly outperforms other unsupervised methods on texts of different sizes, languages, and domains.
Ricardo Campos 0001, Vítor Mangaravite, Arian Pasquali, Alípio Mário Jorge, Celia Nunes, Adam Jatowt
Inf. Sci.1
2019 Document in Context of its Time (DICT): Providing Temporal Context to Support Analysis of Past Documents
abstract
Old documents tend to be difficult to be analyzed and understood, not only for average users but oftentimes for professionals as well. This is due to the context shift, vocabulary evolution and, in general, the lack of precise knowledge about the writing styles in the past. We propose a concept of positioning document in the context of its time, and develop an interactive system to support such an objective. Our system helps users to know whether the vocabulary used by an author in the past were frequent at the time of text creation, whether the author used anachronisms or neologisms, and so on. It also enables detecting terms in text that underwent considerable semantic change and provides more information on the nature of such change. Overall, the proposed tool offers additional knowledge on the writing style and vocabulary choice in documents by drawing from data collected at the time of their creation or at other user-specified time.
Adam Jatowt, Ricardo Campos 0001, Sourav S. Bhowmick, Antoine Doucet
CIKM2
2019 The 2nd International Workshop on Narrative Extraction from Text: Text2Story 2019
Alípio Mário Jorge, Ricardo Campos 0001, Adam Jatowt, Sumit Bhatia
ECIR (2)2
2019 Interactive System for Automatically Generating Temporal Narratives
Arian Pasquali, Vítor Mangaravite, Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt
ECIR (2)3
2019 Information Processing & Management Journal Special Issue on Narrative Extraction from Texts (Text2Story): Preface
Alípio Mário Jorge, Ricardo Campos 0001, Adam Jatowt, Sérgio Nunes 0001
Inf. Process. Manag.2
2018 Every Word has its History: Interactive Exploration and Visualization of Word Sense Evolution
abstract
Human language constantly evolves due to the changing world and the need for easier forms of expression and communication. Our knowledge of language evolution is however still fragmentary despite significant interest of both researchers as well as wider public in the evolution of language. In this paper, we present an interactive framework that permits users study the evolution of words and concepts. The system we propose offers a rich online interface allowing arbitrary queries and complex analytics over large scale historical textual data, letting users investigate changes in meaning, context and word relationships across time.
Adam Jatowt, Ricardo Campos 0001, Sourav S. Bhowmick, Nina Tahmasebi, Antoine Doucet
CIKM2
2018 A Text Feature Based Automatic Keyword Extraction Method for Single Documents
Ricardo Campos 0001, Vítor Mangaravite, Arian Pasquali, Alípio Mário Jorge, Celia Nunes, Adam Jatowt
ECIR1
2018 YAKE! Collection-Independent Automatic Keyword Extractor
Ricardo Campos 0001, Vítor Mangaravite, Arian Pasquali, Alípio Mário Jorge, Celia Nunes, Adam Jatowt
ECIR1
2018 ParsTime: Rule-Based Extraction and Normalization of Persian Temporal Expressions
Behrooz Mansouri, Mohammad Sadegh Zahedi, Ricardo Campos 0001, Mojgan Farhoodi, Masoud Rahgozar
ECIR3
2018 Online Job Search: Study of Users' Search Behavior using Search Engine Query Logs
abstract
Over the last few years, an increasing number of user's and enterprises on the internet has generated a global marketplace for both employers and job seekers. Despite the fact that online job search is now more preferable than traditional methods - leading to better matches between the job seekers and the employer's intents - there is still little insight into how online job searches are different from general web searches. In this paper, we explore the different characteristics of online job search and their differences with general searches, by leveraging search engine query logs. Our experimental results show that job searches have specific attributes which can be used by search engines to increase the quality of the search results.
Behrooz Mansouri, Mohammad Sadegh Zahedi, Ricardo Campos 0001, Mojgan Farhoodi
SIGIR3
2017 Interactive System for Reasoning about Document Age
abstract
Recently, many historical texts have become digitized and made accessible for search and browsing. Professionals who work with collections of such texts often need to verify the correctness of documents' key metadata - their creation dates. In this paper, we demonstrate an interactive system for estimating the age of documents. It may be useful not only for tagging a large number of undated documents, but also for verifying already known timestamps. In order to infer probable dates, we rely on a large scale lexical corpora, Google Books Ngrams. Besides estimating the document creation year, the system also outputs evidences to support age detection and reasoning process and allows testing different hypotheses about document's age.
Adam Jatowt, Ricardo Campos 0001
CIKM2
2017 Learning Temporal Ambiguity in Web Search Queries
abstract
Time has strong influence on web search. The temporal intent of the searcher adds an important dimension to the relevance judgments of web queries. However, lack of understanding their temporal requirements increases the ambiguity of the queries, turning retrieval effectiveness improvements into a complex task. In this paper, we propose an approach to classify web queries into four different categories considering their temporal ambiguity. For each query, we develop features from its search volumes and related queries using Google trends and its related top Wikipedia pages. Our experiment results show that these features can determine temporal ambiguity of a given query with high accuracy. We have demonstrated that a Multilayer Perceptron Networks can achieve better results in classifying temporal class of queries in comparison to other classifiers.
Behrooz Mansouri, Mohammad Sadegh Zahedi, Masoud Rahgozar, Farhad Oroumchian, Ricardo Campos 0001
CIKM5
2017 Identifying top relevant dates for implicit time sensitive queries
Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes
Inf. Retr. J.1
2016 First International Workshop on Recent Trends in News Information Retrieval (NewsIR'16)
Miguel Martinez-Alvarez, Udo Kruschwitz, Gabriella Kazai, Frank Hopfgartner, David P. A. Corney, Ricardo Campos 0001, M-Dyaa Albakour
ECIR6
2016 GTE-Rank: A time-aware search engine to answer time-sensitive queries
Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes
Inf. Process. Manag.1
2015 Time and information retrieval: Introduction to the special issue
Leon Derczynski, Jannik Strötgen, Ricardo Campos 0001, Omar Alonso
Inf. Process. Manag.3
2014 GTE-Rank: Searching for Implicit Temporal Query Results
abstract
Temporal information retrieval has been a topic of great interest in recent years. Despite the efforts that have been conducted so far, most popular search engines remain underdeveloped when it comes to explicitly considering the use of temporal information in their search process. In this paper we present GTE-Rank, an online searching tool that takes time into account when ranking time-sensitive query web search results. GTE-Rank is defined as a linear combination of topical and temporal scores to reflect the relevance of any web page both in topical and temporal dimensions. The resulting system can be explored graphically through a search interface made available for research purposes.
Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes
CIKM1
2014 GTE-Cluster: A Temporal Search Interface for Implicit Temporal Queries
Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes
ECIR1
2012 GTE: a distributional second-order co-occurrence approach to improve the identification of top relevant dates in web snippets
abstract
In this paper, we present an approach to identify top relevant dates in Web snippets with respect to a given implicit temporal query. Our approach is two-fold. First, we propose a generic temporal similarity measure called GTE, which evaluates the temporal similarity between a query and a date. Second, we propose a classification model to accurately relate relevant dates to their corresponding query terms and withdraw irrelevant ones. We suggest two different solutions: a threshold-based classification strategy and a supervised classifier based on a combination of multiple similarity measures. We evaluate both strategies over a set of real-world text queries and compare the performance of our Web snippet approach with a query log approach over the same set of queries. Experiments show that determining the most relevant dates of any given implicit temporal query can be improved with GTE combined with the second order similarity measure InfoSimba, the Dice coefficient and the threshold-based strategy compared to (1) first-order similarity measures and (2) the query log based approach.
Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes
CIKM1
2012 Temporal Web Image Retrieval
Gaël Dias, José G. Moreno 0001, Adam Jatowt, Ricardo Campos 0001
SPIRE4
2012 Disambiguating Implicit Temporal Queries by Clustering Top Relevant Dates in Web Snippets
abstract
With the growing popularity of research in Temporal Information Retrieval (T-IR), a large amount of temporal data is ready to be exploited. The ability to exploit this information can be potentially useful for several tasks. For example, when querying "Football World Cup Germany", it would be interesting to have two separate clusters {1974,2006} corresponding to each of the two temporal instances. However, clustering of search results by time is a non-trivial task that involves determining the most relevant dates associated to a query. In this paper, we propose a first approach to flat temporal clustering of search results. We rely on a second order co-occurrence similarity measure approach which first identifies top relevant dates. Documents are grouped at the year level, forming the temporal instances of the query. Experimental tests were performed using real-world text queries. We used several measures for evaluating the performance of the system and compared our approach with Carrot Web-snippet clustering engine. Both experiments were complemented with a user survey.
Ricardo Campos 0001, Alípio Mário Jorge, Gaël Dias, Celia Nunes
Web Intelligence1
2011 Using k-Top retrieved web snippets to date temporalimplicit queries based on web content analysis
abstract
The World Wide Web (WWW) is a huge information network from which retrieving and organizing quality relevant content remains an open question for mostly all ambiguous queries. As an example, many queries have temporal implicit intents associated with them but they are not inferred by search engines. Inferring the user intentions and the period he has in mind, may therefore play an extremely important role in the improvement of the results. Our work goes in this direction. We aim to introduce a temporal analysis framework for analyzing documents in a temporal dimension in order to identify and understand the temporal nature of any given query, namely implicit ones. Our analysis is not based on metadata, but on the exploitation of temporal information from the content itself, particularly within web snippets, which are interesting pieces of concentrated information, where time clues, especially years, often appear. Our intention is to develop a language-independent solution and to model the degree of relationship between the terms and dates identified. This is the core part of the framework and the basis for both temporal query understanding and search results exploration, such as temporal clustering. We believe that inferring this knowledge is a very important step in the process of adding a temporal dimension to IR systems, thus disambiguating a large class of queries for which search engines continue to fail.
Ricardo Campos 0001
SIGIR1
2006 WISE: Hierarchical Soft Clustering of Web Page Search Results Based on Web Content Mining Techniques
abstract
Typically, search engines are low precision in response to a query, retrieving lots of useless Web pages, and missing some other important ones. In this paper, we study the problem of the hierarchical clustering of Web pages search results. In particular, we propose an architecture called WISE, a meta-search engine that automatically builds clusters of related Web pages embodying one meaning of the query. These clusters are then hierarchically organized and labeled with a phrase representing the key concept of the cluster and the corresponding Web documents. The system which is a Web-based interface (soon available at wise.di.ubi.pt), introduces some interesting new ideas, such as the preselection of the retrieved Web pages, the capacity to statistically detect phrases within documents and the representation of documents based on their most relevant key concepts by using Web content mining techniques. The final step of the system is supported by a graph-based overlapping clustering algorithm which groups the selected documents into a hierarchy of clusters
Ricardo Campos 0001, Gaël Dias, Celia Nunes
Web Intelligence1