VLDB 2026 Research / reviewers in the wild / expert
Alípio Mário Jorge
dblp:j/AlipioMarioJorge · also Alípio Jorge, Alípio M. Jorge
· DBLP profile ↗
72ranked-venue papers in the field
8as first author
34since 2021 · last 2026
0000-0002-5475-1382ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 43 (2 first)Data Mining & Knowledge Discovery · 14 (5 first)Other / Interdisciplinary · 7Database Systems & Data Management · 5 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MiNER: A Two-Stage Pipeline for Metadata Extraction from Municipal Meeting Minutes
Rodrigo Batista, Luís Filipe Cunha, Purificação Silvano, Nuno Guimarães, Alípio Mário Jorge, Evelin Amorim, Ricardo Campos 0001 |
ECIR (2) | 5 |
| 2026 | The 9th International Workshop on Narrative Extraction from Text: Text2Story 2026
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak |
ECIR (3) | 2 |
| 2026 | CitiLink-Minutes: A Multilayer Annotated Dataset of Municipal Meeting Minutes
Ricardo Campos 0001, Ana Filipa Pacheco, Ana Luísa Fernandes, Inês Cantante, Rute Rebouças, Luís Filipe Cunha, José Isidro, José Pedro Evans, Miguel Marques, Rodrigo Batista, Evelin Amorim, Alípio Mário Jorge, Nuno Guimarães, Sérgio Nunes 0001, Antonio Leal-Millán, Purificação Silvano |
ECIR (4) | 12 |
| 2026 | ClaimPT: A Portuguese Dataset of Annotated Claims in News Articles
Ricardo Campos 0001, Raquel Sequeira, Sara Nerea, Inês Cantante, Diogo Folques, Luís Filipe Cunha, João Canavilhas, António Branco, Alípio Mário Jorge, Sérgio Nunes 0001, Nuno Guimarães, Purificação Silvano |
ECIR (4) | 9 |
| 2026 | CitiLink: Enhancing Municipal Transparency and Citizen Engagement Through Searchable Meeting Minutes
José Pedro Evans, José Isidro, Miguel Marques, Afonso Fonseca, Ricardo Morais, João Canavilhas, Arian Pasquali, Purificação Silvano, Alípio Mário Jorge, Nuno Guimarães, Sérgio Nunes 0001, Ricardo Campos 0001 |
ECIR (4) | 10 |
| 2026 | CitiLink-Summ: A Dataset of Discussion Subjects Summaries in European Portuguese Municipal Meeting MinutesabstractMunicipal meeting minutes are formal records documenting the discussions and decisions of local government, yet their content is often lengthy, dense, and difficult for citizens to navigate. Automatic summarization can help address this challenge by producing concise summaries for each discussion subject. Despite its potential, research on summarizing discussion subjects in municipal meeting minutes remains largely unexplored, especially in low-resource languages, where the inherent complexity of these documents adds further challenges. A major bottleneck is the scarcity of datasets containing high-quality, manually crafted summaries, which limits the development and evaluation of effective summarization models for this domain. In this paper, we present CitiLink-Summ, a new corpus of European Portuguese municipal meeting minutes, comprising 120 documents and 2,880 manually hand-written summaries, each corresponding to a distinct discussion subject. Leveraging this dataset, we establish baseline results for automatic summarization in this domain, employing state-of-the-art generative models (e.g., BART, PRIMERA) as well as large language models (LLMs), evaluated with both lexical and semantic metrics such as ROUGE, BLEU, METEOR, and BERTScore. CitiLink-Summ provides the first benchmark for municipal-domain summarization in European Portuguese, offering a valuable resource for advancing NLP research on complex administrative texts. Miguel Marques, Ana Luísa Fernandes, Ana Filipa Pacheco, Rute Rebouças, Inês Cantante, José Isidro, Luís Filipe Cunha, Alípio Mário Jorge, Nuno Guimarães, Sérgio Nunes 0001, António Leal, Purificação Silvano, Ricardo Campos 0001 |
WWW | 8 |
| 2025 | Can LLMs Reliably Label YouTube Videos? A Committee-Based Evaluation
Adriano Mourthé, Carlos Eduardo Ribeiro de Mello, Alípio Mário Jorge |
ASONAM (1) | 3 |
| 2025 | The Temporal Game: A New Perspective on Temporal Relation ExtractionabstractIn this paper we demo the Temporal Game, a novel approach to temporal relation extraction that casts the task as an interactive game. Instead of directly annotating interval-level relations, our approach decomposes them into point-wise comparisons between the start and end points of temporal entities. At each step, players classify a single point relation, and the system applies temporal closure to infer additional relations and enforce consistency. This point-based strategy naturally supports both interval and instant entities, enabling more fine-grained and flexible annotation than any previous approach. The Temporal Game also lays the groundwork for training reinforcement learning agents, by treating temporal annotation as a sequential decision-making task. To showcase this potential, the demo presented in this paper includes a Game mode, in which users annotate texts from the TempEval-3 dataset and receive feedback based on a scoring system, and an Annotation mode, that allows custom documents to be annotated and resulting timeline to be exported. Therefore, this demo serves both as a research tool and an annotation interface. The demo is publicly available at https://temporal-game.inesctec.pt, and the source code is open-sourced to foster further research and community-driven development in temporal reasoning and annotation. Hugo O. Sousa, Ricardo Campos 0001, Alípio Mário Jorge |
CIKM | 3 |
| 2025 | The 8th International Workshop on Narrative Extraction from Texts: Text2Story 2025
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak |
ECIR (5) | 2 |
| 2025 | MedLink: Retrieval and Ranking of Case Reports to Assist Clinical Decision Making
Luís Filipe Cunha, Nuno Guimarães, Alexandra Mendes, Ricardo Campos 0001, Alípio Mário Jorge |
ECIR (5) | 5 |
| 2025 | Leveraging LLMs to Improve Human Annotation Efficiency with INCEpTION
Luís Filipe Cunha, Nana Yu, Purificação Silvano, Ricardo Campos 0001, Alípio Mário Jorge |
ECIR (5) | 5 |
| 2025 | ICDAR 2025 Competition on Automatic Classification of Literary Epochs
Irina Rabaev, Marina Litvak, Roza Bass, Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt |
ICDAR (5) | 5 |
| 2024 | Physio: An LLM-Based Physiotherapy Advisor
Rúben Almeida, Hugo O. Sousa, Luís Filipe Cunha, Nuno Guimarães, Ricardo Campos 0001, Alípio Mário Jorge |
ECIR (5) | 6 |
| 2024 | The 7th International Workshop on Narrative Extraction from Texts: Text2Story 2024
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak |
ECIR (5) | 2 |
| 2024 | ACE-2005-PT: Corpus for Event Extraction in PortugueseabstractEvent extraction is an NLP task that commonly involves identifying the central word (trigger) for an event and its associated arguments in text. ACE-2005 is widely recognised as the standard corpus in this field. While other corpora, like PropBank, primarily focus on annotating predicate-argument structure, ACE-2005 provides comprehensive information about the overall event structure and semantics. However, its limited language coverage restricts its usability. This paper introduces ACE-2005-PT, a corpus created by translating ACE-2005 into Portuguese, with European and Brazilian variants. To speed up the process of obtaining ACE-2005-PT, we rely on automatic translators. This, however, poses some challenges related to automatically identifying the correct alignments between multi-word annotations in the original text and in the corresponding translated sentence. To achieve this, we developed an alignment pipeline that incorporates several alignment techniques: lemmatization, fuzzy matching, synonym matching, multiple translations and a BERT-based word aligner. To measure the alignment effectiveness, a subset of annotations from the ACE-2005-PT corpus was manually aligned by a linguist expert. This subset was then compared against our pipeline results which achieved exact and relaxed match scores of 70.55% and 87.55% respectively. As a result, we successfully generated a Portuguese version of the ACE-2005 corpus, which has been accepted for publication by LDC. Luís Filipe Cunha, Purificação Silvano, Ricardo Campos 0001, Alípio Mário Jorge |
SIGIR | 4 |
| 2024 | Keywords attention for fake news detection using few positive labels
Mariana Caravanti de Souza, Marcos P. S. Gôlo, Alípio Mário Jorge, Evelin Amorim, Ricardo Campos 0001, Ricardo M. Marcacini, Solange Oliveira Rezende |
Inf. Sci. | 3 |
| 2023 | TEI2GO: A Multilingual Approach for Fast Temporal Expression IdentificationabstractTemporal expression identification is crucial for understanding texts written in natural language. Although highly effective systems such as HeidelTime exist, their limited runtime performance hampers adoption in large-scale applications and production environments. In this paper, we introduce the TEI2GO models, matching HeidelTime's effectiveness but with significantly improved runtime, supporting six languages, and achieving state-of-the-art results in four of them. To train the TEI2GO models, we used a combination of manually annotated reference corpus and developed ``Professor HeidelTime'', a comprehensive weakly labeled corpus of news texts annotated with HeidelTime. This corpus comprises a total of $138,069$ documents (over six languages) with $1,050,921$ temporal expressions, the largest open-source annotated dataset for temporal expression identification to date. By describing how the models were produced, we aim to encourage the research community to further explore, refine, and extend the set of models to additional languages and domains. Code, annotations, and models are openly available for community exploration and use. The models are conveniently on HuggingFace for seamless integration and application. Hugo O. Sousa, Ricardo Campos 0001, Alípio Mário Jorge |
CIKM | 3 |
| 2023 | The 6th International Workshop on Narrative Extraction from Texts: Text2Story 2023
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak |
ECIR (3) | 2 |
| 2023 | TweetStream2Story: Narrative Extraction from Tweets in Real Time
Mafalda Castro, Alípio Mário Jorge, Ricardo Campos 0001 |
ECIR (3) | 2 |
| 2023 | Text2Storyline: Generating Enriched Storylines from Text
Francisco Gonçalves, Ricardo Campos 0001, Alípio Mário Jorge |
ECIR (3) | 3 |
| 2023 | ORSUM 2023 - 6th Workshop on Online Recommender Systems and User ModelingabstractModern online platforms for user modeling and recommendation require complex data infrastructures to collect and process data. Some of this data has to be kept to later be used in batches to train personalization models. However, since user activity data can be generated at very fast rates it is also useful to have algorithms able to process data streams online, in real time. Given the continuous and potentially fast change of content, context and user preferences or intents, stream-based models, and their synchronization with batch models can be extremely challenging. Therefore, it is important to investigate methods able to transparently and continuously adapt to the inherent dynamics of user interactions, preferably over long periods of time. Models able to continuously learn from such flows of data are gaining attention in the recommender systems community, and are being increasingly deployed in online platforms. However, many challenges associated with learning from streams need further investigation. João Vinagre, Marie Al-Ghossein, Ladislav Peska, Alípio Mário Jorge, Albert Bifet |
RecSys | 4 |
| 2023 | The 1st International Workshop on Implicit Author Characterization from Texts for Search and Retrieval (IACT'23)abstractThe first edition of the Implicit Author Characterization from Texts for Search and Retrieval (IACT'23) aims at bringing to the forefront the challenges involved in identifying and extracting from texts implicit information about authors (e.g., human or AI) and using it in IR tasks. The IACT workshop provides a common forum to consolidate multi-disciplinary efforts and foster discussions to identify the wide-ranging issues related to the task of extracting implicit author-related information from the textual content, including novel tasks and datasets. We will also discuss the ethical implications of implicit information extraction. In addition, we announce a shared task focused on automatically determining the literary epochs of written books. Marina Litvak, Irina Rabaev, Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt |
SIGIR | 4 |
| 2023 | tieval: An Evaluation Framework for Temporal Information Extraction SystemsabstractTemporal information extraction (TIE) has attracted a great deal of interest over the last two decades. Such endeavors have led to the development of a significant number of datasets. Despite its benefits, having access to a large volume of corpora makes it difficult to benchmark TIE systems. On the one hand, different datasets have different annotation schemes, which hinders the comparison between competitors across different corpora. On the other hand, the fact that each corpus is disseminated in a different format requires a considerable engineering effort for a researcher/practitioner to develop parsers for all of them. These constraints force researchers to select a limited amount of datasets to evaluate their systems which consequently limits the comparability of the systems. Yet another obstacle to the comparability of TIE systems is the evaluation metric employed. While most research works adopt traditional metrics such as precision, recall, and F1, a few others prefer temporal awareness -- a metric tailored to be more comprehensive on the evaluation of temporal systems. Although the reason for the absence of temporal awareness in the evaluation of most systems is not clear, one of the factors that certainly weighs on this decision is the need to implement the temporal closure algorithm, which is neither straightforward to implement nor easily available. All in all, these problems have limited the fair comparison between approaches and consequently, the development of TIE systems. To mitigate these problems, we have developed tieval, a Python library that provides a concise interface for importing different corpora and is equipped with domain-specific operations that facilitate system evaluation. In this paper, we present the first public release of tieval and highlight its most relevant features. The library is available as open source, under MIT License, at PyPI and GitHub. Hugo O. Sousa, Ricardo Campos 0001, Alípio Mário Jorge |
SIGIR | 3 |
| 2022 | Tweet2Story: A Web App to Extract Narratives from Twitter
Vasco Campos, Ricardo Campos 0001, Pedro Mota, Alípio Mário Jorge |
ECIR (2) | 4 |
| 2022 | The 5th International Workshop on Narrative Extraction from Texts: Text2Story 2022
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak |
ECIR (2) | 2 |
| 2022 | ORSUM 2022 - 5th Workshop on Online Recommender Systems and User ModelingabstractModern online systems for user modeling and recommendation need to continuously deal with complex data streams generated by users at very fast rates. This can be overwhelming for systems and algorithms designed to train recommendation models in batches, given the continuous and potentially fast change of content, context and user preferences or intents. Therefore, it is important to investigate methods able to transparently and continuously adapt to the inherent dynamics of user interactions, preferably for long periods of time. Online models that continuously learn from such flows of data are gaining attention in the recommender systems community, given their natural ability to deal with data generated in dynamic, complex environments. User modeling and personalization can particularly benefit from algorithms capable of maintaining models incrementally and online. João Vinagre, Marie Al-Ghossein, Alípio Mário Jorge, Albert Bifet, Ladislav Peska |
RecSys | 3 |
| 2021 | Improving Portuguese Semantic Role Labeling with Transformers and Transfer LearningabstractThe Natural Language Processing task of determining “Who did what to whom” is called Semantic Role Labeling. For English, recent methods based on Transformer models have allowed for major improvements in this task over the previous state of the art. However, for low resource languages, like Portuguese, currently available semantic role labeling models are hindered by scarce training data. In this paper, we explore a model architecture with only a pre-trained Transformer-based model, a linear layer, softmax and Viterbi decoding. We substantially improve the state-of-the-art performance in Portuguese by over 15 F1. Additionally, we improve semantic role labeling results in Portuguese corpora by exploiting cross-lingual transfer learning using multilingual pre-trained models, and transfer learning from dependency parsing in Portuguese, evaluating the various proposed approaches empirically. Sofia Oliveira, Daniel Loureiro, Alípio Mário Jorge |
DSAA | 3 |
| 2021 | Time-Matters: Temporal Unfolding of Texts
Ricardo Campos 0001, Jorge Duque, Tiago Cândido, Gaël Dias, Alípio Mário Jorge, Celia Nunes |
ECIR (2) | 6 |
| 2021 | The 4th International Workshop on Narrative Extraction from Texts: Text2Story 2021
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Mark A. Finlayson |
ECIR (2) | 2 |
| 2021 | TLS-Covid19: A New Annotated Corpus for Timeline Summarization
Arian Pasquali, Ricardo Campos 0001, Alexandre Ribeiro, Brenda Salenave Santana, Alípio Mário Jorge, Adam Jatowt |
ECIR (1) | 5 |
| 2021 | Partially Monotonic Learning for Neural Networks
Joana Trindade, João Vinagre, Kelwin Fernandes, Nuno Paiva, Alípio Mário Jorge |
IDA | 5 |
| 2021 | ORSUM 2021 - 4th Workshop on Online Recommender Systems and User ModelingabstractModern online services continuously generate data at very fast rates. This continuous flow of data encompasses content – e.g. posts, news, products, comments –, but also user feedback – e.g. ratings, views, reads, clicks –, together with context data – user device, spacial or temporal data, user task or activity, weather. This can be overwhelming for systems and algorithms designed to train in batches, given the continuous and potentially fast change of content, context and user preferences or intents. Therefore, it is important to investigate online methods able to transparently adapt to the inherent dynamics of online services. Incremental models that learn from data streams are gaining attention in the recommender systems community, given their natural ability to deal with the continuous flows of data generated in dynamic, complex environments. User modeling and personalization can particularly benefit from algorithms capable of maintaining models incrementally and online. João Vinagre, Alípio Mário Jorge, Marie Al-Ghossein, Albert Bifet |
RecSys | 2 |
| 2021 | A Hybrid Recommender System for Improving Automatic Playlist ContinuationabstractAlthough widely used, the majority of current music recommender systems still focus on recommendations' accuracy, user preferences and isolated item characteristics, without evaluating other important factors, like the joint item selections and the recommendation moment. However, when it comes to playlist recommendations, additional dimensions, as well as the notion of user experience and perception, should be taken into account to improve recommendations' quality. In this work, HybA, a hybrid recommender system for automatic playlist continuation, that combines Latent Dirichlet Allocation and Case-Based Reasoning, is proposed. This system aims to address “similar concepts” rather than similar users. More than generating a playlist based on user requirements, like automatic playlist generation methods, HybA identifies the semantic characteristics of a started playlist and reuses the most similar past ones, to recommend relevant playlist continuations. In addition, support to beyond accuracy dimensions, like increased coherence or diverse items' discovery, is provided. To overcome the semantic gap between music descriptions and user preferences, identify playlist structures and capture songs' similarity, a graph model is used. Experiments on real datasets have shown that the proposed algorithm is able to outperform other state of the art techniques, in terms of accuracy, while balancing between diversity and coherence. Anna Gatzioura, João Vinagre, Alípio Mário Jorge, Miquel Sànchez-Marrè |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | Statistically Robust Evaluation of Stream-Based Recommender SystemsabstractOnline incremental models for recommendation are nowadays pervasive in both the industry and the academia. However, there is not yet a standard evaluation methodology for the algorithms that maintain such models. Moreover, online evaluation methodologies available in the literature generally fall short on the statistical validation of results, since this validation is not trivially applicable to stream-based algorithms. We propose ak-fold validation framework for the pairwise comparison of recommendation algorithms that learn from user feedback streams, using prequential evaluation. Our proposal enables continuous statistical testing on adaptive-size sliding windows over the outcome of the prequential process, allowing practitioners and researchers to make decisions in real time based on solid statistical evidence. We present a set of experiments to gain insights on the sensitivity and robustness of two statistical tests-McNemar's and Wilcoxon signed rank-in a streaming data environment. Our results show that besides allowing a real-time, fine-grained online assessment, the online versions of the statistical tests are at least as robust as the batch versions, and definitely more robust than a simple prequential single-fold approach. João Vinagre, Alípio Mário Jorge, Conceição Rocha, João Gama 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | The 3rd International Workshop on Narrative Extraction from Texts: Text2Story 2020
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia |
ECIR (2) | 2 |
| 2020 | MedLinker: Medical Entity Linking with Neural Representations and Dictionary Matching
Daniel Loureiro, Alípio Mário Jorge |
ECIR (2) | 2 |
| 2020 | Incremental Approach for Automatic Generation of Domain-Specific Sentiment Lexicon
Shamsuddeen Hassan Muhammad, Pavel Brazdil, Alípio Mário Jorge |
ECIR (2) | 3 |
| 2020 | ORSUM - Workshop on Online Recommender Systems and User ModelingabstractModern online web-based systems continuously generate data at very fast rates. This continuous flow of data encompasses web content – e.g. posts, news, products, comments –, but also user feedback – e.g. ratings, views, reads, clicks, thumbs up –, as well as context information – device used, geographic info, social network, current user activity, weather. This is potentially overwhelming for systems and algorithms design to train in offline batches, given the continuous and potentially fast change of content, context and user preferences. Therefore it is important to investigate online methods to be able to transparently adapt to the inherent dynamics of online systems. Incremental models that learn from data streams are gaining attention in the recommender systems community, given their natural ability to deal with data generated in dynamic, complex environments. User modeling and personalization can particularly benefit from algorithms capable of maintaining models incrementally and online, as data is generated. João Vinagre, Alípio Mário Jorge, Marie Al-Ghossein, Albert Bifet |
RecSys | 2 |
| 2020 | YAKE! Keyword extraction from single documents using multiple local featuresabstractAs the amount of generated information grows, reading and summarizing texts of large collections turns into a challenging task. Many documents do not come with descriptive terms, thus requiring humans to generate keywords on-the-fly. The need to automate this kind of task demands the development of keyword extraction systems with the ability to automatically identify keywords within the text. One approach is to resort to machine-learning algorithms. These, however, depend on large annotated text corpora, which are not always available. An alternative solution is to consider an unsupervised approach. In this article, we describe YAKE!, a light-weight unsupervised automatic keyword extraction method which rests on statistical text features extracted from single documents to select the most relevant keywords of a text. Our system does not need to be trained on a particular set of documents, nor does it depend on dictionaries, external corpora, text size, language, or domain. To demonstrate the merits and significance of YAKE!, we compare it against ten state-of-the-art unsupervised approaches and one supervised method. Experimental results carried out on top of twenty datasets show that YAKE! significantly outperforms other unsupervised methods on texts of different sizes, languages, and domains. Ricardo Campos 0001, Vítor Mangaravite, Arian Pasquali, Alípio Mário Jorge, Celia Nunes, Adam Jatowt |
Inf. Sci. | 4 |
| 2019 | The 2nd International Workshop on Narrative Extraction from Text: Text2Story 2019
Alípio Mário Jorge, Ricardo Campos 0001, Adam Jatowt, Sumit Bhatia |
ECIR (2) | 1 |
| 2019 | Interactive System for Automatically Generating Temporal Narratives
Arian Pasquali, Vítor Mangaravite, Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt |
ECIR (2) | 4 |
| 2019 | ORSUM 2019 2nd workshop on online recommender systems and user modelingabstractThe ever-growing nature of user generated data in online systems poses obvious challenges on how we process such data. Typically, this issue is regarded as a scalability problem and has been mainly addressed with distributed algorithms able to train on massive amounts of data in short time windows. However, data is inevitably adding up at high speeds. Eventually one needs to discard or archive some of it. Moreover, the dynamic nature of data in user modeling and recommender systems, such as change of user preferences, and the continuous introduction of new users and items make it increasingly difficult to maintain up-to-date, accurate recommendation models. The objective of this workshop is to bring together researchers and practitioners interested in incremental and adaptive approaches to stream-based user modeling, recommendation and personalization, including algorithms, evaluation issues, incremental content and context mining, privacy and transparency, temporal recommendation or software frameworks for continuous learning. João Vinagre, Alípio Mário Jorge, Albert Bifet, Marie Al-Ghossein |
RecSys | 2 |
| 2019 | Guest Editorial - Special Issue on Data Mining for Geosciences
Alípio Mário Jorge, Rui L. Lopes, Germán Larrazábal, Hamed Nikhalat-Jahromi |
Data Min. Knowl. Discov. | 1 |
| 2019 | Information Processing & Management Journal Special Issue on Narrative Extraction from Texts (Text2Story): Preface
Alípio Mário Jorge, Ricardo Campos 0001, Adam Jatowt, Sérgio Nunes 0001 |
Inf. Process. Manag. | 1 |
| 2019 | Identifying topic relevant hashtags in Twitter streams
Filipe Figueiredo, Alípio Mário Jorge |
Inf. Sci. | 2 |
| 2018 | A Text Feature Based Automatic Keyword Extraction Method for Single Documents
Ricardo Campos 0001, Vítor Mangaravite, Arian Pasquali, Alípio Mário Jorge, Celia Nunes, Adam Jatowt |
ECIR | 4 |
| 2018 | YAKE! Collection-Independent Automatic Keyword Extractor
Ricardo Campos 0001, Vítor Mangaravite, Arian Pasquali, Alípio Mário Jorge, Celia Nunes, Adam Jatowt |
ECIR | 4 |
| 2018 | Forgetting techniques for stream-based matrix factorization in recommender systems
Pawel Matuszyk, João Vinagre, Myra Spiliopoulou, Alípio Mário Jorge, João Gama 0001 |
Knowl. Inf. Syst. | 4 |
| 2017 | Identifying top relevant dates for implicit time sensitive queries
Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes |
Inf. Retr. J. | 3 |
| 2016 | GTE-Rank: A time-aware search engine to answer time-sensitive queries
Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes |
Inf. Process. Manag. | 3 |
| 2015 | Guest editors introduction: special issue of the ECMLPKDD 2015 journal track
Concha Bielza, João Gama 0001, Alípio Mário Jorge, Indre Zliobaite |
Data Min. Knowl. Discov. | 3 |
| 2014 | GTE-Rank: Searching for Implicit Temporal Query ResultsabstractTemporal information retrieval has been a topic of great interest in recent years. Despite the efforts that have been conducted so far, most popular search engines remain underdeveloped when it comes to explicitly considering the use of temporal information in their search process. In this paper we present GTE-Rank, an online searching tool that takes time into account when ranking time-sensitive query web search results. GTE-Rank is defined as a linear combination of topical and temporal scores to reflect the relevance of any web page both in topical and temporal dimensions. The resulting system can be explored graphically through a search interface made available for research purposes. Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes |
CIKM | 3 |
| 2014 | GTE-Cluster: A Temporal Search Interface for Implicit Temporal Queries
Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes |
ECIR | 3 |
| 2014 | Classifying heart sounds using SAX motifs, random forests and text mining techniquesabstractIn this paper we describe an approach to classifying heart sounds (classes Normal, Murmur and Extra-systole) that is based on the discretization of sound signals using the SAX (Symbolic Aggregate Approximation) representation. The ability of automatically classifying heart sounds or at least support human decision in this task is socially relevant to spread the reach of medical care using simple mobile devices or digital stethoscopes. In our approach, sounds are first pre-processed using signal processing techniques (decimate, low-pass filter, normalize, Shannon envelope). Then the pre-processed symbols are transformed into sequences of discrete SAX symbols. These sequences are subject to a process of motif discovery. Frequent sequences of symbols (motifs) are adopted as features. Each sound is then characterized by the frequent motifs that occur in it and their respective frequency. This is similar to the term frequency (TF) model used in text mining. In this paper we compare the TF model with the application of the TFIDF (Term frequency - Inverse Document Frequency) and the use of bi-grams (frequent size two sequences of motifs). Results show the ability of the motifs based TF approach to separate classes and the relative value of the TFIDF and the bi-grams variants. The separation of the Extra-systole class is overly difficult and much better results are obtained for separating the Murmur class. Empirical validation is conducted using real data collected in noisy environments. We have also assessed the cost-reduction potential of the proposed methods by considering a fixed cost model and using a cost sensitive meta algorithm. Elsa Ferreira Gomes, Alípio Mário Jorge, Paulo J. Azevedo |
IDEAS | 2 |
| 2014 | A study of machine learning methods for detecting user interest during web sessionsabstractThe ability to have an automated real time detection of user interest during a web session is very appealing and can be very useful for a number of web intelligence applications. Low level interaction events associated with user interest manifestations form the basis of user interest models. However such data sets present a number of challenges from a machine learning perspective, including the level of noise in the data and class imbalance (given that the majority of content will not be of interest to a user). In this paper we evaluate a large number of machine learning techniques aimed at learning from class imbalanced data using two data sets collected from a real user study. We use the AUC, recall, precision and model complexity to compare the relative merits of these techniques and conclude that useful models with AUC above 0.8 can be obtained using a mix of sampling and cost based methods. Ensemble models can provide further accuracy but make deployment more complex. Alípio Mário Jorge, José Paulo Leal, Sarabjot S. Anand, Hugo Dias |
IDEAS | 1 |
| 2014 | Heart sounds classification using motif based segmentationabstractIn this paper we describe an algorithm for heart sound classification (classes Normal, Murmur and Extrasystole) based on the discretization of sound signals using the SAX (Symbolic Aggregate Approximation) representation. The general strategy is to automatically discover relevant top frequent motifs and relate them with the occurrence of systolic (S1) and diastolic (S2) sounds in the audio signals. The algorithm was tuned using motifs generated from a collection of audio signals obtained from a clinical trial in a hospital. Validation was performed on a separate set of unlabeled audio signals. Results indicate ability to improve the precision of the classification of the classes Normal and Murmur. Soraia Cruz Oliveira, Elsa Ferreira Gomes, Alípio Mário Jorge |
IDEAS | 3 |
| 2013 | Dimensions as Virtual Items: Improving the predictive ability of top-N recommender systems
Marcos Aurélio Domingues, Alípio Mário Jorge, Carlos Soares |
Inf. Process. Manag. | 2 |
| 2012 | GTE: a distributional second-order co-occurrence approach to improve the identification of top relevant dates in web snippetsabstractIn this paper, we present an approach to identify top relevant dates in Web snippets with respect to a given implicit temporal query. Our approach is two-fold. First, we propose a generic temporal similarity measure called GTE, which evaluates the temporal similarity between a query and a date. Second, we propose a classification model to accurately relate relevant dates to their corresponding query terms and withdraw irrelevant ones. We suggest two different solutions: a threshold-based classification strategy and a supervised classifier based on a combination of multiple similarity measures. We evaluate both strategies over a set of real-world text queries and compare the performance of our Web snippet approach with a query log approach over the same set of queries. Experiments show that determining the most relevant dates of any given implicit temporal query can be improved with GTE combined with the second order similarity measure InfoSimba, the Dice coefficient and the threshold-based strategy compared to (1) first-order similarity measures and (2) the query log based approach. Ricardo Campos 0001, Gaël Dias, Alípio Mário Jorge, Celia Nunes |
CIKM | 3 |
| 2012 | Finding Interesting Contexts for Explaining Deviations in Bus Trip Duration Using Distribution Rules
Alípio Mário Jorge, João Mendes-Moreira 0001, Jorge Freire de Sousa, Carlos Soares, Paulo J. Azevedo |
IDA | 1 |
| 2012 | Disambiguating Implicit Temporal Queries by Clustering Top Relevant Dates in Web SnippetsabstractWith the growing popularity of research in Temporal Information Retrieval (T-IR), a large amount of temporal data is ready to be exploited. The ability to exploit this information can be potentially useful for several tasks. For example, when querying "Football World Cup Germany", it would be interesting to have two separate clusters {1974,2006} corresponding to each of the two temporal instances. However, clustering of search results by time is a non-trivial task that involves determining the most relevant dates associated to a query. In this paper, we propose a first approach to flat temporal clustering of search results. We rely on a second order co-occurrence similarity measure approach which first identifies top relevant dates. Documents are grouped at the year level, forming the temporal instances of the query. Experimental tests were performed using real-world text queries. We used several measures for evaluating the performance of the system and compared our approach with Carrot Web-snippet clustering engine. Both experiments were complemented with a user survey. Ricardo Campos 0001, Alípio Mário Jorge, Gaël Dias, Celia Nunes |
Web Intelligence | 2 |
| 2011 | Mining Association Rules for Label Ranking
Cláudio Rebelo de Sá, Carlos Soares, Alípio Mário Jorge, Paulo J. Azevedo, Joaquim Pinto da Costa |
PAKDD (2) | 3 |
| 2011 | Exploiting Additional Dimensions as Virtual Items on Top-N Recommender SystemsabstractTraditionally, recommender systems for the web deal with applications that have two dimensions, users and items. Based on access data that relate these dimensions, a recommendation model can be built and used to identify a set of N items that will be of interest to a certain user. In this paper we propose a multidimensional approach, called DaVI (Dimensions as Virtual Items), that enables the use of common two-dimensional top-N recommender algorithms for the generation of recommendations using additional dimensions (e.g., contextual or background information). We empirically evaluate our approach with two different top-N recommender algorithms, Item-based Collaborative Filtering and Association Rules based, on two real world data sets. The empirical results demonstrate that DaVI enables the application of existing two-dimensional recommendation algorithms to exploit the useful information in multidimensional data. Marcos Aurélio Domingues, Alípio Mário Jorge, Carlos Soares |
Web Intelligence | 2 |
| 2010 | Ensembles of jittered association rule classifiers
Paulo J. Azevedo, Alípio Mário Jorge |
Data Min. Knowl. Discov. | 2 |
| 2009 | The Effect of Varying Parameters and Focusing on Bus Travel Time Prediction
João Mendes-Moreira 0001, Carlos Soares, Alípio Mário Jorge, Jorge Freire de Sousa |
PAKDD | 3 |
| 2008 | The Impact of Contextual Information on the Accuracy of Existing Recommender Systems for Web PersonalizationabstractTraditionally, recommender systems for the Web deal with applications that have two types of entities/dimensions, users and items. With these dimensions, a recommendation model can be built and used to identify a set of N items that will be of interest to a certain user. In this paper we propose a direct method that enriches the information in the access logs with new dimensions. We empirically test this method with two recommender systems, an item-based collaborative filtering technique and association rules, on three data sets. Our results show that while collaborative filtering is not able to take advantage of the new dimensions added, association rules are capable of profiting from our direct method. Marcos Aurélio Domingues, Alípio Mário Jorge, Carlos Soares |
Web Intelligence | 2 |
| 2008 | Incremental Collaborative Filtering for Binary RatingsabstractThe use of collaborative filtering (CF) recommenders on the Web is typically done in environments where data is constantly flowing. In this paper we propose an incremental version of item-based CF for implicit binary ratings, and compare it with a non-incremental one, as well as with an incremental user-based approach. We also study the usage of sparse matrices in these algorithms. We observe that recall and precision tend to improve when we continuously add information to the recommender model, and that the time spent for recommendation does not degrade. Time for updating the similarity matrix is relatively low and motivates the use of the item-based incremental approach. Catarina Miranda, Alípio Mário Jorge |
Web Intelligence | 2 |
| 2007 | Comparing Rule Measures for Predictive Association Rules
Paulo J. Azevedo, Alípio Mário Jorge |
ECML | 2 |
| 2006 | Distribution Rules with Numeric Attributes of Interest
Alípio Mário Jorge, Paulo J. Azevedo, Fernando Pereira 0002 |
PKDD | 1 |
| 2006 | Personalization of E-newsletters Based on Web Log Analysis and ClusteringabstractWe present a methodology for the personalization of e-newsletters based on the analysis of user access logs. To approach the problem we have used clustering on the set of users, described by their Web access patterns. Our work is evaluated using a case study with real data from e-newsletters sent by mail to users of a Web portal, and can be adapted to similar situations. Positive results were obtained, indicating that the methodology is able to automatically select contents for a personalized e-newsletter Carla Carvalho, Alípio Mário Jorge, Carlos Soares |
Web Intelligence | 2 |
| 2006 | Factor Analysis to Support the Visualization and Interpretation of Clusters of Portal UsersabstractClusterings based on many variables are difficult to visualize and interpret. We present a methodology based on factor analysis (FA) which can be used for that purpose. FA generates a small set of variables which encode most of the information in the original variables. We apply the methodology to segment the users of a Web portal, using access log data. It not only makes it simpler to visualize and understand the clusters which are obtained on the original variables but it also helps the analyst in selecting some of the original variables for further analysis of those clusters Carmen Rebelo, Pedro Quelhas Brito, Carlos Soares, Alípio Mário Jorge |
Web Intelligence | 4 |
| 2004 | Hierarchical Clustering for Thematic Browsing and Summarization of Large Sets of Association RulesabstractIn this paper we propose a method for grouping and summarizing large sets of association rules according to the items contained in each rule. We use hierarchical clustering to partition the initial rule set into thematically coherent subsets. This enables the summarization of the rule set by adequately choosing a representative rule for each subset, and helps in the interactive exploration of the rule model by the user. We define the requirements of our approach, and formally show the adequacy of the chosen approach to our aims. Rule clusters can also be used to infer novel interest measures for the rules. Such measures are based on the lexicon of the rules and are complementary to measures based on statistical properties, such as confidence, lift and conviction. We show examples of the application of the proposed techniques. Alípio Mário Jorge |
SDM | 1 |
| 1995 | Learning Recursion with Iterative Bootstrap Induction (Extended Abstract)
Alípio Mário Jorge, Pavel Brazdil |
ECML | 1 |