VLDB 2026 Research / reviewers in the wild / expert
Ricardo Ribeiro 0001
dblp:23/20-1 · also Ricardo Daniel Santos Faro Marques Ribeiro
· DBLP profile ↗
34ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-2058-693XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Large Language Model Vulnerabilities
Meda Racaityte, Hélder Bastos, Ricardo Ribeiro 0001, Luís Nunes |
SECRYPT (1) | 3 |
| 2025 | Exploring Metric Correlations for Legal Text Summarization EvaluationabstractThe rapid advancements in legal text summarization have not been matched by equivalent progress in evaluation metrics capable of assessing the quality of legal summaries. Traditional evaluation approaches, such as ROUGE, remain widely used despite their inability to capture semantic fidelity. While more recent metrics focus on semantic evaluation, their applicability to legal summarization has not been thoroughly tested, and their performance is highly dependent on embedding models and computational resources, particularly for long and complex legal texts. Furthermore, the absence of publicly available datasets with expert annotations hinders the development and validation of domain-specific evaluation methods. In this paper, we address these challenges by introducing the first publicly available dataset of Portuguese legal summaries, annotated by legal experts across multiple dimensions such as Coherence and Relevance. We use this dataset to systematically evaluate several recent evaluation metrics, comparing their performance against ROUGE, the standard metric for summarization tasks. Our analysis, based on Spearman correlation with human judgments, reveals that ROUGE-2 maintains the highest correlation across almost every evaluated dimension, outperforming more recent metrics, including semantic-based approaches. These results emphasize the challenges of adapting new evaluation frameworks to the legal domain and underscore the need for further research into metrics that can better capture domain-specific requirements. Martim Zanatti, Ricardo Ribeiro 0001, H. Sofia Pinto |
ICAIL | 2 |
| 2025 | Mining population opinion about local policeabstractAbstract Sentiment analysis, or opinion mining, is an important task of natural language processing (NLP) that extracts opinions, attitudes, and emotions from text. With the growth of digital platforms like blogs and social networks, opinion mining has become a key tool for organizations to understand public sentiment. In recent research, machine learning and lexicon-based approaches have been applied to analyze sentiments. Our work specifically focuses on national security, where sentiment analysis offers crucial insights into local opinions, helping authorities gauge public mood. As part of our research, we developed the Public Sensing about Police Platform, a prototype system designed to analyze emotions from social networks. This system generates dashboards for law enforcement and security agencies, providing actionable intelligence for public safety. Our findings show that “Hate” was the most common emotion expressed in relation to police interventions, indicating widespread unpopularity of these actions and a resulting sense of insecurity among the public. Kenny Matos, Ricardo Ribeiro 0001 |
Multim. Tools Appl. | 2 |
| 2025 | Towards Cyberbullying Detection: Building, Benchmarking and Longitudinal Analysis of Aggressiveness and Conflicts/Attacks Datasets From TwitterabstractOffense and hate speech are a source of online conflicts which have become common in social media and, as such, their study is a growing topic of research in machine learning and natural language processing. This article presents two Portuguese language offense-related datasets that deepen the study of the subject: an Aggressiveness dataset and a Conflicts/Attacks dataset. While the former is similar to other offense detection related datasets, the latter constitutes a novelty due to the use of the history of the interaction between users. Several studies were carried out to construct and analyze the data in the datasets. The first study included gathering expressions of verbal aggression witnessed by adolescents to guide data extraction for the datasets. The second study included extracting data from Twitter (in Portuguese) that matched the most frequent expressions/words/sentences that were identified in the previous study. The third study consisted in the development of the Aggressiveness dataset, the Conflicts/Attacks dataset, and classification models. In our fourth study, we proposed to examine whether online aggression and conflicts/attacks revealed any trend changes over time with a sample of 86 adolescents. With this study, we also proposed to investigate whether the amount of tweets sent over a period of 273 days was related to online aggression and conflicts/attacks. Finally, we analyzed the percentage of participants who participated in the aggressions and/or attacks/conflicts. Paula Ferreira 0004, Nádia Salgado Pereira, Hugo Rosa, Sofia Oliveira, Luísa Coheur, Sofia Mateus Francisco, Sidclay Bezerra de Souza, Ricardo Ribeiro 0001, João Paulo Carvalho 0001, Paula Paulino, Isabel Trancoso, Ana Margarida Veiga Simão |
IEEE Trans. Affect. Comput. | 8 |
| 2024 | Counter Hate Speech Detection in Youtube Conversations
Pedro Fialho, Ricardo Ribeiro 0001, Fernando Batista, Gil Ramos, António Fonseca, Sérgio Moro, Rita Guerra, Paula Carvalho 0001, Catarina Marques, Cláudia Silva 0002 |
IPMU (3) | 2 |
| 2024 | Unveiling Patterns of Hate Speech in the Portuguese Sphere: A Social Network Analysis Approach
Catarina Pontes, António Fonseca, Sérgio Moro, Fernando Batista, Ricardo Ribeiro 0001, Catarina Marques, Paula Carvalho 0001, Cláudia Silva 0002, Rita Guerra |
IPMU (3) | 5 |
| 2023 | Automatic Recognition of the General-Purpose Communicative Functions Defined by the ISO 24617-2 Standard for Dialog Act Annotation (Extended Abstract)abstractFrom the perspective of a dialog system, the identification of the intention behind the segments in a dialog is important, as it provides cues regarding the information present in the segments and how they should be interpreted. The ISO 24617-2 standard for dialog act annotation defines a hierarchically organized set of general-purpose communicative functions that correspond to different intentions that are relevant in the context of a dialog. In this paper, we explore the automatic recognition of these functions. To do so, we propose to adapt existing approaches to dialog act recognition, so that they can deal with the hierarchical classification problem. More specifically, we propose the use of an end-to-end hierarchical network with cascading outputs and maximum a posteriori path estimation to predict the communicative function at each level of the hierarchy, preserve the dependencies between the functions in the path, and decide at which level to stop. Additionally, we rely on transfer learning processes to address the data scarcity problem. Our experiments on the DialogBank show that this approach outperforms both flat and hierarchical approaches based on multiple classifiers and that each of its components plays an important role in the recognition of general-purpose communicative functions. Eugénio Ribeiro, Ricardo Ribeiro 0001, David Martins de Matos |
IJCAI | 2 |
| 2023 | Semantic similarity for mobile application recommendation under scarce user dataabstractThe More Like This recommendation approach is ubiquitous in multiple domains and consists in recommending items similar to the one currently selected by the user, being particularly relevant when user data is scarce. We studied the impact of using semantic similarity in the context of the More Like This recommendation for mobile applications, by leveraging dense representations in order to infer the similarity between applications, based on their textual fields. Our approach was validated by comparing it to the solution currently in use by Aptoide, a mobile application store, since no benchmarks are available for this specific task. To further evaluate the proposed model, we asked 1262 users to compare the results achieved by both approaches, also allowing us to build an annotated dataset of similar applications. Results show that the semantic representations are able to capture the context of the applications, with more useful recommendations being presented to users, when compared to Aptoide’s current solution. For replication and future research, all the code and data used in this study was made publicly available, including two novel datasets (installed applications for more than one million users, and app user-labeled similarity), the fine-tuned model, and the test platform. João Coelho, Diogo Mano, Beatriz Paula, Carlos Coutinho, João Oliveira 0001, Ricardo Ribeiro 0001, Fernando Batista |
Eng. Appl. Artif. Intell. | 6 |
| 2022 | Hate Speech Dynamics Against African descent, Roma and LGBTQI Communities in PortugalabstractThis paper introduces FIGHT, a dataset containing 63,450 tweets, posted before and after the official declaration of Covid-19 as a pandemic by online users in Portugal. This resource aims at contributing to the analysis of online hate speech targeting the most representative minorities in Portugal, namely the African descent and the Roma communities, and the LGBTQI community, the most commonly reported target of hate speech in social media at the European context. We present the methods for collecting the data, and provide insightful statistics on the distribution of tweets included in FIGHT, considering both the temporal and spatial dimensions. We also analyze the availability over time of tweets targeting the above-mentioned communities, distinguishing public, private and deleted tweets. We believe this study will contribute to better understand the dynamics of online hate speech in Portugal, particularly in adverse contexts, such as a pandemic outbreak, allowing the development of more informed and accurate hate speech resources for Portuguese. Paula Carvalho 0001, Bernardo Cunha Matos, Raquel Bento Santos, Fernando Batista, Ricardo Ribeiro 0001 |
LREC | 5 |
| 2022 | Automatic Recognition of the General-Purpose Communicative Functions Defined by the ISO 24617-2 Standard for Dialog Act AnnotationabstractFrom the perspective of a dialog system, it is important to identify the intention behind the segments in a dialog, since it provides an important cue regarding the information that is present in the segments and how they should be interpreted. ISO 24617-2, the standard for dialog act annotation, defines a hierarchically organized set of general-purpose communicative functions which correspond to different intentions that are relevant in the context of a dialog. We explore the automatic recognition of these communicative functions in the DialogBank, which is a reference set of dialogs annotated according to this standard. To do so, we propose adaptations of existing approaches to flat dialog act recognition that allow them to deal with the hierarchical classification problem. More specifically, we propose the use of an end-to-end hierarchical network with cascading outputs and maximum a posteriori path estimation to predict the communicative function at each level of the hierarchy, preserve the dependencies between the functions in the path, and decide at which level to stop. Furthermore, since the amount of dialogs in the DialogBank is small, we rely on transfer learning processes to reduce overfitting and improve performance. The results of our experiments show that our approach outperforms both a flat one and hierarchical approaches based on multiple classifiers and that each of its components plays an important role towards the recognition of general-purpose communicative functions. Eugénio Ribeiro, Ricardo Ribeiro 0001, David Martins de Matos |
J. Artif. Intell. Res. | 2 |
| 2021 | Gun Model Classification Based on Fired Cartridge Case Head Images with Siamese Networks
Sérgio Valentim, Tiago Fonseca, João Ferreira 0001, Tomás Brandão, Ricardo Ribeiro 0001, Stefan Nae |
ISDA | 5 |
| 2021 | Assessing kinetic meaning of music and dance via deep cross-modal retrieval
Francisco Raposo 0001, David Martins de Matos, Ricardo Ribeiro 0001 |
Neural Comput. Appl. | 3 |
| 2020 | Using Topic Information to Improve Non-exact Keyword-Based Search for Mobile Applications
Eugénio Ribeiro, Ricardo Ribeiro 0001, Fernando Batista, João Oliveira 0001 |
IPMU (1) | 2 |
| 2020 | Mapping the Dialog Act Annotations of the LEGO Corpus into ISO 24617-2 Communicative FunctionsabstractISO 24617-2, the ISO standard for dialog act annotation, sets the ground for more comparable research in the area. However, the amount of data annotated according to it is still reduced, which impairs the development of approaches for automatic recognition. In this paper, we describe a mapping of the original dialog act labels of the LEGO corpus, which have been neglected, into the communicative functions of the standard. Although this does not lead to a complete annotation according to the standard, the 347 dialogs provide a relevant amount of data that can be used in the development of automatic communicative function recognition approaches, which may lead to a wider adoption of the standard. Using the 17 English dialogs of the DialogBank as gold standard, our preliminary experiments have shown that including the mapped dialogs during the training phase leads to improved performance while recognizing communicative functions in the Task dimension. Eugénio Ribeiro, Ricardo Ribeiro 0001, David Martins de Matos |
LREC | 2 |
| 2019 | Deep Dialog Act Recognition using Multiple Token, Segment, and Context Information RepresentationsabstractAutomatic dialog act recognition is a task that has been widely explored over the years. In recent works, most approaches to the task explored different deep neural network architectures to combine the representations of the words in a segment and generate a segment representation that provides cues for intention. In this study, we explore means to generate more informative segment representations, not only by exploring different network architectures, but also by considering different token representations, not only at the word level, but also at the character and functional levels. At the word level, in addition to the commonly used uncontextualized embeddings, we explore the use of contextualized representations, which are able to provide information concerning word sense and segment structure. Character-level tokenization is important to capture intention-related morphological aspects that cannot be captured at the word level. Finally, the functional level provides an abstraction from words, which shifts the focus to the structure of the segment. Additionally, we explore approaches to enrich the segment representation with context information from the history of the dialog, both in terms of the classifications of the surrounding segments and the turn-taking history. This kind of information has already been proved important for the disambiguation of dialog acts in previous studies. Nevertheless, we are able to capture additional information by considering a summary of the dialog history and a wider turn-taking context. By combining the best approaches at each step, we achieve performance results that surpass the previous state-of-the-art on generic dialog act recognition on both the Switchboard Dialog Act Corpus (SwDA) and the ICSI Meeting Recorder Dialog Act Corpus (MRDA), which are two of the most widely explored corpora for the task. Furthermore, by considering both past and future context, similarly to what happens in an annotation scenario, our approach achieves a performance similar to that of a human annotator on SwDA and surpasses it on MRDA. Eugénio Ribeiro, Ricardo Ribeiro 0001, David Martins de Matos |
J. Artif. Intell. Res. | 2 |
| 2019 | An information-theoretic approach to machine-oriented music summarizationabstractMusic summarization allows for higher efficiency in processing, storage, and sharing of datasets. Machine-oriented approaches, being agnostic to human consumption, optimize these aspects even further. Such summaries have already been successfully validated in some MIR tasks. We now generalize previous conclusions by evaluating the impact of generic summarization of music from a probabilistic perspective. We estimate Gaussian distributions for original and summarized songs and compute their relative entropy, in order to measure information loss incurred by summarization. Our results suggest that relative entropy is a good predictor of summarization performance in the context of tasks relying on a bag-of-features model. Based on this observation, we further propose a straightforward yet expressive summarizer, which minimizes relative entropy with respect to the original song, that objectively outperforms previous methods and is better suited to avoid potential copyright issues. Francisco Raposo 0001, David Martins de Matos, Ricardo Ribeiro 0001 |
Pattern Recognit. Lett. | 3 |
| 2018 | Using Fuzzy Fingerprints for Cyberbullying Detection in Social NetworksabstractAs cyberbullying becomes more and more frequent in social networks, automatically detecting it and pro-actively acting upon it becomes of the utmost importance. In this work, we study how a recent technique with proven success in similar tasks, Fuzzy Fingerprints, performs when detecting textual cyberbullying in social networks. Despite being commonly treated as binary classification task, we argue that this is in fact a retrieval problem where the only relevant performance is that of retrieving cyberbullying interactions. Experiments show that the Fuzzy Fingerprints slightly outperforms baseline classifiers when tested in a close to real life scenario, where cyberbullying instances are rarer than those without cyberbullying. Hugo Rosa, João Paulo Carvalho 0001, Pável Calado, Bruno Martins 0001, Ricardo Ribeiro 0001, Luísa Coheur |
FUZZ-IEEE | 5 |
| 2018 | A "Deeper" Look at Detecting Cyberbullying in Social NetworksabstractAs cyberbullying becomes more and more frequent in social networks, automatically detecting it and pro-actively acting upon it becomes of the utmost importance. In this work, a detailed look at the current state-of-the-art in cyberbullying detection reveals that deep learning techniques have seldom been used to tackle this problem, despite growing reputation in other text-based classification tasks. Motivated by neural networks' documented success, three architectures are implemented from similar works: a simple CNN, a hybrid CNN-LSTM and a mixed CNN-LSTM-DNN. In addition, three text representations are trained from three different sources, via the word2vec model: Google-News, Twitter and Formspring. The experiment shows that these models with one of the above embeddings beat other benchmark classifiers (Support Vector Machines and Logistic Regression) both in an unbalanced and balanced version of the same dataset. Hugo Rosa, David Martins de Matos, Ricardo Ribeiro 0001, Luísa Coheur, João Paulo Carvalho 0001 |
IJCNN | 3 |
| 2017 | Stepwise API usage assistance using n-gram language models
André L. Santos 0001, Gonçalo Prendi, Hugo S. Sousa, Ricardo Ribeiro 0001 |
J. Syst. Softw. | 4 |
| 2017 | Event-based summarization using a centrality-as-relevance model
Luís Marujo, Ricardo Ribeiro 0001, Anatole Gershman, David Martins de Matos, João Paulo da Silva Neto, Jaime G. Carbonell |
Knowl. Inf. Syst. | 2 |
| 2016 | SPA: Web-based Platform for easy Access to Speech Processing Modules
Fernando Batista, Pedro Curto, Isabel Trancoso, Alberto Abad, Jaime Ferreira, Eugénio Ribeiro, Helena Moniz, David Martins de Matos, Ricardo Ribeiro 0001 |
LREC | 9 |
| 2016 | Exploring events and distributed representations of text in multi-document summarization
Luís Marujo, Wang Ling, Ricardo Ribeiro 0001, Anatole Gershman, Jaime G. Carbonell, David Martins de Matos, João Paulo da Silva Neto |
Knowl. Based Syst. | 3 |
| 2016 | Summarization of films and documentaries based on subtitles and scripts
Marta Aparício, Paulo Figueiredo, Francisco Raposo 0001, David Martins de Matos, Ricardo Ribeiro 0001, Luís Marujo |
Pattern Recognit. Lett. | 5 |
| 2016 | Using Generic Summarization to Improve Music Information Retrieval TasksabstractIn order to satisfy processing time constraints, many music information retrieval (MIR) tasks process only a segment of the whole music signal. This may lead to decreasing performance, as the most important information for the tasks may not be in the processed segments. We leverage generic summarization algorithms, previously applied to text and speech, to summarize items in music datasets. These algorithms build summaries (both concise and diverse), by selecting appropriate segments from the input signal, also making them good candidates to summarize music. We evaluate the summarization process on binary and multiclass music genre classification tasks, by comparing the accuracy when using summarized datasets against the accuracy when using human-oriented summaries, continuous segments (the traditional method used for addressing the previously mentioned time constraints), and full songs of the original dataset. We show that GRASSHOPPER, LexRank, LSA, MMR, and a Support Sets-based centrality model improve classification performance when compared to selected baselines. We also show that summarized datasets lead to a classification performance whose difference is not statistically significant from using full songs. Furthermore, we make an argument stating the advantages of sharing summarized datasets for future MIR research. Francisco Raposo 0001, Ricardo Ribeiro 0001, David Martins de Matos |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | On the Application of Generic Summarization Algorithms to MusicabstractSeveral generic summarization algorithms were developed in the past and successfully applied in fields such as text and speech summarization. In this paper, we review and apply these algorithms to music. To evaluate their performance, we adopt an extrinsic approach: we compare a Fado genre classifier's performance using truncated contiguous clips against the summaries extracted with those algorithms on two different datasets. We show that Maximal Marginal Relevance (MMR), LexRank, and Latent Semantic Analysis (LSA) all improve classification performance in both datasets used for testing. Francisco Raposo 0001, Ricardo Ribeiro 0001, David Martins de Matos |
IEEE Signal Process. Lett. | 2 |
| 2014 | OpenLogos Semantico-Syntactic Knowledge-Rich Bilingual Dictionaries
Anabela Barreiro, Fernando Batista, Ricardo Ribeiro 0001, Helena Moniz, Isabel Trancoso |
LREC | 3 |
| 2014 | Revising the annotation of a Broadcast News corpus: a linguistic approach
Vera Cabarrão, Helena Moniz, Fernando Batista, Ricardo Ribeiro 0001, Nuno J. Mamede, Hugo Meinedo, Isabel Trancoso, Ana Isabel Mata, David Martins de Matos |
LREC | 4 |
| 2013 | Revisiting Centrality-as-Relevance: Support Sets and Similarity as Geometric Proximity: Extended abstract
Ricardo Ribeiro 0001, David Martins de Matos |
IJCAI | 1 |
| 2013 | Self reinforcement for important passage retrievalabstractIn general, centrality-based retrieval models treat all elements of the retrieval space equally, which may reduce their effectiveness. In the specific context of extractive summarization (or important passage retrieval), this means that these models do not take into account that information sources often contain lateral issues, which are hardly as important as the description of the main topic, or are composed by mixtures of topics. We present a new two-stage method that starts by extracting a collection of key phrases that will be used to help centrality-as-relevance retrieval model. We explore several approaches to the integration of the key phrases in the centrality model. The proposed method is evaluated using different datasets that vary in noise (noisy vs clean) and language (Portuguese vs English). Results show that the best variant achieves relative performance improvements of about 31% in clean data and 18% in noisy data. Ricardo Ribeiro 0001, Luís Marujo, David Martins de Matos, João Paulo da Silva Neto, Anatole Gershman, Jaime G. Carbonell |
SIGIR | 1 |
| 2011 | Centrality-as-Relevance: Support Sets and Similarity as Geometric ProximityabstractIn automatic summarization, centrality-as-relevance means that the most important content of an information source, or a collection of information sources, corresponds to the most central passages, considering a representation where such notion makes sense (graph, spatial, etc.). We assess the main paradigms, and introduce a new centrality-based relevance model for automatic summarization that relies on the use of support sets to better estimate the relevant content. Geometric proximity is used to compute semantic relatedness. Centrality (relevance) is determined by considering the whole input source (and not only local information), and by taking into account the existence of minor topics or lateral subjects in the information sources to be summarized. The method consists in creating, for each passage of the input source, a support set consisting only of the most semantically related passages. Then, the determination of the most relevant content is achieved by selecting the passages that occur in the largest number of support sets. This model produces extractive summaries that are generic, and language- and domain-independent. Thorough automatic evaluation shows that the method achieves state-of-the-art performance, both in written text, and automatically transcribed speech summarization, including when compared to considerably more complex approaches. Ricardo Ribeiro 0001, David Martins de Matos |
J. Artif. Intell. Res. | 1 |
| 2008 | Using prior knowledge to assess relevance in speech summarizationabstractWe explore the use of topic-based automatically acquired prior knowledge in speech summarization, assessing its influence throughout several term weighting schemes. All information is combined using latent semantic analysis as a core procedure to compute the relevance of the sentence-like units of the given input source. Evaluation is performed using the self-information measure, which tries to capture the informativeness of the summary in relation to the summarized input source. The similarity of the output summaries of the several approaches is also analyzed. Ricardo Ribeiro 0001, David Martins de Matos |
SLT | 1 |
| 2004 | Rethinking Reusable Resources
David Martins de Matos, Ricardo Ribeiro 0001, Nuno J. Mamede |
LREC | 2 |
| 2002 | Morphosyntactic Disambiguation for TTS Systems
Ricardo Ribeiro 0001, Luís C. Oliveira, Isabel Trancoso |
LREC | 1 |
| 2000 | Some Language Resources and Tools for Computational Processing of Portuguese at INESC
Luzia Wittmann, Ricardo Ribeiro 0001, Tânia Pêgo, Fernando Batista |
LREC | 2 |