Paula Carvalho 0001

dblp:57/4090 · also Paula C. Carvalho, Paula Cristina Carvalho · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0003-2884-1250ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
YearPublicationVenuePosition
2025 DOMAIN: Explainable Credibility Assessment Tools for Empowering Online Readers Coping with Misinformation
abstract
Despite all the fact-checking initiatives on news and social media aimed at countering misinformation, they remain insufficient to promptly address the wide array of misleading information disseminated by both news and social media outlets. Rather than attempting to identify or filter misleading information, this work advocates new tools for assisting online readers in identifying misinformation among the massive online content pushed every day through multiple platforms. We introduce DOMAIN, an article assessment resource bundle comprising a multidimensional indicator to categorize articles into different types (hard news, soft news, opinion, satire, and conspiracy), a set of explanatory metrics to help users understand the results, a tool for verifying the reliability of the article’s source, and a text summary of the assessment. This work also studies how DOMAIN tools impact online readers, specifically focusing on (i) understanding the extent to which computer-generated assessments influence human perceptions of credibility; (ii) evaluating the effectiveness of automatic article categorization in human assessment of credibility; and (iii) identifying the most relevant explanatory metrics for promoting informed and critical consumption of information.
Danielle Caled, Paula Carvalho 0001, Mário J. Silva
ACM Trans. Web2
2024 Counter Hate Speech Detection in Youtube Conversations
Pedro Fialho, Ricardo Ribeiro 0001, Fernando Batista, Gil Ramos, António Fonseca, Sérgio Moro, Rita Guerra, Paula Carvalho 0001, Catarina Marques, Cláudia Silva 0002
IPMU (3)8
2024 Unveiling Patterns of Hate Speech in the Portuguese Sphere: A Social Network Analysis Approach
Catarina Pontes, António Fonseca, Sérgio Moro, Fernando Batista, Ricardo Ribeiro 0001, Catarina Marques, Paula Carvalho 0001, Cláudia Silva 0002, Rita Guerra
IPMU (3)7
2023 Argumentation models and their use in corpus annotation: Practice, prospects, and challenges
abstract
Abstract The study of argumentation is transversal to several research domains, from philosophy to linguistics, from the law to computer science and artificial intelligence. In discourse analysis, several distinct models have been proposed to harness argumentation, each with a different focus or aim. To analyze the use of argumentation in natural language, several corpora annotation efforts have been carried out, with a more or less explicit grounding on one of such theoretical argumentation models. In fact, given the recent growing interest in argument mining applications, argument-annotated corpora are crucial to train machine learning models in a supervised way. However, the proliferation of such corpora has led to a wide disparity in the granularity of the argument annotations employed. In this paper, we review the most relevant theoretical argumentation models, after which we survey argument annotation projects closely following those theoretical models. We also highlight the main simplifications that are often introduced in practice. Furthermore, we glimpse other annotation efforts that are not so theoretically grounded but instead follow a shallower approach. It turns out that most argument annotation projects make their own assumptions and simplifications, both in terms of the textual genre they focus on and in terms of adapting the adopted theoretical argumentation model for their own agenda. Issues of compatibility among argument-annotated corpora are discussed by looking at the problem from a syntactical, semantic, and practical perspective. Finally, we discuss current and prospective applications of models that take advantage of argument-annotated corpora.
Henrique Lopes Cardoso, Rui Sousa-Silva, Paula Carvalho 0001, Bruno Martins 0001
Nat. Lang. Eng.3
2022 Hate Speech Dynamics Against African descent, Roma and LGBTQI Communities in Portugal
abstract
This paper introduces FIGHT, a dataset containing 63,450 tweets, posted before and after the official declaration of Covid-19 as a pandemic by online users in Portugal. This resource aims at contributing to the analysis of online hate speech targeting the most representative minorities in Portugal, namely the African descent and the Roma communities, and the LGBTQI community, the most commonly reported target of hate speech in social media at the European context. We present the methods for collecting the data, and provide insightful statistics on the distribution of tweets included in FIGHT, considering both the temporal and spatial dimensions. We also analyze the availability over time of tweets targeting the above-mentioned communities, distinguishing public, private and deleted tweets. We believe this study will contribute to better understand the dynamics of online hate speech in Portugal, particularly in adverse contexts, such as a pandemic outbreak, allowing the development of more informed and accurate hate speech resources for Portuguese.
Paula Carvalho 0001, Bernardo Cunha Matos, Raquel Bento Santos, Fernando Batista, Ricardo Ribeiro 0001
LREC1
2022 Annotating Arguments in a Corpus of Opinion Articles
abstract
Interest in argument mining has resulted in an increasing number of argument annotated corpora. However, most focus on English texts with explicit argumentative discourse markers, such as persuasive essays or legal documents. Conversely, we report on the first extensive and consolidated Portuguese argument annotation project focused on opinion articles. We briefly describe the annotation guidelines based on a multi-layered process and analyze the manual annotations produced, highlighting the main challenges of this textual genre. We then conduct a comprehensive inter-annotator agreement analysis, including argumentative discourse units, their classes and relations, and resulting graphs. This analysis reveals that each of these aspects tackles very different kinds of challenges. We observe differences in annotator profiles, motivating our aim of producing a non-aggregated corpus containing the insights of every annotator. We note that the interpretation and identification of token-level arguments is challenging; nevertheless, tasks that focus on higher-level components of the argument structure can obtain considerable agreement. We lay down perspectives on corpus usage, exploiting its multi-faceted nature.
Gil Rocha, Luís Trigo, Henrique Lopes Cardoso, Rui Sousa-Silva, Paula Carvalho 0001, Bruno Martins 0001, Miguel Won
LREC5
2022 Predicting Argument Density from Multiple Annotations
Gil Rocha, Bernardo Leite 0002, Luís Trigo, Henrique Lopes Cardoso, Rui Sousa-Silva, Paula Carvalho 0001, Bruno Martins 0001, Miguel Won
NLDB6
2016 Modelling Context with User Embeddings for Sarcasm Detection in Social Media
abstract
We introduce a deep neural network for automated sarcasm detection.Recent work has emphasized the need for models to capitalize on contextual features, beyond lexical and syntactic cues present in utterances.For example, different speakers will tend to employ sarcasm regarding different subjects and, thus, sarcasm detection models ought to encode such speaker information.Current methods have achieved this by way of laborious feature engineering.By contrast, we propose to automatically learn and then exploit user embeddings, to be used in concert with lexical signals to recognize sarcasm.Our approach does not require elaborate feature engineering (and concomitant data scraping); fitting user embeddings requires only the text from their previous posts.The experimental results show that the our model outperforms a state-of-the-art approach leveraging an extensive set of carefully crafted features.
Silvio Amir, Byron C. Wallace, Paula Carvalho 0001, Mário J. Silva
CoNLL4
2016 Port4NooJ v3.0: Integrated Linguistic Resources for Portuguese NLP
Cristina Mota, Paula Carvalho 0001, Anabela Barreiro
LREC2
2010 Second HAREM: Advancing the State of the Art of Named Entity Recognition in Portuguese
Cláudia Freitas, Cristina Mota, Diana Santos, Hugo Gonçalo Oliveira, Paula Carvalho 0001
LREC5
2004 Portuguese Large-scale Language Resources for NLP Applications
Elisabete Ranchhod, Paula Carvalho 0001, Cristina Mota, Anabela Barreiro
LREC2