EDBT 2026 Demo / reviewers in the wild / expert
Paolo Rosso
dblp:05/3463
· DBLP profile ↗
86ranked-venue papers in the field
4as first author
34since 2021 · last 2026
0000-0002-8922-1242ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 68 (3 first)Database Systems & Data Management · 8Data Mining & Knowledge Discovery · 4Knowledge Engineering, Semantic Web & Information Systems · 3Big Data, Cloud & Distributed Data Systems · 2 (1 first)Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EXIST 2026: Physiological Data for Multimodal Sexism Characterization in Social Media
Laura Plaza, Jorge Carrillo de Albornoz, Elena Gomis-Vicent, Iván Árcos, María Aloy-Mayo, Paolo Rosso, Damiano Spina |
ECIR (4) | 6 |
| 2026 | Zoom In Disparities in Healthcare LLM Q&A
Ipek Baris Schlicht, Burcu Sayin, Zhixue Zhao, Frederik Labonté, Cesare Barbera, Marco Viviani 0001, Paolo Rosso, Lucie Flek |
NLDB | 7 |
| 2025 | EXIST 2025: Learning with Disagreement for Sexism Identification and Characterization in Tweets, Memes, and TikTok Videos
Laura Plaza, Jorge Carrillo de Albornoz, Iván Árcos, Paolo Rosso, Damiano Spina, Enrique Amigó, Julio Gonzalo 0001, Roser Morante |
ECIR (5) | 4 |
| 2025 | Do LLMs Provide Consistent Answers to Health-Related Questions Across Languages?
Ipek Baris Schlicht, Zhixue Zhao, Burcu Sayin, Lucie Flek, Paolo Rosso |
ECIR (3) | 5 |
| 2024 | Overview of PAN 2024: Multi-author Writing Style Analysis, Multilingual Text Detoxification, Oppositional Thinking Analysis, and Generative AI Authorship Verification - Extended Abstract
Janek Bevendorff, Xavier Bonet Casals, Berta Chulvi, Daryna Dementieva, Ashraf Elnagar, Dayne Freitag, Maik Fröbe, Damir Korencic, Maximilian Mayerl, Animesh Mukherjee 0001, Alexander Panchenko, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Alisa Smirnova, Efstathios Stamatatos, Benno Stein 0001, Mariona Taulé, Dmitry Ustalov, Matti Wiegmann, Eva Zangerle |
ECIR (6) | 14 |
| 2024 | Reading Between the Frames: Multi-modal Depression Detection in Videos from Non-verbal Cues
David Gimeno-Gómez, Ana-Maria Bucur, Adrian Cosma, Carlos D. Martínez-Hinarejos, Paolo Rosso |
ECIR (1) | 5 |
| 2024 | EXIST 2024: sEXism Identification in Social neTworks and Memes
Laura Plaza, Jorge Carrillo de Albornoz, Enrique Amigó, Julio Gonzalo 0001, Roser Morante, Paolo Rosso, Damiano Spina, Berta Chulvi, Alba Maeso, Víctor Ruiz |
ECIR (5) | 6 |
| 2024 | Unraveling Disagreement Constituents in Hateful Speech
Giulia Rizzi, Alessandro Astorino, Paolo Rosso, Elisabetta Fersini |
ECIR (4) | 3 |
| 2024 | Towards improving user awareness of search engine biases: A participatory design approachabstractAbstract Bias in news search engines has been shown to influence users' perceptions of a news topic and contribute to the polarisation of society. As a result, there is a need for news search engines that increase user awareness of biases in the search results. While technical approaches have been developed to mitigate biases in search, very few studies have investigated user preferences in interface designs for potentially raising their awareness of biases in news search engines. In this study, we utilized a participatory design methodology to develop eight prototypes with different features that could potentially be used to raise user awareness of biases in news search engines. We conducted three user studies, involving 132 participants with Computer Science backgrounds, to evaluate these prototypes. Our findings indicate the importance of news search engines that (a) inform users of possible biases in the results (bias visualization approach) and (b) allow users to access alternative search results (results‐reranking approach). Our study provides further insights into the strengths and possible risks of each approach, which are important for future research on designing interfaces for raising user awareness of biases in news search engines. Monica Lestari Paramita, Maria Kasinidou, Styliani Kleanthous, Paolo Rosso, Tsvi Kuflik, Frank Hopfgartner |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2024 | Schema-Aware Hyper-Relational Knowledge Graph Embeddings for Link PredictionabstractKnowledge Graph (KG) embeddings have become a powerful paradigm to resolve link prediction tasks for KG completion. The widely adopted triple-based representation, where each triplet$(h,r,t)$links two entities$h$and$t$through a relation$r$, oversimplifies the complex nature of the data stored in a KG, in particular for hyper-relational facts, where each fact contains not only a base triplet$(h,r,t)$, but also the associated key-value pairs$(k,v)$. Even though a few recent techniques tried to learn from such data by transforming a hyper-relational fact into an n-ary representation (i.e., a set of key-value pairs only without triplets), they result in suboptimal models as they are unaware of the triplet structure, which serves as the fundamental data structure in modern KGs and preserves the essential information for link prediction. Moreover, as the KG schema information has been shown to be useful for resolving link prediction tasks, it is thus essential to incorporate the corresponding hyper-relational schema in KG embeddings. Against this background, we propose sHINGE, a schema-aware hyper-relational KG embedding model, which learns from hyper-relational facts directly (without the transformation to the n-ary representation) and their corresponding hyper-relational schema in a KG. Our extensive evaluation shows the superiority of sHINGE on various link prediction tasks over KGs. In particular, compared to a sizeable collection of 21 baselines, sHINGE consistently outperforms the best-performing triple-based KG embedding method, hyper-relational KG embedding method, and schema-aware KG embedding method by 19.1%, 1.8%, and 12.9%, respectively. Yuhuan Lu 0001, Dingqi Yang, Pengyang Wang, Paolo Rosso, Philippe Cudré-Mauroux |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Fast and Slow Thinking: A Two-Step Schema-Aware Approach for Instance Completion in Knowledge GraphsabstractModern Knowledge Graphs (KG) often suffer from an incompleteness issue (i.e., missing facts). By representing a fact as a triplet$(h,r,t)$linking two entities$h$and$t$via a relation$r$, existing KG completion approaches mostly consider a link prediction task to solve this problem, i.e., given two elements of a triplet predicting the missing one, such as$(h,r,?)$. However, this task implicitly has a strong yet impractical assumption on the two given elements in a triplet, which have to be correlated, resulting otherwise in meaningless predictions, such as (Marie Curie,headquarters location, ?). Against this background, this paper studies an instance completion task suggesting$r$-$t$pairs for a given$h$, i.e.,$(h,?,?)$. Inspired by the human psychological principle “fast-and-slow thinking”, we propose a two-step schema-aware approach RETA++ to efficiently solve our instance completion problem. It consists of two components: afastRETA-Filter efficiently filtering candidate$r$-$t$pairs schematically matching the given$h$, and adeliberateRETA-Grader leveraging a KG embedding model scoring each candidate$r$-$t$pair considering the plausibility of both the input triplet and its corresponding schema. RETA++ systematically integrates them by training RETA-Grader on the reduced solution space output by RETA-Filter via a customized negative sampling process, so as to fully benefit from the efficiency of RETA-Filter in solution space reduction and the deliberation of RETA-Grader in scoring candidate triplets. We evaluate our approach against a sizable collection of state-of-the-art techniques on three real-world KG datasets. Results show that RETA-Filter can efficiently reduce the solution space for the instance completion task, outperforming best baseline techniques by 10.61%–84.75% on the reduced solution space size, while also being 1.7×–29.6x faster than these techniques. Moreover, RETA-Grader trained on the reduced solution space also significantly outperforms the best state-of-the-art techniques on the instance completion task by 31.90%–105.02%. Dingqi Yang, Bingqing Qu, Paolo Rosso, Philippe Cudré-Mauroux |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Overview of PAN 2023: Authorship Verification, Multi-author Writing Style Analysis, Profiling Cryptocurrency Influencers, and Trigger Detection - Extended Abstract
Janek Bevendorff, Mara Chinea-Rios, Marc Franco-Salvador, Annina Heini, Erik Körner, Krzysztof Kredens, Maximilian Mayerl, Piotr Pezik, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Magdalena Wolska, Eva Zangerle |
ECIR (3) | 11 |
| 2023 | It's Just a Matter of Time: Detecting Depression with Time-Enriched Multimodal Transformers
Ana-Maria Bucur, Adrian Cosma, Paolo Rosso, Liviu P. Dinu |
ECIR (1) | 3 |
| 2023 | Overview of EXIST 2023: sEXism Identification in Social NeTworks
Laura Plaza, Jorge Carrillo de Albornoz, Roser Morante, Enrique Amigó, Julio Gonzalo 0001, Damiano Spina, Paolo Rosso |
ECIR (3) | 7 |
| 2023 | Multilingual Detection of Check-Worthy Claims Using World Languages and Adapter Fusion
Ipek Baris Schlicht, Lucie Flek, Paolo Rosso |
ECIR (1) | 3 |
| 2023 | How Challenging is Multimodal Irony Detection?
Manuj Malik, David Tomás 0001, Paolo Rosso |
NLDB | 3 |
| 2023 | Cross-Domain and Cross-Language Irony Detection: The Impact of Bias on Models' Generalization
Reynier Ortega Bueno, Paolo Rosso, Elisabetta Fersini |
NLDB | 2 |
| 2023 | Recognizing misogynous memes: Biased models and tricky archetypesabstractWarning: This paper contains examples of language and images which may be offensive. Misogyny is a form of hate against women and has been spreading exponentially through the Web, especially on social media platforms. Hateful content towards women can be conveyed not only by text but also using visual and/or audio sources or their combination, highlighting the necessity to address it from a multimodal perspective. One of the predominant forms of multimodal content against women is represented by memes, which are images characterized by pictorial content with an overlaying text introduced a posteriori. Its main aim is originally to be funny and/or ironic, making misogyny recognition in memes even more challenging. In this paper, we investigated 4 unimodal and 3 multimodal approaches to determine which source of information contributes more to the detection of misogynous memes. Moreover, a bias estimation technique is proposed to identify specific elements that compose a meme that could lead to unfair models, together with a bias mitigation strategy based on Bayesian Optimization. The proposed method is able to push the prediction probabilities towards the correct class for up to 61.43% of the cases. Finally, we identified the most challenging archetypes of memes that are still far to be properly recognized, highlighting the most relevant open research directions. Giulia Rizzi, Francesca Gasparini, Aurora Saibene, Paolo Rosso, Elisabetta Fersini |
Inf. Process. Manag. | 4 |
| 2023 | Systematic keyword and bias analyses in hate speech detectionabstractHate speech detection refers broadly to the automatic identification of language that may be considered discriminatory against certain groups of people. The goal is to help online platforms to identify and remove harmful content. Humans are usually capable of detecting hatred in critical cases, such as when the hatred is non-explicit, but how do computer models address this situation? In this work, we aim to contribute to the understanding of ethical issues related to hate speech by analysing two transformer-based models trained to detect hate speech. Our study focuses on analysing the relationship between these models and a set of hateful keywords extracted from the three well-known datasets. For the extraction of the keywords, we propose a metric that takes into account the division among classes to favour the most common words in hateful contexts. In our experiments, we first compared the overlap between the extracted keywords with the words to which the models pay the most attention in decision-making. On the other hand, we investigate the bias of the models towards the extracted keywords. For the bias analysis, we characterize and use two metrics and evaluate two strategies to try to mitigate the bias. Surprisingly, we show that over 50% of the salient words of the models are not hateful and that there is a higher number of hateful words among the extracted keywords. However, we show that the models appear to be biased towards the extracted keywords. Experimental results suggest that fitting models with hateful texts that do not contain any of the keywords can reduce bias and improve the performance of the models. Gretel Liz De la Peña Sarracén, Paolo Rosso |
Inf. Process. Manag. | 2 |
| 2023 | Revisiting Embedding Based Graph Analyses: Hyperparameters Matter!abstractGraph embeddings have been widely used for many graph analysis tasks. Mainstream factorization-based and graph-sampling-based embedding learning schemes both involve many hyperparameters and design choices. However, existing techniques often adopt some heuristics for these hyperparameters and design choices with little investigation into their impact, making it unclear what is the exact performance gains of these techniques on graph analysis tasks. Against this background, this paper presents a systematic study on the impact of an extensive list of hyperparameters for both factorization-based and graph-sampling-based graph embedding techniques for homogeneous graphs. We design generalized factorization-based and graph-sampling-based techniques involving these hyperparameters, and conduct a comprehensive set of experiments with over 3,000 embedding models trained and evaluated per dataset. We reveal that much of the performance gains are indeed due to optimal hyperparameter settings/design choices rather than the sophistication of embedding models; appropriate hyperparameter settings for typical embedding techniques can outperform a sizeable collection of 18 state-of-the-art graph embedding techniques by 0.30-35.41% across different tasks. Moreover, we find that there is no one-size-fits-all hyperparameter setting across tasks, but we can indeed provide a list of task-specific practical recommendations for these hyperparameter settings/design choices, which we believe can serve as important guidelines for future research on embedding based graph analyses. Dingqi Yang, Bingqing Qu, Rana Hussein, Paolo Rosso, Philippe Cudré-Mauroux, Jie Liu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Overview of PAN 2022: Authorship Verification, Profiling Irony and Stereotype Spreaders, Style Change Detection, and Trigger Detection - Extended Abstract
Janek Bevendorff, Berta Chulvi, Elisabetta Fersini, Annina Heini, Mike Kestemont, Krzysztof Kredens, Maximilian Mayerl, Reyner Ortega-Bueno, Piotr Pezik, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Magdalena Wolska, Eva Zangerle |
ECIR (2) | 12 |
| 2022 | Unsupervised Ranking and Aggregation of Label Descriptions for Zero-Shot Classifiers
Angelo Basile, Marc Franco-Salvador, Paolo Rosso |
NLDB | 3 |
| 2022 | Investigating Topic-Agnostic Features for Authorship Tasks in Spanish Political Speeches
Silvia Corbara, Berta Chulvi, Paolo Rosso, Alejandro Moreo |
NLDB | 3 |
| 2022 | Detecting Early Signs of Depression in the Conversational Domain: The Role of Transfer Learning in Low-Resource Scenarios
Petr Lorenc, Ana Sabina Uban, Paolo Rosso, Jan Sedivý |
NLDB | 3 |
| 2022 | Convolutional Graph Neural Networks for Hate Speech Detection in Data-Poor Settings
Gretel Liz De la Peña Sarracén, Paolo Rosso |
NLDB | 2 |
| 2022 | The impact of psycholinguistic patterns in discriminating between fake news spreaders and fact checkersabstractFake news is a threat to society. A huge amount of fake news is posted every day on social networks which is read, believed and sometimes shared by a number of users. On the other hand, with the aim to raise awareness, some users share posts that debunk fake news by using information from fact-checking websites. In this paper, we are interested in exploring the role of various psycholinguistic characteristics in differentiating between users that tend to share fake news and users that tend to debunk them. Psycholinguistic characteristics represent the different linguistic information that can be used to profile users and can be extracted or inferred from users’ posts. We present the CheckerOrSpreader model that uses a Convolution Neural Network (CNN) to differentiate between spreaders and checkers of fake news. The experimental results showed that CheckerOrSpreader is effective in classifying a user as a potential spreader or checker. Our analysis showed that checkers tend to use more positive language and a higher number of terms that show causality compared to spreaders who tend to use a higher amount of informal language, including slang and swear words. Anastasia Giahanou, Bilal Ghanem, Esteban A. Ríssola, Paolo Rosso, Fabio Crestani, Daniel L. Oberski |
Data Knowl. Eng. | 4 |
| 2021 | ConvTab: A Context-Preserving, Convolutional Model for Ad-Hoc Table RetrievalabstractAd-hoc table retrieval, also known as table search, is the problem of finding tables relevant to a search query. This search query can be a keyword or a table itself, referred to as keyword-based and table-based search, respectively. With the vast amounts of tabular data available online, it has become essential for users to identify relevant tables that meet their search criteria. In this regard, there has been a wide variety of research on this problem using pure lexical features, semantic representation, embeddings, as well as intrinsic and extrinsic features of the tables. However, one of the significant limitations of most of the existing methods is that they do not keep the table’s structure and the globalized context intact when building semantic representations of tabular data. Deriving motivation from this fact, we propose an effective approach based on Convolutional Neural Networks (CNNs) – ConvTab – to train the embeddings of tabular data. Our approach is divided into two phases. First, we leverage the discriminating power of CNNs to train a table classifier. Next, the representations learned from this model are used to generate semantic features for query-table similarity. These query-table similarity features are then used as input to the learning algorithm. We evaluate our approach on the table retrieval task using standard NDCG, MAP, and MRR metrics. Experiments reveal that ConvTab significantly outperforms the state of the art in ad-hoc table retrieval by 16.9% and 8.37% using NDCG at cutoffs 5 and 20, respectively. For reproducibility purposes, we share our model as well as all details of our implementation1. Vibhav Agarwal, Akansha Bhardwaj, Paolo Rosso, Philippe Cudré-Mauroux |
IEEE BigData | 3 |
| 2021 | Overview of PAN 2021: Authorship Verification, Profiling Hate Speech Spreaders on Twitter, and Style Change Detection - Extended Abstract
Janek Bevendorff, Berta Chulvi, Gretel Liz De la Peña Sarracén, Mike Kestemont, Enrique Manjavacas, Ilia Markov, Maximilian Mayerl, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Magdalena Wolska, Eva Zangerle |
ECIR (2) | 10 |
| 2021 | Profiling Fake News Spreaders: Personality and Visual Information Matter
Riccardo Cervero, Paolo Rosso, Gabriella Pasi |
NLDB | 2 |
| 2021 | On the Generalization of Figurative Language Detection: The Case of Irony and Sarcasm
Lorenzo Famiglini, Elisabetta Fersini, Paolo Rosso |
NLDB | 3 |
| 2021 | On the Explainability of Automatic Predictions of Mental Disorders from Social Media Data
Ana Sabina Uban, Berta Chulvi, Paolo Rosso |
NLDB | 3 |
| 2021 | RETA: A Schema-Aware, End-to-End Solution for Instance Completion in Knowledge GraphsabstractKnowledge Graph (KG) completion has been widely studied to tackle the incompleteness issue (i.e., missing facts) in modern KGs. A fact in a KG is represented as a triplet (h, r, t) linking two entities h and t via a relation r. Existing work mostly consider link prediction to solve this problem, i.e., given two elements of a triplet predicting the missing one, such as (h, r, ?). This task has, however, a strong assumption on the two given elements in a triplet, which have to be correlated, resulting otherwise in meaningless predictions, such as (Marie Curie, headquarters location, ?). In addition, the KG completion problem has also been formulated as a relation prediction task, i.e., when predicting relations r for a given entity h. Without predicting t, this task is however a step away from the ultimate goal of KG completion. Against this background, this paper studies an instance completion task suggesting r-t pairs for a given h, i.e., (h, ?, ?). We propose an end-to-end solution called RETA (as it suggests the Relation and Tail for a given head entity) consisting of two components: a RETA-Filter and RETA-Grader. More precisely, our RETA-Filter first generates candidate r-t pairs for a given h by extracting and leveraging the schema of a KG; our RETA-Grader then evaluates and ranks the candidate r-t pairs considering the plausibility of both the candidate triplet and its corresponding schema using a newly-designed KG embedding model. We evaluate our methods against a sizable collection of state-of-the-art techniques on three real-world KG datasets. Results show that our RETA-Filter generates of high-quality candidate r-t pairs, outperforming the best baseline techniques while reducing by 10.61%-84.75% the candidate size under the same candidate quality guarantees. Moreover, our RETA-Grader also significantly outperforms state-of-the-art link prediction techniques on the instance completion task by 16.25%-65.92% across different datasets. Paolo Rosso, Dingqi Yang, Natalia Ostapuk, Philippe Cudré-Mauroux |
WWW | 1 |
| 2021 | Detecting ethnicity-targeted hate speech in Russian social media texts
Ekaterina V. Pronoza, Polina Panicheva, Olessia Koltsova, Paolo Rosso |
Inf. Process. Manag. | 4 |
| 2021 | The impact of emotional signals on credibility assessmentabstractFake news is considered one of the main threats of our society. The aim of fake news is usually to confuse readers and trigger intense emotions to them in an attempt to be spread through social networks. Even though recent studies have explored the effectiveness of different linguistic patterns for fake news detection, the role of emotional signals has not yet been explored. In this paper, we focus on extracting emotional signals from claims and evaluating their effectiveness on credibility assessment. First, we explore different methodologies for extracting the emotional signals that can be triggered to the users when they read a claim. Then, we present emoCred, a model that is based on a long-short term memory model that incorporates emotional signals extracted from the text of the claims to differentiate between credible and non-credible ones. In addition, we perform an analysis to understand which emotional signals and which terms are the most useful for the different credibility classes. We conduct extensive experiments and a thorough analysis on real-world datasets. Our results indicate the importance of incorporating emotional signals in the credibility assessment problem. Anastasia Giahanou, Paolo Rosso, Fabio Crestani |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2020 | The Battle Against Online Harmful Information: The Cases of Fake News and Hate SpeechabstractSocial media have given the opportunity to users to express their opinions online in a fast and easy way. The ease of generating content online and the anonymity that social media provide have increased the amount of harmful content that is published. This tutorial will focus on the topic of online harmful information. First, we will analyse and explain the different types of online harmful information with a particular focus on fake news and hate speech. In addition, we will explain the different computational approaches proposed in the literature for the detection of fake news and hate speech. Next, we will present details regarding the evaluation process, datasets and shared tasks and finally, we will discuss future directions in the field of online harmful information detection. Anastasia Giahanou, Paolo Rosso |
CIKM | 2 |
| 2020 | Multimodal Multi-image Fake News DetectionabstractRecent years have seen a large increase in the amount of false information that is posted online. Fake news are created and propagated in order to deceive users and manipulate opinions and subsequently have a negative impact on the society. The automatic detection of fake news is very challenging since some of those news are created in sophisticated ways containing text or images that have been deliberately modified. Combining information from different modalities can be very useful for determining which of the online articles are fake. In this paper, we propose a multimodal multi-image system that combines information from different modalities in order to detect fake news posted online. In particular, our system combines textual, visual and semantic information. For the textual representation we use the Bidirectional Encoder Representations from Transformers (BERT) to better capture the underlying semantic and contextual meaning of the text. For the visual representation we extract image tags from multiple images that the articles contain using the VGG-16 model. The semantic representation refers to the text-image similarity calculated using the cosine similarity between the title and image tags embeddings. Our experimental results on a real world dataset show that combining features from the different modalities is effective for fake news detection. In particular, our multimodal multi-image system significantly outperforms the BERT baseline by 4.19% and SpotFake by 5.39%. Anastasia Giahanou, Guobiao Zhang, Paolo Rosso |
DSAA | 3 |
| 2020 | Shared Tasks on Authorship Analysis at PAN 2020
Janek Bevendorff, Bilal Ghanem, Anastasia Giahanou, Mike Kestemont, Enrique Manjavacas, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Günther Specht, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Eva Zangerle |
ECIR (2) | 8 |
| 2020 | Irony Detection in a Multilingual Context
Bilal Ghanem, Jihen Karoui, Farah Benamara, Paolo Rosso, Véronique Moriceau |
ECIR (2) | 4 |
| 2020 | The Role of Personality and Linguistic Patterns in Discriminating Between Fake News Spreaders and Fact CheckersabstractUsers play a critical role in the creation and propagation of fake news online by consuming and sharing articles with inaccurate information either intentionally or unintentionally. Fake news are written in a way to confuse readers and therefore understanding which articles contain fabricated information is very challenging for non-experts. Given the difficulty of the task, several fact checking websites have been developed to raise awareness about which articles contain fabricated information. As a result of those platforms, several users are interested to share posts that cite evidence with the aim to refute fake news and warn other users. These users are known as fact checkers . However, there are users who tend to share false information, who can be characterised as potential fake news spreaders . In this paper, we propose the CheckerOrSpreader model that can classify a user as a potential fact checker or a potential fake news spreader. Our model is based on a Convolutional Neural Network (CNN) and combines word embeddings with features that represent users’ personality traits and linguistic patterns used in their tweets. Experimental results show that leveraging linguistic patterns and personality traits can improve the performance in differentiating between checkers and spreaders. Anastasia Giahanou, Esteban A. Ríssola, Bilal Ghanem, Fabio Crestani, Paolo Rosso |
NLDB | 5 |
| 2020 | Beyond Triplets: Hyper-Relational Knowledge Graph Embedding for Link PredictionabstractKnowledge Graph (KG) embeddings are a powerful tool for predicting missing links in KGs. Existing techniques typically represent a KG as a set of triplets, where each triplet (h, r, t) links two entities h and t through a relation r, and learn entity/relation embeddings from such triplets while preserving such a structure. However, this triplet representation oversimplifies the complex nature of the data stored in the KG, in particular for hyper-relational facts, where each fact contains not only a base triplet (h, r, t), but also the associated key-value pairs (k, v). Even though a few recent techniques tried to learn from such data by transforming a hyper-relational fact into an n-ary representation (i.e., a set of key-value pairs only without triplets), they result in suboptimal models as they are unaware of the triplet structure, which serves as the fundamental data structure in modern KGs and preserves the essential information for link prediction. To address this issue, we propose HINGE, a hyper-relational KG embedding model, which directly learns from hyper-relational facts in a KG. HINGE captures not only the primary structural information of the KG encoded in the triplets, but also the correlation between each triplet and its associated key-value pairs. Our extensive evaluation shows the superiority of HINGE on various link prediction tasks over KGs. In particular, HINGE consistently outperforms not only the KG embedding methods learning from triplets only (by 0.81-41.45% depending on the link prediction tasks and settings), but also the methods learning from hyper-relational facts using the n-ary representation (by 13.2-84.1%). Paolo Rosso, Dingqi Yang, Philippe Cudré-Mauroux |
WWW | 1 |
| 2019 | Revisiting Text and Knowledge Graph Joint Embeddings: The Amount of Shared Information Matters!abstractJointly learning embeddings from text and a Knowledge Graph benefits both word and entity/relation embeddings by taking advantage of both large-scale unstructured content (text) and high-quality structured data (the Knowledge Graph). Current techniques leverage anchors to associate entities in the Knowledge Graph to corresponding words in the text corpus; these anchors are then used to generate additional learning samples during the embedding learning process. However, we show in this paper that such techniques yield suboptimal results, as they fail to control the amount of shared information between the two data sources during the joint learning process. Moreover, the additional learning samples often incur significant computational overhead. Aiming at releasing the power of such joint embeddings, we propose JOINER, a new joint text and Knowledge Graph embedding method using regularization. JOINER not only preserves co-occurrence between words in a text corpus and relations between entities in a Knowledge Graph, it also provides the flexibility to control the amount of information shared between the two data sources via regularization. Our method does not generate additional learning samples, which makes it computationally efficient. Our extensive empirical evaluation on real datasets shows the superiority of JOINER across different evaluation tasks, including analogical reasoning, link prediction, and relation extraction. Compared to state-of-the-art techniques generating additional learning samples from a set of anchors, our method yields better results (with up to 4.3% absolute improvement) and significantly less computational overhead (76% less learning time overhead). Paolo Rosso, Dingqi Yang, Philippe Cudré-Mauroux |
IEEE BigData | 1 |
| 2019 | A Decade of Shared Tasks in Digital Text Forensics at PAN
Martin Potthast, Paolo Rosso, Efstathios Stamatatos, Benno Stein 0001 |
ECIR (2) | 2 |
| 2019 | NodeSketch: Highly-Efficient Graph Embeddings via Recursive SketchingabstractEmbeddings have become a key paradigm to learn graph representations and facilitate downstream graph analysis tasks. Existing graph embedding techniques either sample a large number of node pairs from a graph to learn node embeddings via stochastic optimization, or factorize a high-order proximity/adjacency matrix of the graph via expensive matrix factorization. However, these techniques usually require significant computational resources for the learning process, which hinders their applications on large-scale graphs. Moreover, the cosine similarity preserved by these techniques shows suboptimal efficiency in downstream graph analysis tasks, compared to Hamming similarity, for example. To address these issues, we propose NodeSketch, a highly-efficient graph embedding technique preserving high-order node proximity via recursive sketching. Specifically, built on top of an efficient data-independent hashing/sketching technique, NodeSketch generates node embeddings in Hamming space. For an input graph, it starts by sketching the self-loop-augmented adjacency matrix of the graph to output low-order node embeddings, and then recursively generates k-order node embeddings based on the self-loop-augmented adjacency matrix and (k-1)-order node embeddings. Our extensive evaluation compares NodeSketch against a sizable collection of state-of-the-art techniques using five real-world graphs on two graph analysis tasks. The results show that NodeSketch achieves state-of-the-art performance compared to these techniques, while showing significant speedup of 9x-372x in the embedding learning process and 1.19x-1.68x speedup when performing downstream graph analysis tasks. Dingqi Yang, Paolo Rosso, Bin Li 0015, Philippe Cudré-Mauroux |
KDD | 2 |
| 2019 | Leveraging Emotional Signals for Credibility DetectionabstractThe spread of false information on the Web is one of the main problems of our society. Automatic detection of fake news posts is a hard task since they are intentionally written to mislead the readers and to trigger intense emotions to them in an attempt to be disseminated in the social networks. Even though recent studies have explored different linguistic patterns of false claims, the role of emotional signals has not yet been explored. In this paper, we study the role of emotional signals in fake news detection. In particular, we propose an LSTM model that incorporates emotional signals extracted from the text of the claims to differentiate between credible and non-credible ones. Experiments on real world datasets show the importance of emotional signals for credibility assessment. Anastasia Giahanou, Paolo Rosso, Fabio Crestani |
SIGIR | 2 |
| 2019 | Stance polarity in political debates: A diachronic perspective of network homophily and conversations on Twitter
Mirko Lai, Marcella Tambuscio, Viviana Patti, Giancarlo Ruffo, Paolo Rosso |
Data Knowl. Eng. | 5 |
| 2019 | Irony detection via sentiment-based transfer learning
Xiuzhen Zhang 0001, Jeffrey Chan, Paolo Rosso |
Inf. Process. Manag. | 4 |
| 2018 | Emotional Influence Prediction of News Posts
Anastasia Giahanou, Paolo Rosso, Ida Mele, Fabio Crestani |
ICWSM | 2 |
| 2018 | Automatic Identification and Classification of Misogynistic Language on Twitter
Maria Anzovino, Elisabetta Fersini, Paolo Rosso |
NLDB | 3 |
| 2018 | HYPLAG: Hybrid Arabic Text Plagiarism Detection System
Bilal Ghanem, Labib Arafeh, Paolo Rosso, Fernando Sánchez-Vega |
NLDB | 3 |
| 2018 | String Kernels for Polarity Classification: A Study Across Different Languages
Rosa M. Giménez-Pérez, Marc Franco-Salvador, Paolo Rosso |
NLDB | 3 |
| 2018 | Stance Evolution and Twitter Interactions in an Italian Political Debate
Mirko Lai, Viviana Patti, Giancarlo Ruffo, Paolo Rosso |
NLDB | 4 |
| 2018 | Identifying and Classifying Influencers in Twitter only with Textual Information
Victoria Nebot, Francisco M. Rangel Pardo, Rafael Berlanga Llavori, Paolo Rosso |
NLDB | 4 |
| 2018 | Early Commenting Features for Emotional Reactions Prediction
Anastasia Giahanou, Paolo Rosso, Ida Mele, Fabio Crestani |
SPIRE | 2 |
| 2017 | Continuous space models for CLIR
Parth Gupta, Rafael E. Banchs, Paolo Rosso |
Inf. Process. Manag. | 3 |
| 2016 | MultiLingMine 2016: Modeling, Learning and Mining for Cross/Multilinguality
Dino Ienco, Mathieu Roche, Salvatore Romeo, Paolo Rosso, Andrea Tagarelli |
ECIR | 4 |
| 2016 | Arabic WordNet: New Content and New ApplicationsabstractThe Arabic WordNet project has provided the Arabic Natural Language Processing (NLP) community with the first WordNet-compliant resource.It allowed new possibilities in terms of building sophisticated NLP applications related to this Semitic language.In this paper, we present the new content added to this resource, using semi-automatic techniques, and validated by Arabic native-speaker lexicographers.We also present how this content helps in the implementation of new Arabic NLP applications, especially for Question Answering (QA) systems.The obtained results show the contribution of the added content.The resource, fully transformed into the standard Lexical Markup Framework (LMF), is made available for the community. Yasser Regragui, Lahsen Abouenour, Fettoum Krieche, Karim Bouzoubaa, Paolo Rosso |
GWC | 5 |
| 2016 | A systematic study of knowledge graph analysis for cross-language plagiarism detection
Marc Franco-Salvador, Paolo Rosso, Manuel Montes-y-Gómez |
Inf. Process. Manag. | 2 |
| 2016 | On the impact of emotions on author profiling
Francisco M. Rangel Pardo, Paolo Rosso |
Inf. Process. Manag. | 2 |
| 2016 | Emotion and sentiment in social and expressive media: Introduction to the special issue
Paolo Rosso, Cristina Bosco, Rossana Damiano, Viviana Patti, Erik Cambria |
Inf. Process. Manag. | 1 |
| 2016 | Comparing and combining Content- and Citation-based approaches for plagiarism detectionabstractThe vast amount of scientific publications available online makes it easier for students and researchers to reuse text from other authors and makes it harder for checking the originality of a given text. Reusing text without crediting the original authors is considered plagiarism. A number of studies have reported the prevalence of plagiarism in academia. As a consequence, numerous institutions and researchers are dedicated to devising systems to automate the process of checking for plagiarism. This work focuses on the problem of detecting text reuse in scientific papers. The contributions of this paper are twofold: (a) we survey the existing approaches for plagiarism detection based on content, based on content and structure, and based on citations and references; and (b) we compare content and citation‐based approaches with the goal of evaluating whether they are complementary and if their combination can improve the quality of the detection. We carry out experiments with real data sets of scientific papers and concluded that a combination of the methods can be beneficial. Solange de L. Pertile, Viviane Pereira Moreira, Paolo Rosso |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2015 | Detecting positive and negative deceptive opinions using PU-learning
Donato Hernández-Fusilier, Manuel Montes-y-Gómez, Paolo Rosso, Rafael Guzmán-Cabrera |
Inf. Process. Manag. | 3 |
| 2014 | A Comparison of Approaches for Measuring Cross-Lingual Similarity of Wikipedia Articles
Alberto Barrón-Cedeño, Monica Lestari Paramita, Paul D. Clough, Paolo Rosso |
ECIR | 4 |
| 2014 | Query expansion for mixed-script information retrievalabstractFor many languages that use non-Roman based indigenous scripts (e.g., Arabic, Greek and Indic languages) one can often find a large amount of user generated transliterated content on the Web in the Roman script. Such content creates a monolingual or multi-lingual space with more than one script which we refer to as the Mixed-Script space. IR in the mixed-script space is challenging because queries written in either the native or the Roman script need to be matched to the documents written in both the scripts. Moreover, transliterated content features extensive spelling variations. In this paper, we formally introduce the concept of Mixed-Script IR, and through analysis of the query logs of Bing search engine, estimate the prevalence and thereby establish the importance of this problem. We also give a principled solution to handle the mixed-script term matching and spelling variation where the terms across the scripts are modelled jointly in a deep-learning architecture and can be compared in a low-dimensional abstract space. We present an extensive empirical analysis of the proposed method along with the evaluation results in an ad-hoc retrieval setting of mixed-script IR where the proposed method achieves significantly better results (12% increase in MRR and 29% increase in MAP) compared to other state-of-the-art baselines. Parth Gupta, Kalika Bali, Rafael E. Banchs, Monojit Choudhury, Paolo Rosso |
SIGIR | 5 |
| 2014 | An efficient Particle Swarm Optimization approach to cluster short texts
Leticia C. Cagnina, Marcelo Luis Errecalde, Diego Ingaramo, Paolo Rosso |
Inf. Sci. | 4 |
| 2014 | On the difficulty of automatically detecting irony: beyond a simple case of negation
Antonio Reyes, Paolo Rosso |
Knowl. Inf. Syst. | 2 |
| 2013 | Cross-Language Plagiarism Detection Using a Multilingual Semantic Network
Marc Franco-Salvador, Parth Gupta, Paolo Rosso |
ECIR | 3 |
| 2013 | Identifying subjective statements in news titles using a personal sense annotation frameworkabstractSubjective language contains information about private states. The goal of subjective language identification is to determine that a private state is expressed, without considering its polarity or specific emotion. A component of word meaning, “Personal Sense,” has clear potential in the field of subjective language identification, as it reflects a meaning of words in terms of unique personal experience and carries personal characteristics. In this paper we investigate how Personal Sense can be harnessed for the purpose of identifying subjectivity in news titles. In the process, we develop a new Personal Sense annotation framework for annotating and classifying subjectivity, polarity, and emotion. The Personal Sense framework yields high performance in a fine‐grained subsentence subjectivity classification. Our experiments demonstrate lexico‐syntactic features to be useful for the identification of subjectivity indicators and the targets that receive the subjective Personal Sense. Polina Panicheva, John Cardiff, Paolo Rosso |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2012 | From humor recognition to irony detection: The figurative language of social media
Antonio Reyes, Paolo Rosso, Davide Buscaldi |
Data Knowl. Eng. | 2 |
| 2012 | Introduction to the Special Section on Search and Mining User-Generated ContentabstractThe primary goal of this special section of ACM Transactions on Intelligent Systems and Technology is to foster research in the interplay between Social Media, Data/Opinion Mining and Search, aiming to reflect the actual developments in technologies that exploit user-generated content. José Carlos Cortizo, Francisco M. Carrero, Iván Cantador, José Antonio Troyano Jiménez, Paolo Rosso |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2011 | Overview of the third international workshop on search and mining user-generated contentsabstractIn this paper, we provide an overview of the 3rd International Workshop on Search and Mining User-generated Contents, held in conjunction with the 20th ACM International Conference on Information and Knowledge Management. We present the motivation and goals of the workshop, and some statistics and details about accepted papers and keynotes. Iván Cantador, José Carlos Cortizo, Francisco M. Carrero, José Antonio Troyano Jiménez, Paolo Rosso, Markus Schedl |
CIKM | 5 |
| 2011 | Towards the Detection of Cross-Language Source Code Reuse
Enrique Flores, Alberto Barrón-Cedeño, Paolo Rosso, Lidia Moreno |
NLDB | 3 |
| 2011 | POS Tagging in Amazighe Using Support Vector Machines and Conditional Random Fields
Mohamed Outahajala, Yassine Benajiba, Paolo Rosso, Lahbib Zenkouar |
NLDB | 3 |
| 2010 | Overview of the 2nd international workshop on search and mining user-generated contentsabstractThis overview introduces the aim of the SMUC 2010 workshop, as well as the list of papers presented in the workshop. José Carlos Cortizo, Francisco M. Carrero, Iván Cantador, José Antonio Troyano Jiménez, Paolo Rosso |
CIKM | 5 |
| 2010 | On the Extension of Arabic Wordnet Named Entities and Its Impact on Question / Answering
Lahsen Abouenour, Karim Bouzoubaa, Paolo Rosso |
KEOD | 3 |
| 2010 | Identifying Writers' Background by Comparing Personal Sense Thesauri
Polina Panicheva, John Cardiff, Paolo Rosso |
NLDB | 3 |
| 2010 | An Automatic Definition Extraction in Arabic Language
Omar Trigui, Lamia Hadrich Belguith, Paolo Rosso |
NLDB | 3 |
| 2010 | Answering questions with an n-gram based passage retrieval engine
Davide Buscaldi, Paolo Rosso, José Manuel Gómez Soriano, Emilio Sanchis Arnal |
J. Intell. Inf. Syst. | 2 |
| 2010 | Automatic Ontology Matching via Upper Ontologies: A Systematic Evaluationabstract“Ontology matching” is the process of finding correspondences between entities belonging to different ontologies. This paper describes a set of algorithms that exploit upper ontologies as semantic bridges in the ontology matching process and presents a systematic analysis of the relationships among features of matched ontologies (number of simple and composite concepts, stems, concepts at the top level, common English suffixes and prefixes, and ontology depth), matching algorithms, used upper ontologies, and experiment results. This analysis allowed us to state under which circumstances the exploitation of upper ontologies gives significant advantages with respect to traditional approaches that do no use them. We run experiments with SUMO-OWL (a restricted version of SUMO), OpenCyc, and DOLCE. The experiments demonstrate that when our “structural matching method via upper ontology” uses an upper ontology large enough (OpenCyc, SUMO-OWL), the recall is significantly improved while preserving the precision obtained without upper ontologies. Instead, our “nonstructural matching method” via OpenCyc and SUMO-OWL improves the precision and maintains the recall. The “mixed method” that combines the results of structural alignment without using upper ontologies and structural alignment via upper ontologies improves the recall and maintains the F-measure independently of the used upper ontology. Viviana Mascardi, Angela Locoro, Paolo Rosso |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2009 | On Automatic Plagiarism Detection Based on n-Grams Comparison
Alberto Barrón-Cedeño, Paolo Rosso |
ECIR | 2 |
| 2009 | Characterizing Weblog Corpora
Fernando Pérez-Téllez, David Pinto 0001, John Cardiff, Paolo Rosso |
NLDB | 4 |
| 2009 | On the Assessment of Text Corpora
David Pinto 0001, Paolo Rosso, Héctor Jiménez-Salazar |
NLDB | 2 |
| 2009 | The Impact of Semantic and Morphosyntactic Ambiguity on Automatic Humour Recognition
Antonio Reyes, Davide Buscaldi, Paolo Rosso |
NLDB | 3 |
| 2009 | Using the Web as corpus for self-training text categorization
Rafael Guzmán-Cabrera, Manuel Montes-y-Gómez, Paolo Rosso, Luis Villaseñor-Pineda |
Inf. Retr. | 3 |
| 2008 | A conceptual density-based approach for the disambiguation of toponymsabstractNowadays, a huge quantity of information is stored in digital format. A great portion of this information is constituted by textual and unstructured documents, where geographical references are usually given by means of place names. A common problem with textual information retrieval is represented by polysemous words, that is, words can have more than one sense. This problem is present also in the geographical domain: place names may refer to different locations in the world. In this paper we investigate the use of our word sense disambiguation technique in the geographical domain, with the aim of resolving ambiguous place names. Our technique is based on WordNet conceptual density. Due to the lack of a reference corpus tagged with WordNet senses, we carried out the experiments over a set of 1,210 place names extracted from the SemCor corpus that we named GeoSemCor and made publicly available. We compared our method with the most‐frequent baseline and the enhanced‐Lesk method, which previously has not been tested in large contexts. The results show that a better precision can be achieved by using a small context (phrase level), whereas a greater coverage can be obtained by using large contexts (document level). The proposed method should be tested with other corpora, due to the fact that our experiments evidenced the excessive bias towards the most‐frequent sense of the GeoSemCor. Davide Buscaldi, Paolo Rosso |
Int. J. Geogr. Inf. Sci. | 2 |
| 2007 | Biomedical Named Entity Recognition: A Poor Knowledge HMM-Based Approach
Ferran Plà, Antonio Molina, Paolo Rosso |
NLDB | 4 |
| 2005 | An Approach to Clustering Abstracts
Mikhail Alexandrov, Alexander F. Gelbukh, Paolo Rosso |
NLDB | 3 |