Jose M. Such

dblp:18/4101 · also Jose Miguel Such, Jose Such · DBLP profile ↗
← Back
12ranked-venue papers in the field
2as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 5Data Mining & Knowledge Discovery · 3Database Systems & Data Management · 2 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 2 (1 first)
YearPublicationVenuePosition
2025 Cross-Partisan Interactions on Twitter
abstract
Many social media studies argue that social media creates echo chambers where some users only interact with peers of the same political orientation. However, recent studies suggest that a substantial amount of Cross-Partisan Interactions (CPIs) do exist --- even within echo chambers, but they may be toxic. There is no consensus about how such interactions occur and when they lead to healthy or toxic dialogue. In this paper, we study a comprehensive Twitter dataset that consists of 3 million tweets from 2020 related to the U.S. context to understand the dynamics behind CPIs. We investigate factors that are more associated with such interactions, including how users engage in CPIs, which topics are more contentious, and what are the stances associated with healthy interactions. We find that CPIs are significantly influenced by the nature of the topics being discussed, with politically charged events acting as strong catalysts. The political discourse and pre-established political views sway how users participate in CPIs, but the direction in which users go is nuanced. While Democrats engage in cross-partisan interactions slightly more frequently, these interactions often involve more negative and nonconstructive stances compared to their intra-party interactions. In contrast, Republicans tend to maintain a more consistent tone across interactions. Although users are more likely to engage in CPIs with popular accounts in general, this is less common among Republicans who often engage in CPIs with accounts with a low number of followers for personal matters. Our study has implications beyond Twitter as identifying topics with low toxicity and high CPI can help highlight potential opportunities for reducing polarization while topics with high toxicity and low CPI may action targeted interventions when moderating harm.
Yusuf Mücahit Çetinkaya, Vahid Ghafouri, Guillermo Suarez-Tangil, Jose M. Such, Tugrulcan Elmas
ICWSM4
2024 Differences in the Toxic Language of Cross-Platform Communities
abstract
Cross-platform communities are social media communities that have a presence on multiple online platforms. One active community on both Reddit and Discord is dankmemes. Our study aims to examine differences in harmful language usage across different platforms in a community. We scrape 15 communities that are active on both Reddit and Discord. We then identify and compare differences in type and level of toxicity, in the topics of the harmful discourse, in the temporal evolution of toxicity and its attribution to users, and in the moderation strategies communities across platforms. Our results show that most communities exhibit differences in toxicity depending on the platform. We see that toxicity is rooted in the different subcultures as well as in the way in which the platforms operate and their administrators moderate content. However, we note that in general terms Discord is significantly more toxic than Reddit. We offer a detailed analysis of the topics and types of communities in which this happens and why, which will help moderators and policymakers shape their strategies to mitigate the harm on the Web. In particular, we propose practical and effective strategies that Discord can implement to improve its platform moderation.
Ashwini Kumar Singh, Vahid Ghafouri, Jose M. Such, Guillermo Suarez-Tangil
ICWSM3
2023 AI in the Gray: Exploring Moderation Policies in Dialogic Large Language Models vs. Human Answers in Controversial Topics
abstract
The introduction of ChatGPT and the subsequent improvement of Large Language Models (LLMs) have prompted more and more individuals to turn to the use of ChatBots, both for information and assistance with decision-making. However, the information the user is after is often not formulated by these ChatBots objectively enough to be provided with a definite, globally accepted answer.
Vahid Ghafouri, Vibhor Agarwal, Nishanth Sastry, Jose M. Such, Guillermo Suarez-Tangil
CIKM5
2023 Discovering and Interpreting Biased Concepts in Online Communities
abstract
Language carries implicit human biases, functioning both as a reflection and a perpetuation of stereotypes that people carry with them. Recently, ML-based NLP methods such as word embeddings have been shown to learn such language biases with striking accuracy. This capability of word embeddings has been successfully exploited as a tool to quantify and study human biases. However, previous studies only consider a predefined set of biased concepts to attest (e.g., whether gender is more or less associated with particular jobs), or just discover biased words without helping to understand their meaning at the conceptual level. As such, these approaches can be either unable to find biased concepts that have not been defined in advance, or the biases they find are difficult to interpret and study. This could make existing approaches unsuitable to discover and interpret biases in online communities, as such communities may carry different biases than those in mainstream culture. This paper improves upon, extends, and evaluates our previous data-driven method to automatically discover and help interpret biased concepts encoded in word embeddings. We apply this approach to study the biased concepts present in the language used in online communities and experimentally show the validity and stability of our method.
Xavier Ferrer Aran, Tom van Nuenen, Natalia Criado, Jose M. Such
IEEE Trans. Knowl. Data Eng.4
2022 Measuring Alexa Skill Privacy Practices across Three Years
abstract
Smart Voice Assistants are transforming the way users interact with technology. This transformation is mostly fostered by the proliferation of voice-driven applications (called skills) offered by third-party developers through an online market. We see how the number of skills has rocked in recent years, with the Amazon Alexa skill ecosystem growing from just 135 skills in early 2016 to about 125k skills in early 2021. Along with the growth in skills, there is increasing concern over the risks that third-party skills pose to users’ privacy. In this paper, we perform a systematic and longitudinal measurement study of the Alexa marketplace. We shed light on how this ecosystem evolves using data collected across three years between 2019 and 2021. We demystify developers’ data disclosure practices and present an overview of the third-party ecosystem. We see how the research community continuously contribute to the market’s sanitation, but the Amazon vetting process still requires significant improvement. We perform a responsible disclosure process reporting 675 skills with privacy issues to both Amazon and all affected developers, out of which 246 skills suffer from important issues (i.e., broken traceability). We see that 107 out of the 246 (43.5%) skills continue to display broken traceability almost one year after being reported. As a result, the overall state of affairs has improved in the ecosystem over the years. Yet, newly submitted skills and unresolved known issues pose an endemic risk.
Jide S. Edu, Xavier Ferrer Aran, Jose M. Such, Guillermo Suarez-Tangil
WWW3
2021 Discovering and Categorising Language Biases in Reddit
Xavier Ferrer Aran, Tom van Nuenen, Jose M. Such, Natalia Criado
ICWSM3
2017 REACT: REcommending Access Control decisions To social media users
abstract
The problems that social media users have in appropriately controlling access to their content has been well documented in previous research. A promising method of providing assistance to users is by learning from the access control decisions made by them and making future recommendations. In this paper, we present REACT, a learning mechanism which utilizes information available in the social network in conjunction with information about the content to be shared to provide users with access control recommendations. We demonstrate the highly accurate performance of REACT through a detailed empirical evaluation and also discuss ways of personalizing it for different users in order to improve performance even further.
Gaurav Misra, Jose M. Such
ASONAM2
2017 A Privacy Assessment of Social Media Aggregators
abstract
Social Media Aggregator (SMA) applications present a platform enabling users to manage multiple Social Networking Sites (SNS) in one convenient application, which results in a unique concentration of data from several SNS accounts in addition to the user's mobile phone data available to them. In this paper, we provide a detailed privacy assessment of 13 popular SMAs from 3 app stores by using a three-step methodology by inspecting the mobile data and social media data accessed by these applications, checking for privacy policies and their compliance with distributors' vetting policies and performing a qualitative assessment of traceability between privacy policies and the actual transparency and control mechanisms offered to users by the apps' interfaces. Our results demonstrate a variation in data accessed by the individual applications, an absence of privacy policies for 5 of the SMAs evaluated, and a lack of traceability between privacy policies and transparency and control of interface operations.
Gaurav Misra, Jose M. Such, Lauren Gill
ASONAM2
2016 Non-sharing communities? An empirical study of community detection for access control decisions
abstract
Social media users often find it difficult to make appropriate access control decisions which govern how they share their information with a potentially large audience on these platforms. Community detection algorithms have been previously put forth as a solution which can help users by automatically partitioning their friend network. These partitions can then be used by the user as a basis for making access control decisions. Previous works which leverage communities for enhancing access control mechanisms assume that members of the same community will have the same access to a user's content, but whether or to what extent this assumption is correct is a lingering question. In this paper, we empirically evaluate a goodness of fit between the communities created by implementing 8 community detection algorithms on the friend networks of users and the access control decisions made by them during a user study. We also analyze whether personal characteristics of the users or the nature of the content play a role in the performance of the algorithms. The results indicate that community detection algorithms may be useful for creating default access control policies for users who exhibit a relatively more static access control behaviour. For users showing great variation in their access control decisions across the board (both in terms of number and actual members), we found that community detection algorithms performed poorly.
Gaurav Misra, Jose M. Such, Hamed Balogun
ASONAM2
2016 Resolving Multi-Party Privacy Conflicts in Social Media
abstract
Items shared through Social Media may affect more than one user's privacy-e.g., photos that depict multiple users, comments that mention multiple users, events in which multiple users are invited, etc. The lack of multi-party privacy management support in current mainstream Social Media infrastructures makes users unable to appropriately control to whom these items are actually shared or not. Computational mechanisms that are able to merge the privacy preferences of multiple users into a single policy for an item can help solve this problem. However, merging multiple users' privacy preferences is not an easy task, because privacy preferences may conflict, so methods to resolve conflicts are needed. Moreover, these methods need to consider how users' would actually reach an agreement about a solution to the conflict in order to propose solutions that can be acceptable by all of the users affected by the item to be shared. Current approaches are either too demanding or only consider fixed ways of aggregating privacy preferences. In this paper, we propose the first computational mechanism to resolve conflicts for multi-party privacy management in Social Media that is able to adapt to different situations by modelling the concessions that users make to reach a solution to the conflicts. We also present results of a user study in which our proposed mechanism outperformed other existing approaches in terms of how many times each approach matched users' behaviour.
Jose M. Such, Natalia Criado
IEEE Trans. Knowl. Data Eng.1
2015 Implicit Contextual Integrity in Online Social Networks
Natalia Criado, Jose M. Such
Inf. Sci.2
2012 Self-disclosure decision making based on intimacy and privacy
Jose M. Such, Agustín Espinosa Minguet, Ana García-Fornes, Carles Sierra
Inf. Sci.1