Xavier Ferrer Aran

dblp:120/5874 · also Xavier Ferrer 0001 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
4since 2021 · last 2023
0000-0001-7278-7373ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2023 SkillVet: Automated Traceability Analysis of Amazon Alexa Skills
abstract
Skills, are essential components in Smart Personal Assistants (SPA). The number of skills has grown rapidly, dominated by a changing environment that has no clear business model. Skills can access personal information and this may pose a risk to users. However, there is little information about how this ecosystem works, let alone the tools that can facilitate its study. In this article, we present the largest systematic measurement of the Amazon Alexa skill ecosystem to date. We study developers’ practices in this ecosystem, including how they collect and justify the need for sensitive information, by designing a methodology to identify over-privileged skills with broken privacy policies. We collect 199,295 Alexa skills and uncover that around 43% of the skills (and 50% of the developers) that request these permissions follow bad privacy practices, including (partially) broken data permissions traceability. In order to perform this kind of analysis at scale, we presentSkillVetthat leverages machine learning and natural language processing techniques, and generates high-accuracy prediction sets. We report several concerning practices, including how developers can bypass Alexa's permission system through account linking and conversational skills, and offer recommendations on how to improve transparency, privacy and security. Resulting from the responsible disclosure we did, 13% of the reported issues no longer pose a threat at submission time.
Jide S. Edu, Xavier Ferrer Aran, Jose M. Such, Guillermo Suarez-Tangil
IEEE Trans. Dependable Secur. Comput.2
2023 Discovering and Interpreting Biased Concepts in Online Communities
abstract
Language carries implicit human biases, functioning both as a reflection and a perpetuation of stereotypes that people carry with them. Recently, ML-based NLP methods such as word embeddings have been shown to learn such language biases with striking accuracy. This capability of word embeddings has been successfully exploited as a tool to quantify and study human biases. However, previous studies only consider a predefined set of biased concepts to attest (e.g., whether gender is more or less associated with particular jobs), or just discover biased words without helping to understand their meaning at the conceptual level. As such, these approaches can be either unable to find biased concepts that have not been defined in advance, or the biases they find are difficult to interpret and study. This could make existing approaches unsuitable to discover and interpret biases in online communities, as such communities may carry different biases than those in mainstream culture. This paper improves upon, extends, and evaluates our previous data-driven method to automatically discover and help interpret biased concepts encoded in word embeddings. We apply this approach to study the biased concepts present in the language used in online communities and experimentally show the validity and stability of our method.
Xavier Ferrer Aran, Tom van Nuenen, Natalia Criado, Jose M. Such
IEEE Trans. Knowl. Data Eng.1
2022 Measuring Alexa Skill Privacy Practices across Three Years
abstract
Smart Voice Assistants are transforming the way users interact with technology. This transformation is mostly fostered by the proliferation of voice-driven applications (called skills) offered by third-party developers through an online market. We see how the number of skills has rocked in recent years, with the Amazon Alexa skill ecosystem growing from just 135 skills in early 2016 to about 125k skills in early 2021. Along with the growth in skills, there is increasing concern over the risks that third-party skills pose to users’ privacy. In this paper, we perform a systematic and longitudinal measurement study of the Alexa marketplace. We shed light on how this ecosystem evolves using data collected across three years between 2019 and 2021. We demystify developers’ data disclosure practices and present an overview of the third-party ecosystem. We see how the research community continuously contribute to the market’s sanitation, but the Amazon vetting process still requires significant improvement. We perform a responsible disclosure process reporting 675 skills with privacy issues to both Amazon and all affected developers, out of which 246 skills suffer from important issues (i.e., broken traceability). We see that 107 out of the 246 (43.5%) skills continue to display broken traceability almost one year after being reported. As a result, the overall state of affairs has improved in the ecosystem over the years. Yet, newly submitted skills and unresolved known issues pose an endemic risk.
Jide S. Edu, Xavier Ferrer Aran, Jose M. Such, Guillermo Suarez-Tangil
WWW2
2021 Discovering and Categorising Language Biases in Reddit
Xavier Ferrer Aran, Tom van Nuenen, Jose M. Such, Natalia Criado
ICWSM1
2016 Concept Discovery and Argument Bundles in the Experience Web
Xavier Ferrer Aran, Enric Plaza
ICCBR1
2015 Aspect Selection for Social Recommender Systems
Yoke Yie Chen, Xavier Ferrer Aran, Nirmalie Wiratunga, Enric Plaza
ICCBR2
2014 Sentiment and Preference Guided Social Recommendation
Yoke Yie Chen, Xavier Ferrer Aran, Nirmalie Wiratunga, Enric Plaza
ICCBR2