Haris Bin Zia

dblp:213/7568 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
7since 2021 · last 2024
0000-0002-9826-9731ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Mastodoner: A Command-line Tool and Python Library for Public Data Collection from Mastodon
abstract
This paper introduces Mastodoner, a command-line tool and Python library aimed at simplifying access to public data on Mastodon, a prominent player in the Fediverse --- a decentralized network of interconnected social media platforms. Mastodoner addresses the challenges posed by Mastodon's decentralized nature by providing a unified interface for data collection, instance discovery, and secure data sharing. Through examples and demonstrations, this paper illustrates Mastodoner's capabilities in facilitating researchers' access to and analysis of public Mastodon data, thus advancing research in decentralized social media analytics. The tool and documentation are available at: https://github.com/harisbinzia/mastodoner.
Haris Bin Zia, Ignacio Castro, Gareth Tyson
CIKM1
2024 Collecting and Analyzing Public Data from Mastodon
abstract
Understanding online behaviors, communities, and trends through social media analytics is becoming increasingly important. Recent changes in the accessibility of platforms like Twitter have made Mastodon a valuable alternative for researchers. In this tutorial, we will explore methods for collecting and analyzing public data from Mastodon, a decentralized micro-blogging social network. Participants will learn about the architecture of Mastodon, techniques and best practices for data collection, and various analytical methods to derive insights from the collected data. This session aims to equip researchers with the skills necessary to harness the potential of Mastodon data in computational social science and social data science research.
Haris Bin Zia, Ignacio Castro, Gareth Tyson
CIKM1
2024 Fediverse Migrations: A Study of User Account Portability on the Mastodon Social Network
abstract
The advent of regulation, such as the Digital Markets Act, will foster greater interoperability across competing digital platforms. In such regulatory environments, decentralized platforms like Mastodon have pioneered the principles of social data portability. Such platforms are composed of thousands of independent servers, each of which hosts their own social community. To enable transparent interoperability, users can easily migrate their accounts from one server provider to another. In this paper, we examine 8,745 users who switch their server instances in Mastodon. We use this as a case study to examine account portability behavior more broadly. We explore the factors that affect users' decision to switch instances, as well as the impact of switching on their social media engagement and discussion topics. This leads us to build a classifier to show that switching is predictable, with an F1 score of 0.891. We argue that Mastodon serves as an early exemplar of a social media platform that advocates account interoperability and portability. We hope that this study can bring unique insights to a wider and open digital world in the future.
Haris Bin Zia, Jiahui He 0001, Ignacio Castro, Gareth Tyson
IMC1
2023 Flocking to Mastodon: Tracking the Great Twitter Migration
abstract
The acquisition of Twitter by Elon Musk has spurred controversy and uncertainty among Twitter users. The move raised both praise and concerns, particularly regarding Musk's views on free speech. As a result, a large number of Twitter users have looked for alternatives to Twitter. Mastodon, a decentralized micro-blogging social network, has attracted the attention of many users and the general media. In this paper, we analyze the migration of 136,009 users from Twitter to Mastodon. We inspect the impact that this has on the wider Mastodon ecosystem, particularly in terms of user-driven pressure towards centralization. We further explore factors that influence users to migrate, highlighting the effect of users' social networks. Finally, we inspect the behavior of individual users, showing how they utilize both Twitter and Mastodon in parallel. We find a clear difference in the topics discussed on the two platforms. This leads us to build classifiers to explore if migration is predictable. Through feature analysis, we find that the content of tweets as well as the number of URLs, the number of likes, and the length of tweets are effective metrics for the prediction of user migration.
Jiahui He 0001, Haris Bin Zia, Ignacio Castro, Aravindh Raman, Nishanth Sastry, Gareth Tyson
IMC2
2023 Will Admins Cope? Decentralized Moderation in the Fediverse
abstract
As an alternative to Twitter and other centralized social networks, the Fediverse is growing in popularity. The recent, and polemical, takeover of Twitter by Elon Musk has exacerbated this trend. The Fediverse includes a growing number of decentralized social networks, such as Pleroma or Mastodon, that share the same subscription protocol (ActivityPub). Each of these decentralized social networks is composed of independent instances that are run by different administrators. Users, however, can interact with other users across the Fediverse regardless of the instance they are signed up to. The growing user base of the Fediverse creates key challenges for the administrators, who may experience a growing burden. In this paper, we explore how large that overhead is, and whether there are solutions to alleviate the burden. We study the overhead of moderation on the administrators. We observe a diversity of administrator strategies, with evidence that administrators on larger instances struggle to find sufficient resources. We then propose a tool, WatchGen, to semi-automate the process.
Anaobi Ishaku Hassan, Aravindh Raman, Ignacio Castro, Haris Bin Zia, Damilola Ibosiola, Gareth Tyson
WWW4
2022 Improving Zero-Shot Cross-Lingual Hate Speech Detection with Pseudo-Label Fine-Tuning of Transformer Language Models
Haris Bin Zia, Ignacio Castro, Arkaitz Zubiaga, Gareth Tyson
ICWSM1
2021 Exploring content moderation in the decentralised web: the pleroma case
abstract
Decentralising the Web is a desirable but challenging goal. One particular challenge is achieving decentralised content moderation in the face of various adversaries (e.g. trolls). To overcome this challenge, many Decentralised Web (DW) implementations rely on federation policies. Administrators use these policies to create rules that ban or modify content that matches specific rules. This, however, can have unintended consequences for many users. In this paper, we present the first study of federation policies on the DW, their in-the-wild usage, and their impact on users. We identify how these policies may negatively impact "innocent" users and outline possible solutions to avoid this problem in the future.
Anaobi Ishaku Hassan, Aravindh Raman, Ignacio Castro, Haris Bin Zia, Emiliano De Cristofaro, Nishanth Sastry, Gareth Tyson
CoNEXT4
2020 SimplifyUR: Unsupervised Lexical Text Simplification for Urdu
abstract
This paper presents the first attempt at Automatic Text Simplification (ATS) for Urdu, the language of 170 million people worldwide. Being a low-resource language in terms of standard linguistic resources, recent text simplification approaches that rely on manually crafted simplified corpora or lexicons such as WordNet are not applicable to Urdu. Urdu is a morphologically rich language that requires unique considerations such as proper handling of inflectional case and honorifics. We present an unsupervised method for lexical simplification of complex Urdu text. Our method only requires plain Urdu text and makes use of word embeddings together with a set of morphological features to generate simplifications. Our system achieves a BLEU score of 80.15 and SARI score of 42.02 upon automatic evaluation on manually crafted simplified corpora. We also report results for human evaluations for correctness, grammaticality, meaning-preservation and simplicity of the output. Our code and corpus are publicly available to make our results reproducible.
Namoos Hayat Qasmi, Haris Bin Zia, Awais Athar, Agha Ali Raza
LREC2
2018 Urdu Word Segmentation using Conditional Random Fields (CRFs)
abstract
State-of-the-art Natural Language Processing algorithms rely heavily on efficient word segmentation. Urdu is amongst languages for which word segmentation is a complex task as it exhibits space omission as well as space insertion issues. This is partly due to the Arabic script which although cursive in nature, consists of characters that have inherent joining and non-joining attributes regardless of word boundary. This paper presents a word segmentation system for Urdu which uses a Conditional Random Field sequence modeler with orthographic, linguistic and morphological features. Our proposed model automatically learns to predict white space as word boundary as well as Zero Width Non-Joiner (ZWNJ) as sub-word boundary. Using a manually annotated corpus, our model achieves F1 score of 0.97 for word boundary identification and 0.85 for sub-word boundary identification tasks. We have made our code and corpus publicly available to make our results reproducible.
Haris Bin Zia, Agha Ali Raza, Awais Athar
COLING1
2018 Rapid Collection of Spontaneous Speech Corpora Using Telephonic Community Forums
Agha Ali Raza, Awais Athar, Shan Randhawa, Zain Tariq, Bilal Saleem, Haris Bin Zia, Umar Saif, Ronald Rosenfeld
INTERSPEECH6
2018 PronouncUR: An Urdu Pronunciation Lexicon Generator
Haris Bin Zia, Agha Ali Raza, Awais Athar
LREC1