EDBT 2026 Demo / reviewers in the wild / expert
Gareth Tyson
dblp:85/85
· DBLP profile ↗
54ranked-venue papers in the field
3as first author
33since 2021 · last 2026
0000-0003-3010-791XORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 40 (2 first)Data Mining & Knowledge Discovery · 12 (1 first)Other / Interdisciplinary · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Content Moderation with LLMs: A Reddit Case Study on Evaluating and Refining Human DecisionsabstractLarge Language Models (LLMs) offer significant potential for assisting with the design and implementation of social platform moderation. This study evaluates their efficacy as both a replacement to and an augmentation for human moderators. Using Reddit as a case study, we first demonstrate that LLMs can effectively replicate human moderation decisions, achieving 83.9% agreement. Through a mix of LLMs and human annotations, we then evaluate real moderator decisions, uncovering substantial error rates: 15.2% of removals and 13% of approvals are estimated as incorrect, primarily stemming from moderators citing the incorrect rules (84.3% of errors). This motivates us to propose RuleSharpener, a tool that uses LLMs to diagnose the root causes of moderation errors (e.g. ambiguous rules) and generate clearer, more actionable guidelines. Our evaluation shows that RuleSharpener increases the accuracy of identifying the specific rules violated by violation posts by 38.0%. Our work demonstrates how LLMs can augment human moderation, refine community policies, and reduce operational burdens, offering a better solution for platform governance on the web. Jiahui He 0001, Yiluo Wei, Gareth Tyson |
WWW | 3 |
| 2026 | Open or Blocked Skies? Community Moderation Practices in BlueskyabstractContent moderation is a major challenge for online platforms. While user-driven blocking is a common tool, its dynamics are usually hidden as moderation data is private. Bluesky makes moderation actions public-by-design, providing an unprecedented opportunity to study a community-driven moderation ecosystem at scale. We leverage this transparency to (1) map the ecosystem of moderation blocking actions across 34 million users, including both individual blocks and the through blocklists, (2) identify the signals that correlate with blocking, and (3) measure the consequences of these actions. We demonstrate that community blocking is widespread, with a volume several orders of magnitude higher than official takedowns, and affects the visibility of more than 90 % of Bluesky content. The blocked accounts represent the most active, popular, toxic, and politically inclined users. However, different blocklists target different types of accounts and behaviors. Finally, blocking does not decrease the popularity and activity of the blocked users and has a limited effect on the social graph. By quantifying its dynamics and trade-offs, our study provides empirical grounding for designing future moderation systems that are transparent, pluralistic, and resistant to centralized control. Taken together, this study provides the first large-scale, quantitative analysis of a community-driven moderation ecosystem, demonstrating how individual and collective interventions influence user behavior. Saidu Sokoto, Leonhard Balduf, Onur Ascigil, Gareth Tyson, Ignacio Castro, Björn Scheuermann 0001, Andrea Baronchelli, Michal Król |
WWW | 4 |
| 2026 | SkyCL: Swift Continuous Learning with Kinship-Awareness for Multi-Drone Video Analytics under Drastic Drift
Yuanzheng Tan, Qing Li 0006, Junkun Peng, Gareth Tyson, Zhenhui Yuan, Tingting Yang 0001, Yong Jiang 0001 |
WWW | 5 |
| 2026 | Understanding the Consequences of VTuber Reincarnation
Yiluo Wei, Gareth Tyson |
WWW | 2 |
| 2025 | Bootstrapping Social Networks: Lessons from Bluesky Starter PacksabstractMicroblogging is a crucial mode of online communication. However, launching a new microblogging platform remains challenging, largely due to network effects. This has resulted in entrenched (and undesirable) dominance by established players, such as X/Twitter. To overcome these network effects, Bluesky, an emerging microblogging platform, introduced starter packs — curated lists of accounts that users can follow with a single click. We ask if starter packs have the potential to tackle the critical problem of social bootstrapping in new online social networks. We assess whether starter packs have indeed been helpful in supporting Bluesky growth. Our dataset includes 25.05 × 10⁶ users and 335.42 × 10³ starter packs with 1.73 × 10⁶ members, covering the entire lifecycle of Bluesky. We study the usage of these starter packs, their ability to drive network and activity growth, and their potential downsides. We also quantify the benefits of starter packs for members and creators on user visibility and activity while identifying potential challenges. By evaluating starter packs’ effectiveness and limitations, we contribute to the broader discourse on platform growth strategies and competitive innovation in the social media landscape. Leonhard Balduf, Saidu Sokoto, Andrea Baronchelli, Ignacio Castro, Michal Król, Gareth Tyson, George Pavlou, Björn Scheuermann 0001, Onur Ascigil |
ICWSM | 6 |
| 2025 | Examining the Makeup of Media Trigger Warnings OnlineabstractIn today’s digital landscape, the prevalence of sensitive online content has made trigger warnings essential. These warnings inform viewers that the content they are about to see contains sensitive artifacts (e.g. violence). This paper studies the use of trigger warnings, exploiting data from two major platforms: Does the Dog Die, a crowdsourcing trigger warnings platform, and IMDb, a media database. We first study how different media types (e.g. films, video games, and TV shows) are labeled with varying trigger warnings and the different co-occurrence patterns among different trigger warnings. We also discover controversy surrounding certain trigger warnings, with inconsistent opinions stated by different people. We further show that different jurisdictions (e.g. USA vs. UK) assign different content ratings (e.g. R-18) for the same media, even when the same trigger warnings are present. Finally, we develop automatic detectors to identify trigger warnings from IMDb text. We achieve F1 scores exceeding 0.7 for all 10 selected trigger warnings. Peixian Zhang, Yupeng He, Ehsan ul Haq, Gareth Tyson |
ICWSM | 4 |
| 2025 | Virtual Stars, Real Fans: Understanding the VTuber EcosystemabstractLivestreaming by VTubers --- animated 2D/3D avatars controlled by real individuals --- have recently garnered substantial global followings and achieved significant monetary success. Despite prior research highlighting the importance of realism in audience engagement, VTubers deliberately conceal their identities, cultivating dedicated fan communities through virtual personas. While previous studies underscore that building a core fan community is essential to a streamer's success, we lack an understanding of the characteristics of viewers of this new type of streamer. Gaining a deeper insight into these viewers is critical for VTubers to enhance audience engagement, foster a more robust fan base, and attract a larger viewership. To address this gap, we conduct a comprehensive analysis of VTuber viewers on Bilibili, a leading livestreaming platform where nearly all VTubers in China stream. By compiling a first-of-its-kind dataset covering 2.7M livestreaming sessions, we investigate the characteristics, engagement patterns, and influence of VTuber viewers. Our research yields several valuable insights, which we then leverage to develop a tool to ''recommend'' future subscribers to VTubers. By reversing the typical approach of recommending streams to viewers, this tool assists VTubers in pinpointing potential future fans to pay more attention to, and thereby effectively growing their fan community. Yiluo Wei, Gareth Tyson |
WWW | 2 |
| 2024 | The Emergence of Threads: The Birth of a New Social Network
Peixian Zhang, Yupeng He, Ehsan ul Haq, Jiahui He 0001, Gareth Tyson |
ASONAM (3) | 5 |
| 2024 | Exploring the Capability of ChatGPT to Reproduce Human Labels for Social Computing Tasks
Peixian Zhang, Ehsan ul Haq, Pan Hui 0001, Gareth Tyson |
ASONAM (3) | 5 |
| 2024 | Mastodoner: A Command-line Tool and Python Library for Public Data Collection from MastodonabstractThis paper introduces Mastodoner, a command-line tool and Python library aimed at simplifying access to public data on Mastodon, a prominent player in the Fediverse --- a decentralized network of interconnected social media platforms. Mastodoner addresses the challenges posed by Mastodon's decentralized nature by providing a unified interface for data collection, instance discovery, and secure data sharing. Through examples and demonstrations, this paper illustrates Mastodoner's capabilities in facilitating researchers' access to and analysis of public Mastodon data, thus advancing research in decentralized social media analytics. The tool and documentation are available at: https://github.com/harisbinzia/mastodoner. Haris Bin Zia, Ignacio Castro, Gareth Tyson |
CIKM | 3 |
| 2024 | Collecting and Analyzing Public Data from MastodonabstractUnderstanding online behaviors, communities, and trends through social media analytics is becoming increasingly important. Recent changes in the accessibility of platforms like Twitter have made Mastodon a valuable alternative for researchers. In this tutorial, we will explore methods for collecting and analyzing public data from Mastodon, a decentralized micro-blogging social network. Participants will learn about the architecture of Mastodon, techniques and best practices for data collection, and various analytical methods to derive insights from the collected data. This session aims to equip researchers with the skills necessary to harness the potential of Mastodon data in computational social science and social data science research. Haris Bin Zia, Ignacio Castro, Gareth Tyson |
CIKM | 3 |
| 2024 | Decentralised Moderation for Interoperable Social Networks: A Conversation-Based Approach for Pleroma and the FediverseabstractThe recent development of decentralised and interoperable social networks (such as the "fediverse") creates new challenges for content moderators. This is because millions of posts generated on one server can easily "spread" to another, even if the recipient server has very different moderation policies. An obvious solution would be to leverage moderation tools to automatically tag (and filter) posts that contravene moderation policies, e.g. related to toxic speech. Recent work has exploited the conversational context of a post to improve this automatic tagging, e.g. using the replies to a post to help classify if it contains toxic speech. This has shown particular potential in environments with large training sets that contain complete conversations. This, however, creates challenges in a decentralised context, as a single conversation may be fragmented across multiple servers. Thus, each server only has a partial view of an entire conversation because conversations are often federated across servers in a non-synchronized fashion. To address this, we propose a decentralised conversation-aware content moderation approach suitable for the fediverse. Our approach employs a graph deep learning model (GraphNLI) trained locally on each server. The model exploits local data to train a model that combines post and conversational information captured through random walks to detect toxicity. We evaluate our approach with data from Pleroma, a major decentralised and interoperable micro-blogging network containing 2 million conversations. Our model effectively detects toxicity on larger instances, exclusively trained using their local post information (0.8837 macro-F1). Yet, we show that this approach does not perform well on smaller instances that do not possess sufficient local training data. Thus, in cases where a server contains insufficient data, we strategically retrieve information (posts or model parameters) from other servers to reconstruct larger conversations and improve results. With this, we show that we can attain a macro-F1 of 0.8826. Our approach has considerable scope to improve moderation in decentralised and interoperable social networks such as Pleroma or Mastodon. Vibhor Agarwal, Aravindh Raman, Nishanth Sastry, Ahmed M. Abdelmoniem, Gareth Tyson, Ignacio Castro |
ICWSM | 5 |
| 2024 | Temporal Network Analysis of Email Communication Patterns in a Long Standing HierarchyabstractAn important concept in organisational behaviour is how hierarchy affects the voice of individuals, whereby members of a given organisation exhibit differing power relations based on their hierarchical position. Although there have been prior studies of the relationship between hierarchy and voice, they tend to focus on more qualitative small-scale methods and do not account for structural aspects of the organisation. This paper develops large-scale computational techniques utilising temporal network analysis to measure the effect that organisational hierarchy has on communication patterns throughout an organisation, focusing on the structure of pairwise interactions between individuals. To this end, we focus on one major organisation as a case study --- the Internet Engineering Task Force (IETF) --- a major technical standards development organisation for the Internet. A particularly useful feature of the IETF is a transparent hierarchy, where participants take on explicit roles (e.g., Area Directors, Working Group Chairs), and because its processes are open we have visibility into the communication of people at different hierarchy levels over a long time period. Exploiting this, we utilise a temporal network dataset of 989,911 email interactions among 23,741 participants to study how hierarchy impacts communication patterns. We show that the middle levels of the IETF are growing in terms of their dominance in communications. Higher levels consistently experience a higher proportion of incoming communication than lower levels, with higher levels initiating more communications too. We find that, overall, communication tends to flow "up" the hierarchy more than "down". Finally, we find that communication with higher-levels is associated with future communication more than for lower-levels, which we interpret as "facilitation". We conclude by discussing the implications this has on patterns within the wider IETF and the impact our analysis can have for other organisations. Matthew Russell Barnes, Mladen Karan, Stephen McQuistin, Colin Perkins, Gareth Tyson, Matthew Purver, Ignacio Castro, Richard G. Clegg |
ICWSM | 5 |
| 2024 | Making the Pick: Understanding Professional Editor Comment Curation in Online NewsabstractOnline comments within news articles are a key way people share opinions. Discovering insightful comments can, however, be challenging for readers. A solution to this problem is using comment curation, whereby professional editors select the highest quality comments manually --- referred to as ''editor-picks''. This paper studies the growing use of professional editor-curation for user-generated comments. We focus on the New York Times as a case study, using a dataset covering 80k articles. We study the characteristics of editor-pick comments, highlighting how editor criteria vary across news sections (e.g. sports, entertainment). We find that editor-pick comments tend to be longer, more relevant to the article, positive in sentiment, and contain low toxicity. Our analysis further reveals that editors within different news sections exhibit differing criteria when they perform comment selection. Thus, we finally propose a set of models that can automatically identify good candidate editor-picks. Our ultimate goal is to reduce editor and journalistic workload, increasing productivity and the quality of curated comments. Yupeng He, Yimeng Gu, Ravi Shekhar, Ignacio Castro, Gareth Tyson |
ICWSM | 5 |
| 2024 | A Study of Partisan News Sharing in the Russian Invasion of UkraineabstractSince the Russian invasion of Ukraine, a large volume of biased and partisan news has been spread via social media platforms. As this may lead to wider societal issues, we argue that understanding how partisan news sharing impacts users' communication is crucial for better governance of online communities. In this paper, we perform a measurement study of partisan news sharing. We aim to characterize the role of such sharing in influencing users' communications. Our analysis covers an eight-month dataset across six Reddit communities related to the Russian invasion. We first perform an analysis of the temporal evolution of partisan news sharing. We confirm that the invasion stimulates discussion in the observed communities, accompanied by an increased volume of partisan news sharing. Next, we characterize users' response to such sharing. We observe that partisan bias plays a role in narrowing its propagation. More biased media is less likely to be spread across multiple subreddits. However, we find that partisan news sharing attracts more users to engage in the discussion, by generating more comments. We then built a predictive model to identify users likely to spread partisan news. The prediction is challenging though, with 61.57% accuracy on average. Our centrality analysis on the commenting network further indicates that the users who disseminate partisan news possess lower network influence in comparison to those who propagate neutral news. Ehsan ul Haq, Gareth Tyson, Lik-Hang Lee, Yuyang Wang 0002, Pan Hui 0001 |
ICWSM | 3 |
| 2024 | Understanding and Improving Content Moderation in Web3 PlatformsabstractThere have been numerous recent attempts to “decentralize” social media platforms, loosely referred to as Web3. Such ideas, often underpinned by blockchain solutions, offer decentralized equivalents of well-known services (e.g., forums, social networks, video sharing sites, microblogs). One particularly challenging function to implement in such a design is content moderation, due to the lack of central control. Consequently, they often rely on user-controlled moderation, whereby each user must create their own personal block list to filter out content they do not wish to see. This paper presents a first study of user-controlled moderation on one exemplar Web3 social microblogging platform called memo.cash. Based on a dataset covering 391K posts, we study the factors that lead users to “mute” each other. We find that the most crucial factor is the platform action count, rather than the presence of things like hate speech. We also show that the followership network plays a pivotal role in determining their visibility on the platform, further influencing their muting behavior. This leads us to design tooling to automate the muting process on a per-user basis. We model this as a recommendation problem, and experiment with a number of state-of-the-art recommender engines. We show that our system can generate effective personalized mute lists for users. Wenrui Zuo, Raul J. Mondragón, Aravindh Raman, Gareth Tyson |
ICWSM | 4 |
| 2024 | Global Prosperity or Local Monopoly? Understanding the Geography of App PopularityabstractApp stores allow developers to globally distribute their apps to gain more users and attention. In the highly competitive market of app stores, developers need to cater to a large number of users spanning multiple countries. We posit that the characteristics of diverse geographical, linguistic, cultural, societal, and economic environments may impact the adoption of apps. In this paper, we take the first step to characterize popular apps across over 150 countries worldwide, and explore the potential correlations to a number of underlying factors including geography, language as well as cultural, societal, and economic dimensions. Our study is based on a longitudinal (one-year) dataset of daily app popularity from the iOS app stores, covering 154 regions around the world. We reveal that app popularity shows great diversity across the world, while similarities exist among countries that share geographical proximity and linguistic convergence. The differences in app popularity across regions can be further correlated with the cultural model and socioeconomic indices we adopt. On top of the dataset and findings, we implement a prediction task that contributes to app distribution, helping developers choose the right market to distribute and promote their apps. To the best of our knowledge, we are the first to attempt to provide a global understanding of the characteristics of app popularity across the mobile app ecosystem. Our observations can benefit stakeholders in the ecosystem, striving to improve app uptake. Liu Wang 0002, Conghui Zheng, Haoyu Wang 0001, Xiapu Luo, Gareth Tyson, Yi Wang 0004, Shangguang Wang |
MSR | 5 |
| 2024 | Unveiling the Paradox of NFT ProsperityabstractUnlike fungible tokens (e.g., cryptocurrency), a Non-Fungible Token (NFT) is unique and indivisible. As such, they can be used to authenticate ownership of digital assets (e.g., a photo) in a decentralized fashion. Given that NFTs have generated significant media attention since 2021, we perform a large-scale measurement study of the NFT ecosystem. We collect over 242M transfer logs and over 97M marketplace transactions until Aug 1st, 2023, by far the largest NFT dataset, to the best of our knowledge. We characterize the on-chain behavior of NFTs and their trading across five major marketplaces. We find that, although the NFT ecosystem is growing rapidly, it is driven by a relatively small set of dominant centralized players, with suspicious trade activities, e.g., over 23% of the monetary volume is generated by malicious wash trading and the ecosystem has experienced over 157K cases of NFT arbitrage, with a total sum of over \25M profit. Our observations motivate the need for more research efforts in the NFT security analysis. Pengcheng Xia 0001, Gareth Tyson, Xiapu Luo, Lei Wu 0012, Yajin Zhou, Wei Cai 0002, Haoyu Wang 0001 |
WWW | 5 |
| 2024 | APT-Pipe: A Prompt-Tuning Tool for Social Data Annotation using ChatGPTabstractRecent research has highlighted the potential of LLMs, like ChatGPT, for performing label annotation on social computing data. However, it is already well known that performance hinges on the quality of the input prompts. To address this, there has been a flurry of research into prompt tuning --- techniques and guidelines that attempt to improve the quality of prompts. Yet these largely rely on manual effort and prior knowledge of the dataset being annotated. To address this limitation, we propose APT-Pipe, an automated prompt-tuning pipeline. APT-Pipe aims to automatically tune prompts to enhance ChatGPT's text classification performance on any given dataset. We implement APT-Pipe and test it across twelve distinct text classification datasets. We find that prompts tuned by APT-Pipe help ChatGPT achieve higher weighted F1-score on nine out of twelve experimented datasets, with an improvement of 7.01% on average. We further highlight APT-Pipe's flexibility as a framework by showing how it can be extended to support additional tuning mechanisms. Zhizhuo Yin, Gareth Tyson, Ehsan ul Haq, Lik-Hang Lee, Pan Hui 0001 |
WWW | 3 |
| 2023 | Understanding Characteristics of Catalyst Users in the WallStreetBets CommunityabstractWallStreetBets (WSB), a Reddit community, impacted stock markets during the 2021 GameStop Short Squeeze. We examine the content and user properties that influence engagement in WSB. Despite WSB's association with emojis and informal terms, engagement among community members depends on more than surface-level factors. Although emojis are commonly used, they are not as effective at fostering interactions among users. Community members engage more with posts that have longer and topic-specific text. Simply producing a high volume of posts is not enough to attract an audience. Consistent topical focus, reciprocal interactions, and previous authorship of catalyst posts influence engagement. WSB posts, regardless of length, generally remain relevant to the community's theme of stock trading. Our findings provide insights into WSB engagement patterns and can be useful for downstream research, such as financial predictive tasks using WSB data. Ehsan ul Haq, Haodi Weng, Gareth Tyson, Lik-Hang Lee, Reza Hadi Mogavi, Tristan Braud, Pan Hui 0001 |
ASONAM | 5 |
| 2023 | Echo Chambers within the Russo-Ukrainian War: The Role of Bipartisan UsersabstractThe ongoing Russia-Ukraine war has been extensively discussed on social media. One commonly observed problem in such discourse is the emergence of echo chambers, where users are rarely exposed to opinions outside their own worldview. Prior literature on this topic has assumed that such users hold a single consistent view. However, recent work has revealed that complex topics often trigger bipartisanship among certain people. With this in mind, we study the presence of echo chambers on Twitter related to the Russo-Ukrainian war. We measure their presence and identify an important subset of bipartisan users who vary their opinion during the invasion. We then explore the role they play in the communications graph and their impact on echo chambers. Peixian Zhang, Ehsan ul Haq, Pan Hui 0001, Gareth Tyson |
ASONAM | 5 |
| 2023 | Lady and the Tramp Nextdoor: Online Manifestations of Real-World Inequalities in the Nextdoor Social NetworkabstractFrom health to education, income impacts a huge range of life choices. Earlier research has leveraged data from online social networks to study precisely this impact. In this paper, we ask the opposite question: do different levels of income result in different online behaviors? We demonstrate it does. We present the first large-scale study of Nextdoor, a popular location-based social network. We collect 2.6 Million posts from 64,283 neighborhoods in the United States and 3,325 neighborhoods in the United Kingdom, to examine whether online discourse reflects the income and income inequality of a neighborhood. We show that posts from neighborhoods with different incomes indeed differ, e.g. richer neighborhoods have a more positive sentiment and discuss crimes more, even though their actual crime rates are much lower. We then show that user-generated content can predict both income and inequality. We train multiple machine learning models and predict both income (R2=0.841) and inequality (R2=0.77). Waleed Iqbal, Vahid Ghafouri, Gareth Tyson, Guillermo Suarez-Tangil, Ignacio Castro |
ICWSM | 3 |
| 2023 | Will Admins Cope? Decentralized Moderation in the FediverseabstractAs an alternative to Twitter and other centralized social networks, the Fediverse is growing in popularity. The recent, and polemical, takeover of Twitter by Elon Musk has exacerbated this trend. The Fediverse includes a growing number of decentralized social networks, such as Pleroma or Mastodon, that share the same subscription protocol (ActivityPub). Each of these decentralized social networks is composed of independent instances that are run by different administrators. Users, however, can interact with other users across the Fediverse regardless of the instance they are signed up to. The growing user base of the Fediverse creates key challenges for the administrators, who may experience a growing burden. In this paper, we explore how large that overhead is, and whether there are solutions to alleviate the burden. We study the overhead of moderation on the administrators. We observe a diversity of administrator strategies, with evidence that administrators on larger instances struggle to find sufficient resources. We then propose a tool, WatchGen, to semi-automate the process. Anaobi Ishaku Hassan, Aravindh Raman, Ignacio Castro, Haris Bin Zia, Damilola Ibosiola, Gareth Tyson |
WWW | 6 |
| 2023 | Not Seen, Not Heard in the Digital World! Measuring Privacy Practices in Children's AppsabstractThe digital age has brought a world of opportunity to children. Connectivity can be a game-changer for some of the world’s most marginalized children. However, while legislatures around the world have enacted regulations to protect children’s online privacy, and app stores have instituted various protections, privacy in mobile apps remains a growing concern for parents and wider society. In this paper, we explore the potential privacy issues and threats that exist in these apps. We investigate 20195 mobile apps from the Google Play store that are designed particularly for children (Family apps) or include children in their target user groups (Normal apps). Using both static and dynamic analysis, we find that 4.47% of Family apps request location permissions, even though collecting location information from children is forbidden by the Play store, and 81.25% of Family apps use trackers (which are not allowed in children’s apps). Even major developers with 40+ kids apps on the Play store use ad trackers. Furthermore, we find that most permission request notifications are not well designed for children, and 19.25% apps have inconsistent content age ratings across the different protection authorities. Our findings suggest that, despite significant attention to children’s privacy, a large gap between regulatory provisions, app store policies, and actual development practices exist. Our research sheds light for government policymakers, app stores, and developers. Ruoxi Sun 0001, Minhui Xue 0001, Gareth Tyson, Shuo Wang 0012, Seyit Ahmet Çamtepe, Surya Nepal |
WWW | 3 |
| 2023 | Cashing in on Contacts: Characterizing the OnlyFans EcosystemabstractAdult video-sharing has undergone dramatic shifts. New platforms that directly interconnect (often amateur) producers and consumers now allow content creators to promote material across the web and directly monetize the content they produce. OnlyFans is the most prominent example of this new trend. OnlyFans is a content subscription service where creators earn money from users who subscribe to their material. In contrast to prior adult platforms, OnlyFans emphasizes creator-consumer interaction for audience accumulation and maintenance. This results in a wide cross-platform ecosystem geared towards bringing consumers to creators’ accounts. In this paper, we inspect this emerging ecosystem, focusing on content creators and the third-party platforms they connect to. Pelayo Vallina, Ignacio Castro, Gareth Tyson |
WWW | 3 |
| 2023 | BiSR: Bidirectionally Optimized Super-Resolution for Mobile Video StreamingabstractThe user experience of mobile web video streaming is often impacted by insufficient and dynamic network bandwidth. In this paper, we design Bidirectionally Optimized Super-Resolution (BiSR) to improve the quality of experience (QoE) for mobile web users under limited bandwidth. BiSR exploits a deep neural network (DNN)-based model to super-resolve key frames efficiently without changing the inter-frame spatial-temporal information. We then propose a downscaling DNN and a mobile-specific optimized lightweight super-resolution DNN to enhance the performance. Finally, a novel reinforcement learning-based adaptive bitrate (ABR) algorithm is proposed to verify the performance of BiSR on real network traces. Our evaluation, using a full system implementation, shows that BiSR saves 26% of bitrate compared to the traditional H.264 codec and improves the SSIM of video by 3.7% compared to the prior state-of-the-art. Overall, BiSR enhances the user-perceived quality of experience by up to 30.6%. Qian Yu 0011, Qing Li 0006, Gareth Tyson, Wanxin Shi, Jianhui Lv, Zhenhui Yuan, Peng Zhang 0104, Yulong Lan |
WWW | 4 |
| 2023 | Set in Stone: Analysis of an Immutable Web3 Social Media PlatformabstractThere has been growing interest in the so-called “Web3” movement. This loosely refers to a mix of decentralized technologies, often underpinned by blockchain technologies. Among these, Web3 social media platforms have begun to emerge. These store all social interaction data (e.g., posts) on a public ledger, removing the need for centralized data ownership and management. But this comes at a cost, which some argue is prohibitively expensive. As an exemplar within this growing ecosytem, we explore memo.cash, a microblogging service built on the Bitcoin Cash (BCH) blockchain. We gather data for 24K users, 317K posts, 2.57M user actions, which have facilitated $6.75M worth of transactions. A particularly unique feature is that users must pay BCH tokens for each interaction (e.g., posting, following). We study how this may impact the social makeup of the platform. We therefore study memo.cash as both a social network and a transaction platform. Wenrui Zuo, Aravindh Raman, Raul J. Mondragón, Gareth Tyson |
WWW | 4 |
| 2022 | Exploring Mental Health Communications among Instagram CoachesabstractThere has been a significant expansion in the use of online social networks (OSNs) to support people experiencing mental health issues. This paper studies the role of Instagram influencers who specialize in coaching people with mental health issues. Using a dataset of 97k posts, we characterize such users' linguistic and behavioural features. We explore how these observations impact audience engagement (as measured by likes). We show that the support provided by these accounts varies based on their self-declared professional identities. For instance, Instagram accounts that declare themselves as Authors offer less support than accounts that label themselves as a Coach. We show that increasing information support in general communication positively affects user engagement. However, the effect of vocabulary on engagement is not consistent across the Instagram account types. Our findings shed light on this understudied topic and guide how mental health practitioners can improve outreach. Ehsan ul Haq, Lik-Hang Lee, Gareth Tyson, Reza Hadi Mogavi, Tristan Braud, Pan Hui 0001 |
ASONAM | 3 |
| 2022 | The Web We Weave: Untangling the Social Graph of the IETF
Prashant Khare, Mladen Karan, Stephen McQuistin, Colin Perkins, Gareth Tyson, Matthew Purver, Patrick G. T. Healey, Ignacio Castro |
ICWSM | 5 |
| 2022 | Improving Zero-Shot Cross-Lingual Hate Speech Detection with Pseudo-Label Fine-Tuning of Transformer Language Models
Haris Bin Zia, Ignacio Castro, Arkaitz Zubiaga, Gareth Tyson |
ICWSM | 4 |
| 2022 | Human-Avatar Interaction in Metaverse: Framework for Full-Body InteractionabstractThe metaverse is a network of shared virtual environments where people can interact synchronously through their avatars. To enable this, it is necessary to accurately capture and recreate (physical) human motion. This is used to render avatars correctly, reflecting the motion of their corresponding users. In large-scale environments this must be done in real-time. This paper proposes a human-avatar framework with full-body motion capture. Its goal is to deliver high-accuracy capture with low computational and network overheads. It relies on a lightweight Octree data structure to record and transmit motion to other users. We conduct a user study with 22 participants and perform a preliminary evaluation of its scalability. Our user study shows that Octree with Inverse Kinematic achieves the best trade-off, achieving low delay and high accuracy. Our proposed solution delivers the lowest delay, with an average of 67ms in an environment of 8 concurrent users. It attains a 55.7% improvement over the prior techniques. Kit-Yung Lam, Ahmad Yousef Alhilal, Lik-Hang Lee, Gareth Tyson, Pan Hui 0001 |
MMAsia | 5 |
| 2022 | Jettisoning Junk Messaging in the Era of End-to-End Encryption: A Case Study of WhatsAppabstractWhatsApp is a popular messaging app used by over a billion users around the globe. Due to this popularity, understanding misbehavior on WhatsApp is an important issue. The sending of unwanted junk messages by unknown contacts via WhatsApp remains understudied by researchers, in part because of the end-to-end encryption offered by the platform. We address this gap by studying junk messaging on a multilingual dataset of 2.6M messages sent to 5K public WhatsApp groups in India. We characterise both junk content and senders. We find that nearly 1 in 10 messages is unwanted content sent by junk senders, and a number of unique strategies are employed to reflect challenges faced on WhatsApp, e.g., the need to change phone numbers regularly. We finally experiment with on-device classification to automate the detection of junk, whilst respecting end-to-end encryption. Pushkal Agarwal, Aravindh Raman, Damilola Ibosiola, Nishanth Sastry, Gareth Tyson, Venkata Rama Kiran Garimella |
WWW | 5 |
| 2022 | Modeling and Optimizing the Scaling Performance in Distributed Deep Learning TrainingabstractDistributed Deep Learning (DDL) is widely used to accelerate deep neural network training for various Web applications. In each iteration of DDL training, each worker synchronizes neural network gradients with other workers. This introduces communication overhead and degrades the scaling performance. In this paper, we propose a recursive model, OSF (Scaling Factor considering Overlap), for estimating the scaling performance of DDL training of neural network models, given the settings of the DDL system. OSF captures two main characteristics of DDL training: the overlap between computation and communication, and the tensor fusion for batching updates. Measurements on a real-world DDL system show that OSF obtains a low estimation error (ranging from 0.5% to 8.4% for different models). Using OSF, we identify the factors that degrade the scaling performance, and propose solutions to effectively mitigate their impacts. Specifically, the proposed adaptive tensor fusion improves the scaling performance by 32.2%∼ 150% compared to the constant tensor fusion buffer size. Tianhao Miao, Qinghua Wu 0004, Zhenyu Li 0001, Guangxin He, Jiaoren Wu, Shengzhuo Zhang, Xingwu Yang, Gareth Tyson, Gaogang Xie |
WWW | 9 |
| 2020 | A First Look at COVID-19 Messages on WhatsApp in PakistanabstractThe worldwide spread of COVID-19 has prompted extensive online discussions, creating an ‘infodemic’ on social media platforms such as WhatsApp and Twitter. However, the information shared on these platforms is prone to be unreliable and/or misleading. In this paper, we present the first analysis of COVID-19 discourse on public WhatsApp groups from Pakistan. Building on a large scale annotation of thousands of messages containing text and images, we identify the main categories of discussion. We focus on COVID-19 messages and understand the different types of images/text messages being propagated. By exploring user behavior related to COVID messages, we inspect how misinformation is spread. Finally, by quantifying the flow of information across WhatsApp and Twitter, we show how information spreads across platforms and how WhatsApp acts as a source for much of the information shared on Twitter. Rana Tallal Javed, Mirza Elaaf Shuja, Junaid Qadir 0001, Waleed Iqbal, Gareth Tyson, Ignacio Castro, Venkata Rama Kiran Garimella |
ASONAM | 6 |
| 2020 | Impersonation on Social Media: A Deep Neural Approach to Identify Ingenuine ContentabstractImpersonators are playing an important role in the production and propagation of the content on Online Social Networks, notably on Instagram. These entities are nefarious fake accounts that intend to disguise a legitimate account by making similar profiles and then striking social media by fake content, which makes it considerably harder to understand which posts are genuinely produced. In this study, we focus on three important communities with legitimate verified accounts. Among them, we identify a collection of 2.2K impersonator profiles with nearly 10k generated posts, 68K comments, and 90K likes. Then, based on profile characteristics and user behaviours, we cluster them into two collections of `bot' and `fan'. In order to separate the impersonator-generated post from genuine content, we propose a Deep Neural Network architecture that measures `profiles' and `posts' features to predict the content type: `bot-generated', `fan-generated', or `genuine' content. Our study shed light into this interesting phenomena and provides interesting observation on bot-generated content that can help us to understand the role of impersonators in the production of fake content on Instagram. Koosha Zarei, Reza Farahbakhsh, Noël Crespi, Gareth Tyson |
ASONAM | 4 |
| 2020 | Characterising and Detecting Sponsored Influencer Posts on InstagramabstractRecent years have seen a new form of advertisement campaigns emerge: those involving so-called social media influencers. These influencers accept money in return for promoting products via their social media feeds. We gather a large-scale Instagram dataset covering thousands of accounts advertising products, and create a categorisation based on the number of users they reach. We then provide a detailed analysis of the types of products being advertised by these accounts, their potential reach, and the engagement they receive from their followers. Based on our findings, we train machine learning models to distinguish sponsored content from non-sponsored, and identify cases where people are generating sponsored posts without labelling them. Koosha Zarei, Damilola Ibosiola, Reza Farahbakhsh, Zafar Gilani, Venkata Rama Kiran Garimella, Noël Crespi, Gareth Tyson |
ASONAM | 7 |
| 2020 | Characterising User Content on a Multi-Lingual Social Network
Pushkal Agarwal, Venkata Rama Kiran Garimella, Sagar Joglekar 0001, Nishanth Sastry, Gareth Tyson |
ICWSM | 5 |
| 2020 | Mobile App SquattingabstractDomain squatting, the adversarial tactic where attackers register domain names that mimic popular ones, has been observed for decades. However, there has been growing anecdotal evidence that this style of attack has spread to other domains. In this paper, we explore the presence of squatting attacks in the mobile app ecosystem. In “App Squatting”, attackers release apps with identifiers (e.g., app name or package name) that are confusingly similar to those of popular apps or well-known Internet brands. This paper presents the first in-depth measurement study of app squatting showing its prevalence and implications. We first identify 11 common deformation approaches of app squatters and propose “AppCrazy”, a tool for automatically generating variations of app identifiers. We have applied AppCrazy to the top-500 most popular apps in Google Play, generating 224,322 deformation keywords which we then use to test for app squatters on popular markets. Through this, we confirm the scale of the problem, identifying 10,553 squatting apps (an average of over 20 squatting apps for each legitimate one). Our investigation reveals that more than 51% of the squatting apps are malicious, with some being extremely popular (up to 10 million downloads). Meanwhile, we also find that mobile app markets have not been successful in identifying and eliminating squatting apps. Our findings demonstrate the urgency to identify and prevent app squatting abuses. To this end, we have publicly released all the identified squatting apps, as well as our tool AppCrazy. Yangyu Hu, Haoyu Wang 0001, Li Li 0029, Gareth Tyson, Ignacio Castro, Yao Guo 0001, Lei Wu 0012, Guoai Xu |
WWW | 5 |
| 2019 | Who Watches the Watchmen: Exploring Complaints on the WebabstractUnder increasing scrutiny, many web companies now offer bespoke mechanisms allowing any third party to file complaints (e.g., requesting the de-listing of a URL from a search engine). While this self-regulation might be a valuable web governance tool, it places huge responsibility within the hands of these organisations that demands close examination. We present the first large-scale study of web complaints (over 1 billion URLs). We find a range of complainants, largely focused on copyright enforcement. Whereas the majority of organisations are occasional users of the complaint system, we find a number of bulk senders specialised in targeting specific types of domain. We identify a series of trends and patterns amongst both the domains and complainants. By inspecting the availability of the domains, we also observe that a sizeable portion go offline shortly after complaints are generated. This paper sheds critical light on how complaints are issued, who they pertain to and which domains go offline after complaints are issued. Damilola Ibosiola, Ignacio Castro, Gianluca Stringhini, Steve Uhlig, Gareth Tyson |
WWW | 5 |
| 2019 | The Chain of Implicit Trust: An Analysis of the Web Third-party Resources LoadingabstractThe Web is a tangled mass of interconnected services, where websites import a range of external resources from various third-party domains. The latter can also load resources hosted on other domains. For each website, this creates a dependency chain underpinned by a form of implicit trust between the first-party and transitively connected third-parties. The chain can only be loosely controlled as first-party websites often have little, if any, visibility on where these resources are loaded from. This paper performs a large-scale study of dependency chains in the Web, to find that around 50% of first-party websites render content that they did not directly load. Although the majority (84.91%) of websites have short dependency chains (below 3 levels), we find websites with dependency chains exceeding 30. Using VirusTotal, we show that 1.2% of these third-parties are classified as suspicious - although seemingly small, this limited set of suspicious third-parties have remarkable reach into the wider ecosystem. Muhammad Ikram 0001, Rahat Masood, Gareth Tyson, Mohamed Ali Kâafar, Noha Loizon, Roya Ensafi |
WWW | 3 |
| 2019 | PYTHIA: a Framework for the Automated Analysis of Web Hosting EnvironmentsabstractA common approach when setting up a website is to utilize third party Web hosting and content delivery networks. Without taking this trend into account, any measurement study inspecting the deployment and operation of websites can be heavily skewed. Unfortunately, the research community lacks generalizable tools that can be used to identify how and where a given website is hosted. Instead, a number of ad hoc techniques have emerged, e.g., using Autonomous System databases, domain prefixes for CNAME records. In this work we propose Pythia , a novel lightweight approach for identifying Web content hosted on third-party infrastructures, including both traditional Web hosts and content delivery networks. Our framework identifies the organization to which a given Web page belongs, and it detects which Web servers are self-hosted and which ones leverage third-party services to provide contents. To test our framework we run it on 40,000 URLs and evaluate its accuracy, both by comparing the results with similar services and with a manually validated groundtruth. Our tool achieves an accuracy of 90% and detects that under 11% of popular domains are self-hosted. We publicly release our tool to allow other researchers to reproduce our findings, and to apply it to their own studies. Srdjan Matic, Gareth Tyson, Gianluca Stringhini |
WWW | 2 |
| 2019 | A Large-scale Behavioural Analysis of Bots and Humans on TwitterabstractRecent research has shown a substantial active presence of bots in online social networks (OSNs). In this article, we perform a comparative analysis of the usage and impact of bots and humans on Twitter—one of the largest OSNs in the world. We collect a large-scale Twitter dataset and define various metrics based on tweet metadata. Using a human annotation task, we assign “bot” and “human” ground-truth labels to the dataset and compare the annotations against an online bot detection tool for evaluation. We then ask a series of questions to discern important behavioural characteristics of bots and humans using metrics within and among four popularity groups. From the comparative analysis, we draw clear differences and interesting similarities between the two entities. Zafar Gilani, Reza Farahbakhsh, Gareth Tyson, Jon Crowcroft |
ACM Trans. Web | 3 |
| 2018 | WhatApp Doc? A First Look at WhatsApp Public Group Data
Venkata Rama Kiran Garimella, Gareth Tyson |
ICWSM | 2 |
| 2018 | Movie Pirates of the Caribbean: Exploring Illegal Streaming Cyberlockers
Damilola Ibosiola, Benjamin A. Steer, Álvaro García-Recuero, Gianluca Stringhini, Steve Uhlig, Gareth Tyson |
ICWSM | 6 |
| 2018 | Facebook (A)Live?: Are Live Social Broadcasts Really Broadcasts?abstractThe era of live-broadcast is back but with two major changes. First, unlike traditional TV broadcasts, content is now streamed over the Internet enabling it to reach a wider audience. Second, due to various user-generated content platforms it has become possible for anyone to get involved, streaming their own content to the world. This emerging trend of going live usually happens via social platforms, where users perform live social broadcasts predominantly from their mobile devices, allowing their friends (and the general public) to engage with the stream in real-time. With the growing popularity of such platforms, the burden on the current Internet infrastructure is therefore expected to multiply. With this in mind, we explore one such prominent platform - Facebook Live. We gather 3TB of data, representing one month of global activity and explore the characteristics of live social broadcast. From this, we derive simple yet effective principles which can decrease the network burden. We then dissect global and hyper-local properties of the video while on-air, by capturing the geography of the broadcasters or the users who produce the video and the viewers or the users who interact with it. Finally, we study the social engagement while the video is live and distinguish the key aspects when the same video goes on-demand. A common theme throughout the paper is that, despite its name, many attributes of Facebook Live deviate from both the concepts of live and broadcast. Aravindh Raman, Gareth Tyson, Nishanth Sastry |
WWW | 2 |
| 2018 | Exploring and Analysing the African Web EcosystemabstractIt is well known that internet infrastructure deployment is progressing at a rapid pace in the African continent. A flurry of recent research has quantified this, highlighting the expansion of its underlying connectivity network. However, improving the infrastructure is not useful without appropriately provisioned services to exploit it. This article measures the availability and utilisation of web infrastructure in Africa. Whereas others have explored web infrastructure in developed regions, we shed light on practices in developing regions. To achieve this, we apply a comprehensive measurement methodology to collect data from a variety of sources. We first focus on Google to reveal that its content infrastructure in Africa is, indeed, expanding. That said, we find that much of its web content is still served from the US and Europe, despite being the most popular website in many African countries. We repeat the same analysis across a number of other regionally popular websites to find that even top African websites prefer to host their content abroad. To explore the reasons for this, we evaluate some of the major bottlenecks facing content delivery networks (CDNs) in Africa. Amongst other factors, we find a lack of peering between the networks hosting our probes, preventing the sharing of CDN servers, as well as poorly configured DNS resolvers. Finally, our mapping of middleboxes in the region reveals that there is a greater presence of transparent proxies in Africa than in Europe or the US. We conclude the work with a number of suggestions for alleviating the issues observed. Rodérick Fanou, Gareth Tyson, Eder Leão Fernandes, Pierre François, Francisco Valera, Arjuna Sathiaseelan |
ACM Trans. Web | 2 |
| 2017 | Of Bots and Humans (on Twitter)abstractRecent research has shown a substantial active presence of bots in online social networks (OSNs). In this paper we utilise our previous work (Stweeler) to comparatively analyse the usage and impact of bots and humans on Twitter, one of the largest OSNs in the world. We collect a large-scale Twitter dataset and define various metrics based on tweet metadata. Using a human annotation task we assign 'bot' and 'human' ground-truth labels to the dataset, and compare the annotations against an online bot detection tool for evaluation. We then ask a series of questions to discern important behavioural characteristics of bots and humans using metrics within and among four popularity groups. From the comparative analysis we draw differences and interesting similarities between the two entities, thus paving the way for reliable classification of bots, and studying automated political infiltration and advertisement campaigns. Zafar Gilani, Reza Farahbakhsh, Gareth Tyson, Liang Wang 0009, Jon Crowcroft |
ASONAM | 3 |
| 2017 | Fake it till you make it: Fishing for CatfishesabstractMany adult content websites incorporate social networking features. Although these are popular, they raise significant challenges, including the potential for users to "catfish", i.e., to create fake profiles to deceive other users. This paper takes an initial step towards automated catfish detection. We explore the characteristics of the different age and gender groups, identifying a number of distinctions. Through this, we train models based on user profiles and comments, via the ground truth of specially verified profiles. When applying our models for age and gender estimation to unverified profiles, 38% of profiles are classified as lying about their age, and 25% are predicted to be lying about their gender. The results suggest that women have a greater propensity to catfish than men. Our preliminary work has notable implications on operators of such online social networks, as well as users who may worry about interacting with catfishes. Walid Magdy, Yehia El-khatib, Gareth Tyson, Sagar Joglekar 0001, Nishanth Sastry |
ASONAM | 3 |
| 2017 | Unbiased Sampling of Social Media Networks for Well-connected SubgraphsabstractSampling social graphs is critical for studying things like information diffusion. However, it is often necessary to laboriously obtain unbiased and well-connected datasets because existing survey algorithms are unable to generate well-connected samples, and current random-walk based unbiased sampling algorithms adopt rejection sampling, which heavily undermines performance. This paper proposes a novel random-walk based algorithm which implements Unbiased Sampling using Dummy Edges (USDE). It injects dummy edges between nodes, on which the walkers would otherwise experience excessive rejections before moving out from such nodes. We propose a rejection probability estimation algorithm to facilitate the construction of dummy edges and the computation of moving probabilities. Finally, we apply USDE in two real-life social media: Twitter and Sina Weibo. The results demonstrate that USDE generates well-connected samples, and outperforms existing approaches in terms of sampling efficiency and quality of samples. Dong Wang 0027, Zhenyu Li 0001, Gareth Tyson, Zhenhua Li 0001, Gaogang Xie |
ASONAM | 3 |
| 2017 | Exploring HTTP Header Manipulation In-The-WildabstractHeaders are a critical part of HTTP, and it has been shown that they are increasingly subject to middlebox manipulation. Although this is well known, little is understood about the general regional and network trends that underpin these manipulations. In this paper, we collect data on thousands of networks to understand how they intercept HTTP headers in-the-wild. Our analysis reveals that 25% of measured ASes modify HTTP headers. Beyond this, we witness distinct trends among different regions and AS types; e.g., we observe high numbers of cache headers in poorly connected regions. Finally, we perform an in-depth analysis of the types of manipulations and how they differ across regions. Gareth Tyson, Félix Cuadrado, Ignacio Castro, Vasile Claudiu Perta, Arjuna Sathiaseelan, Steve Uhlig |
WWW | 1 |
| 2016 | A first look at user activity on tinderabstractMobile dating apps have become a popular means to meet potential partners. Although several exist, one recent addition stands out amongst all others. Tinder presents its users with pictures of people geographically nearby, whom they can either like or dislike based on first impressions. If two users like each other, they are allowed to initiate a conversation via the chat feature. In this paper we use a set of curated profiles to explore the behaviour of men and women in Tinder. We reveal differences between the way men and women interact with the app, highlighting the strategies employed. Women attain large numbers of matches rapidly, whilst men only slowly accumulate matches. Most notably, our results indicate that a little effort in grooming profiles, especially for male users, goes a long way in attracting attention. Gareth Tyson, Vasile Claudiu Perta, Hamed Haddadi 0001, Michael C. Seto |
ASONAM | 1 |
| 2016 | Pub Crawling at Scale: Tapping Untappd to Explore Social Drinking
Martin J. Chorley, Luca Rossi 0004, Gareth Tyson, Matthew J. Williams |
ICWSM | 3 |
| 2016 | Pushing the Frontier: Exploring the African Web EcosystemabstractIt is well known that Africa's mobile and fixed Internet infrastructure is progressing at a rapid pace. A flurry of recent research has quantified this, highlighting the expansion of its underlying connectivity network. However, improving the infrastructure is not useful without appropriately provisioned services to utilise it. This paper measures the availability of web content infrastructure in Africa. Whereas others have explored web infrastructure in developed regions, we shed light on practices in developing regions. To achieve this, we apply a comprehensive measurement methodology to collect data from a variety of sources. We focus on a large content delivery network to reveal that Africa's content infrastructure is, indeed, expanding. However, we find much web content is still served from the US and Europe. We discover that many of the problems faced are actually caused by significant inter-AS delays in Africa, which contribute to local ISPs not sharing their cache capacity. We discover that a related problem is the poor DNS configuration used by some ISPs, which confounds the attempts of providers to optimise their delivery. We then explore a number of other websites to show that large web infrastructure deployments are a rarity in Africa and that even regional websites host their services abroad. We conclude by making suggestions for improvements. Rodérick Fanou, Gareth Tyson, Pierre François, Arjuna Sathiaseelan |
WWW | 2 |
| 2015 | Are People Really Social in Porn 2.0?
Gareth Tyson, Yehia El-khatib, Nishanth Sastry, Steve Uhlig |
ICWSM | 1 |