Rui Zhao 0009

dblp:26/2578-9 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0003-2993-2023ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 'They've Stolen My GPL-Licensed Model!': Toward Standardized and Transparent Model Licensing
abstract
As model parameter sizes scale into the billions and training consumes zettaFLOPs of computation, the reuse of Machine Learning (ML) assets and collaborative development have become increasingly prevalent in the ML community. These ML assets, including models, datasets, and software, may originate from various sources and be published under different licenses, which govern the use and distribution of licensed works and their derivatives. However, commonly chosen licenses, such as GPL and Apache, are software-specific and are not clearly defined or bounded in the context of model publishing. Meanwhile, the reused assets may also be under free-content licenses and model licenses, which pose a potential risk of license noncompliance and rights infringement within the model production workflow. In this paper, we address these challenges along two lines: 1) For ML workflow compliance, we propose ModelGo (MG) Analyzer, a tool that incorporates a vocabulary for ML workflow management and encoded license rules, enabling ontological reasoning to analyze rights granting and compliance issues. 2) For standardized model publishing, we introduce ModelGo Licenses, a set of modell-specific licenses that provide flexible options to meet the diverse needs of the ML community. MG Analyzer is built on Turtle language and Notation3 reasoning engine, envisioned as a first step toward Linked Open Data for ML workflow management. We have also encoded our proposed model licenses into rules and demonstrated the effects of GPL and other commonly used licenses in model publishing, along with the flexibility advantages of our licenses, through comparisons and experiments.
Moming Duan, Rui Zhao 0009, Linshan Jiang, Nigel Shadbolt, Bingsheng He
WWW2
2025 Libertas: Privacy-Preserving Collaborative Computation for Decentralised Personal Data Stores
abstract
Data and their processing have become an indispensable aspect for our society. Insights drawn from collective data make invaluable contribution to scientific, societal and communal research and business. However, there are increasing worries about privacy issues and data misuse, prompting the emergence of decentralised personal data stores (PDS) like Solid. However, existing PDS frameworks face challenges in ensuring data privacy when performing collective computation to combine data from multiple users. At a glance, Secure Multi-Party Computation (MPC) offers input secrecy protection while performing collective computation without relying on any single party. However, issues emerge when directly applying MPC in the context of PDS, particularly due to key factors like autonomy and decentralisation. In this work, we discuss the essence of this issue, identify the potential solution, and introduce a modular system architecture, Libertas, to integrate MPC with PDS like Solid, without requiring protocol-level changes. We introduce the paradigm shift from an 'omniscient' view to individual-based, user-centric view of trust and security, and discuss the threat model of Libertas. Two realistic use cases for collaborative data processing are used for evaluation, both for technical feasibility and empirical benchmark, highlighting its effectiveness in empowering gig workers and generating differentially private synthetic data. The results of our experiments underscore Libertas' linear scalability and provide valuable insights into compute optimisations, thereby advancing the state-of-the-art in privacy-preserving data processing practices. By offering practical solutions for maintaining both individual autonomy and privacy in collaborative data processing environments, Libertas contributes significantly to the ongoing discourse on privacy protection in data-driven decision-making contexts.
Rui Zhao 0009, Naman Goel, Nitin Agrawal 0002, Jun Zhao 0003, Jake M. L. Stein, Wael S. Albayaydh, Ruben Verborgh, Reuben Binns, Tim Berners-Lee, Nigel Shadbolt
Proc. ACM Hum. Comput. Interact.1
2024 Perennial Semantic Data Terms of Use for Decentralized Web
abstract
In today's digital landscape, the Web has become increasingly centralized, raising concerns about user privacy violations. Decentralized Web architectures, such as Solid, offer a promising solution by empowering users with better control over their data in their personal 'Pods'. However, a significant challenge remains: users must navigate numerous applications to decide which application can be trusted with access to their data Pods. This often involves reading lengthy and complex Terms of Use agreements, a process that users often find daunting or simply ignore. This compromises user autonomy and impedes detection of data misuse. We propose a novel formal description of Data Terms of Use (DToU), along with a DToU reasoner. Users and applications specify their own parts of the DToU policy with local knowledge, covering permissions, requirements, prohibitions and obligations. Automated reasoning verifies compliance, and also derives policies for output data. This constitutes a "perennial'' DToU language, where the policy authoring only occurs once, and we can conduct ongoing automated checks across users, applications and activity cycles. Our solution is built on Turtle, Notation 3 and RDF Surfaces, for the language and the reasoning engine. It ensures seamless integration with other semantic tools for enhanced interoperability. We have successfully integrated this language into the Solid framework, and conducted performance benchmark. We believe this work demonstrates a practicality of a perennial DToU language and the potential of a paradigm shift to how users interact with data and applications in a decentralized Web, offering both improved privacy and usability.
Rui Zhao 0009, Jun Zhao 0003
WWW1
2024 Trouble in Paradise? Understanding Mastodon Admin's Motivations, Experiences, and Challenges Running Decentralised Social Media
abstract
Decentralised social media platforms are increasingly being recognised as viable alternatives to their centralised counterparts. Among these, Mastodon stands out as a popular alternative, offering a citizen-powered option distinct from larger and centralised platforms like Twitter/X. However, the future path of Mastodon remains uncertain, particularly in terms of its challenges and the long-term viability of a more citizen-powered internet. In this paper, following a pre-study survey, we conducted semi-structured interviews with 16 Mastodon instance administrators, including those who host instances to support marginalised and stigmatised communities, to understand their motivations and lived experiences of running decentralised social media. Our research indicates that while decentralised social media offers significant potential in supporting the safety, identity and privacy needs of marginalised and stigmatised communities, they also face considerable challenges in content moderation, community building and governance. We emphasise the importance of considering the community's values and diversity when designing future support mechanisms.
Zhilin Zhang 0004, Jun Zhao 0003, Ge Wang 0004, Samantha-Kaye Johnston, George Chalhoub, Tala Ross, Claudine Tinsman, Rui Zhao 0009, Max Van Kleek, Nigel Shadbolt
Proc. ACM Hum. Comput. Interact.9
2023 'You are you and the app. There's nobody else.': Building Worker-Designed Data Institutions within Platform Hegemony
abstract
Information asymmetries create extractive, often harmful relationships between platform workers (e.g., Uber or Deliveroo drivers) and their algorithmic managers. Recent HCI studies have put forward more equitable platform designs but leave open questions about the social and technical infrastructures required to support them without the cooperation of platforms. We conducted a participatory design study in which platform workers deconstructed and re-imagined Uber’s schema for driver data. We analyzed the data structures and social institutions participants proposed, focusing on the stakeholders, roles, and strategies for mitigating conflicting interests of privacy, personal agency, and utility. Using critical theory, we reflected on the capability of participatory design to generate bottom-up collective data infrastructures. Based on the plurality of alternative institutions participants produced and their aptitude to navigate data stewardship decisions, we propose user-configurable tools for lightweight data institution building, as an alternative to redesigning existing platforms or delegating control to centralized trusts.
Jake M. L. Stein, Vidminas Vizgirda, Max Van Kleek, Reuben Binns, Jun Zhao 0003, Rui Zhao 0009, Naman Goel, George Chalhoub, Wael S. Albayaydh, Nigel Shadbolt
CHI6
2021 An Automated Framework for Supporting Data-Governance Rule Compliance in Decentralized MIMO Contexts
abstract
We propose Dr.Aid, a logic-based AI framework for automated compliance checking of data governance rules over data-flow graphs. The rules are modelled using a formal language based on situation calculus and are suitable for decentralized contexts with multi-input-multi-output (MIMO) processes. Dr.Aid models data rules and flow rules and checks compliance by reasoning about the propagation, combination, modification and application of data rules over the data flow graphs. Our approach is driven and evaluated by real-world datasets using provenance graphs from data-intensive research.
Rui Zhao 0009
IJCAI1
2021 Dr.Aid: Supporting Data-governance Rule Compliance for Decentralized Collaboration in an Automated Way
abstract
Collaboration across institutional boundaries is widespread and increasing today. It depends on federations sharing data that often have governance rules or external regulations restricting their use. However, the handling of data governance rules (aka. data-use policies) remains manual, time-consuming and error-prone, limiting the rate at which collaborations can form and respond to challenges and opportunities, inhibiting citizen science and reducing data providers' trust in compliance. Using an automated system to facilitate compliance handling reduces substantially the time needed for such non-mission work, thereby accelerating collaboration and improving productivity. We present a framework, Dr.Aid, that helps individuals, organisations and federations comply with data rules, using automation to track which rules are applicable as data is passed between processes and as derived data is generated. It encodes data-governance rules using a formal language and performs reasoning on multi-input-multi-output data-flow graphs in decentralised contexts. We test its power and utility by working with users performing cyclone tracking and earthquake modelling to support mitigation and emergency response. We query standard provenance traces to detach Dr.Aid from details of the tools and systems they are using, as these inevitably vary across members of a federation and through time. We evaluate the model in three aspects by encoding real-life data-use policies from diverse fields, showing its capability for real-world usage and its advantages compared with traditional frameworks. We argue that this approach will lead to more agile, more productive and more trustworthy collaborations and show that the approach can be adopted incrementally. This, in-turn, will allow more appropriate data policies to emerge opening up new forms of collaboration.
Rui Zhao 0009, Malcolm P. Atkinson 0001, Petros Papapanagiotou, Federica Magnoni, Jacques D. Fleuriot
Proc. ACM Hum. Comput. Interact.1
2019 Towards a Computer-Interpretable Actionable Formal Model to Encode Data Governance Rules
abstract
With the needs of science and business, data sharing and re-use has become an intensive activity for various areas. In many cases, governance imposes rules concerning data use, but there is no existing computational technique to help data-users comply with such rules. We argue that intelligent systems can be used to improve the situation, by recording provenance records during processing, encoding the rules and performing reasoning. We present our initial work, designing formal models for data rules and flow rules and the reasoning system, as the first step towards helping data providers and data users sustain productive relationships.
Rui Zhao 0009, Malcolm P. Atkinson 0001
eScience1