Jun Zhao 0003

dblp:47/2026-3 · DBLP profile ↗
← Back
46ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0001-6935-9028ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 25 · 3 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-authorSecurity and privacy · 2 · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Supporting Parents' Playfulness With Their Preschool Children Using Generative AI
abstract
Playful interactions between parents and their young children are highly beneficial for the parent-child relationship and the child’s social, emotional, and cognitive development. However, parents may struggle to initiate these interactions. To support them, we have developed In-Between, an app that takes brief input from the parent about their situation and uses generative AI to devise three suggestions to kick off a playful interaction: a story starter, a playful question, and a game idea. We share some reflections from initial conversations with parents and outline the research plan moving forwards.
Isobel Voysey, Trisha-mae Capistrano, Jun Zhao 0003
IDC3
2026 Agency in Child-AI Interaction: A Review of How It Is Conceptualised, Studied, and Supported in HCI
abstract
Children’s lives are increasingly intertwined with AI systems, from recommender algorithms to generative models, raising concerns about potential impacts on children’s agency. Although supporting human agency, autonomy, and empowerment is a widely shared HCI goal, we lack clear definitions of these concepts in designing child-AI interaction. Through a review of 25 recent HCI studies, we find agency is rarely explicitly defined and its conceptualisation varies across something children innately possess and something to be developed. Our literature mapping shows that researchers observed agency through children’s planning and self-regulation, asserting control over AI systems, and critique and re-design of the status quo. Conditions reported by researchers that enable or constrain agency span epistemic conditions, interactional design, social context, and motivational orientation. Our review highlights gaps in research on designing for children’s agency. We advocate for conceptual clarity by drawing upon existing frameworks and highlight the importance of considering children’s agency through a relational lens.
Isobel Voysey, Vidminas Vizgirda, Sarah Turner, Leslye Denisse Dias Duran, Zaki Pauzi, Manolis Mavrikis, Carina Prunkl, Jun Zhao 0003
IDC8
2026 Attitudes, Imagined Roles, and Governance Boundaries for AI in Decentralized Social Media
abstract
Decentralised social media (DSM) platforms such as Mastodon offer community-governed alternatives to corporate social networks but place substantial governance burdens on volunteer operators. As interest grows in applying artificial intelligence (AI) to support this work, little is known about whether DSM operators want AI, what roles they consider appropriate, and what governance boundaries they require. We conducted semi-structured interviews with 20 operators across Mastodon, Pixelfed, PeerTube, Lemmy, Pleroma, and Funkwhale, using generative feature probes and speculative scenarios to explore their perceptions of AI. Operators rejected AI as an autonomous actor, instead envisioning it as governance infrastructure that provides contextual intelligence, supports cross-instance coordination, and sustains community and moderator well-being. They also articulated strict boundaries rooted in DSM values, including human accountability, reversibility, transparency, community-centred configuration, and strong data-governance constraints. We contribute empirical insights and design implications for AI compatible with decentralised, federated social media.
Zhilin Zhang 0004, Jun Zhao 0003, Ge Wang 0004, Sruthi Viswanathan, Tala Ross, Samantha-Kaye Johnston, Hayoun Noh, Max Van Kleek, Nigel Shadbolt
CHI2
2026 Editorial for the special issue on child-centred AI
Jun Zhao 0003, Grace C. Lin, Jason C. Yip 0001, Zhen Bai 0002, Ayça Atabey, Ge Wang 0004, Kaiwen Sun 0001
Int. J. Hum. Comput. Stud.1
2025 FamiData Hub: A Speculative Design Exploration with Families on Smart Home Datafication
Ge Wang 0004, Jun Zhao 0003, Max Van Kleek, Roy D. Pea, Nigel Shadbolt
CHI2
2025 Governance of Generative AI in Creative Work: Consent, Credit, Compensation, and Beyond
Lin Kyi, Amruta Mahuli, Michael Six Silberman, Reuben Binns, Jun Zhao 0003, Asia J. Biega
CHI5
2025 Libertas: Privacy-Preserving Collaborative Computation for Decentralised Personal Data Stores
abstract
Data and their processing have become an indispensable aspect for our society. Insights drawn from collective data make invaluable contribution to scientific, societal and communal research and business. However, there are increasing worries about privacy issues and data misuse, prompting the emergence of decentralised personal data stores (PDS) like Solid. However, existing PDS frameworks face challenges in ensuring data privacy when performing collective computation to combine data from multiple users. At a glance, Secure Multi-Party Computation (MPC) offers input secrecy protection while performing collective computation without relying on any single party. However, issues emerge when directly applying MPC in the context of PDS, particularly due to key factors like autonomy and decentralisation. In this work, we discuss the essence of this issue, identify the potential solution, and introduce a modular system architecture, Libertas, to integrate MPC with PDS like Solid, without requiring protocol-level changes. We introduce the paradigm shift from an 'omniscient' view to individual-based, user-centric view of trust and security, and discuss the threat model of Libertas. Two realistic use cases for collaborative data processing are used for evaluation, both for technical feasibility and empirical benchmark, highlighting its effectiveness in empowering gig workers and generating differentially private synthetic data. The results of our experiments underscore Libertas' linear scalability and provide valuable insights into compute optimisations, thereby advancing the state-of-the-art in privacy-preserving data processing practices. By offering practical solutions for maintaining both individual autonomy and privacy in collaborative data processing environments, Libertas contributes significantly to the ongoing discourse on privacy protection in data-driven decision-making contexts.
Rui Zhao 0009, Naman Goel, Nitin Agrawal 0002, Jun Zhao 0003, Jake M. L. Stein, Wael S. Albayaydh, Ruben Verborgh, Reuben Binns, Tim Berners-Lee, Nigel Shadbolt
Proc. ACM Hum. Comput. Interact.4
2024 CHAITok: A Proof-of-Concept System Supporting Children's Sense of Data Autonomy on Social Media
abstract
Social media has become a primary source of entertainment and education for children globally. While much attention has been given to children’s online well-being, a pressing concern often goes unnoticed: the pervasive data harvesting underlying social media and its manipulative impact on undermining children’s autonomy. In this paper, we present CHAITok, an Android mobile application designed to enhance children’s sense of autonomy over their data on social media. Through 27 user study sessions with 109 children aged 10–13, we offer insights into the current lack of data autonomy among children regarding their online information, and how we can foster children’s sense of data autonomy through a socio-technical journey. Our findings inspire design recommendations to respect children’s values, support children’s evolving autonomy, and design for children’s digital rights. We emphasize data autonomy as a fundamental right for children, call for further research, design innovation, and policy changes on this critical issue.
Ge Wang 0004, Jun Zhao 0003, Samantha-Kaye Johnston, Zhilin Zhang 0004, Max Van Kleek, Nigel Shadbolt
CHI2
2024 KOALA Hero Toolkit: A New Approach to Inform Families of Mobile Datafication Risks
abstract
Children today are deeply immersed in the online world, where their activities are routinely tracked, analysed, and monetised. This exposes them to various datafication risks, including harmful profiling, micro-targeting and behavioural engineering. Most existing measures focus on immediate online threats, rather than informing children about these implicit risks. In this paper, we present The KOALA Hero Toolkit, a hybrid toolkit designed to help children and parents jointly understand the datafication risks posed by their mobile apps. Through user studies involving 17 families we evaluate how the toolkit influenced families’ thought processes, perceptions and decision-making regarding mobile datafication risks. Our findings show that KOALA Hero supports families’ critical thinking and promotes family engagement. We identify future design recommendations for family support, featuring ideas such as integrating triggering moments and bonding moments in toolkit designs. This work provides timely inputs on global efforts aimed at addressing datafication risks and underscores the importance of strengthening legislative and policy enforcement of ethical data governance.
Ge Wang 0004, Jun Zhao 0003, Konrad Kollnig, Adrien Zier, Blanche Duron, Zhilin Zhang 0004, Max Van Kleek, Nigel Shadbolt
CHI2
2024 Perennial Semantic Data Terms of Use for Decentralized Web
abstract
In today's digital landscape, the Web has become increasingly centralized, raising concerns about user privacy violations. Decentralized Web architectures, such as Solid, offer a promising solution by empowering users with better control over their data in their personal 'Pods'. However, a significant challenge remains: users must navigate numerous applications to decide which application can be trusted with access to their data Pods. This often involves reading lengthy and complex Terms of Use agreements, a process that users often find daunting or simply ignore. This compromises user autonomy and impedes detection of data misuse. We propose a novel formal description of Data Terms of Use (DToU), along with a DToU reasoner. Users and applications specify their own parts of the DToU policy with local knowledge, covering permissions, requirements, prohibitions and obligations. Automated reasoning verifies compliance, and also derives policies for output data. This constitutes a "perennial'' DToU language, where the policy authoring only occurs once, and we can conduct ongoing automated checks across users, applications and activity cycles. Our solution is built on Turtle, Notation 3 and RDF Surfaces, for the language and the reasoning engine. It ensures seamless integration with other semantic tools for enhanced interoperability. We have successfully integrated this language into the Solid framework, and conducted performance benchmark. We believe this work demonstrates a practicality of a perennial DToU language and the potential of a paradigm shift to how users interact with data and applications in a decentralized Web, offering both improved privacy and usability.
Rui Zhao 0009, Jun Zhao 0003
WWW2
2024 Privacy in Chinese iOS apps and impact of the personal information protection law
abstract
Privacy in apps is a topic of widespread interest because many apps collect and share large amounts of highly sensitive information. In response, the Chinese legislator introduced a range of new data protection laws over recent years, notably the Personal Information Protection Law (PIPL) in 2021. So far, there exists limited research on the impacts of these new laws on apps’ privacy practices. To address this gap, this paper analyses data collection in pairs of 634 Chinese iOS apps, one version from early 2020 and one from late 2021. Our work finds that many more apps now implement consent. Yet, those end-users that decline consent will often be forced to exit the app. Fewer apps now collect data without consent but many still integrate tracking libraries. Market concentration in app data collection has seen limited change. At the same time, there exists a larger number of influential and equal market participants than in the West. Among them, Apple was the only relevant foreign company. We see our findings characteristic of a first iteration at Chinese data regulation with room for improvement. With the help of enhanced technological capabilities, we expect increased enforcement of the new data rules. There is also room to refine the new laws and make them more targeted at mobile apps and the online sphere, particularly through clear and up-to-date technical specifications for software developers. As such, our findings could also be motivation for non-Chinese policy- and lawmakers to enhance their own data protection regimes.
Konrad Kollnig, Lu Zhang 0073, Jun Zhao 0003, Nigel Shadbolt
Comput. Law Secur. Rev.3
2024 Trouble in Paradise? Understanding Mastodon Admin's Motivations, Experiences, and Challenges Running Decentralised Social Media
abstract
Decentralised social media platforms are increasingly being recognised as viable alternatives to their centralised counterparts. Among these, Mastodon stands out as a popular alternative, offering a citizen-powered option distinct from larger and centralised platforms like Twitter/X. However, the future path of Mastodon remains uncertain, particularly in terms of its challenges and the long-term viability of a more citizen-powered internet. In this paper, following a pre-study survey, we conducted semi-structured interviews with 16 Mastodon instance administrators, including those who host instances to support marginalised and stigmatised communities, to understand their motivations and lived experiences of running decentralised social media. Our research indicates that while decentralised social media offers significant potential in supporting the safety, identity and privacy needs of marginalised and stigmatised communities, they also face considerable challenges in content moderation, community building and governance. We emphasise the importance of considering the community's values and diversity when designing future support mechanisms.
Zhilin Zhang 0004, Jun Zhao 0003, Ge Wang 0004, Samantha-Kaye Johnston, George Chalhoub, Tala Ross, Claudine Tinsman, Rui Zhao 0009, Max Van Kleek, Nigel Shadbolt
Proc. ACM Hum. Comput. Interact.2
2023 12 Ways to Empower: Designing for Children's Digital Autonomy
abstract
In recent years, growing research has been made on supporting children to become more autonomous in the digital environment around them. However, there has been little consensus regarding the conceptualisation of digital autonomy for children in the HCI community and how best they can be supported. Through a systematic review of autonomy-supportive designs within HCI research, this paper makes three contributions: a landscape overview of the existing conceptualisation of Digital Autonomy for children within HCI; a framework of 12 distinct design mechanisms for supporting children’s digital autonomy, clustered into 5 categories by their common mechanisms; and an identification of 5 critical design considerations for future support of children’s digital autonomy. Our findings provide a critical understanding of current support for children’s digital autonomy in HCI. We highlight the importance of considering children’s digital autonomy from multi-perspectives and suggest critical factors and gaps to be considered for future autonomy-supportive designs.
Ge Wang 0004, Jun Zhao 0003, Max Van Kleek, Nigel Shadbolt
CHI2
2023 'You are you and the app. There's nobody else.': Building Worker-Designed Data Institutions within Platform Hegemony
abstract
Information asymmetries create extractive, often harmful relationships between platform workers (e.g., Uber or Deliveroo drivers) and their algorithmic managers. Recent HCI studies have put forward more equitable platform designs but leave open questions about the social and technical infrastructures required to support them without the cooperation of platforms. We conducted a participatory design study in which platform workers deconstructed and re-imagined Uber’s schema for driver data. We analyzed the data structures and social institutions participants proposed, focusing on the stakeholders, roles, and strategies for mitigating conflicting interests of privacy, personal agency, and utility. Using critical theory, we reflected on the capability of participatory design to generate bottom-up collective data infrastructures. Based on the plurality of alternative institutions participants produced and their aptitude to navigate data stewardship decisions, we propose user-configurable tools for lightweight data institution building, as an alternative to redesigning existing platforms or delegating control to centralized trusts.
Jake M. L. Stein, Vidminas Vizgirda, Max Van Kleek, Reuben Binns, Jun Zhao 0003, Rui Zhao 0009, Naman Goel, George Chalhoub, Wael S. Albayaydh, Nigel Shadbolt
CHI5
2023 'Treat me as your friend, not a number in your database': Co-designing with Children to Cope with Datafication Online
abstract
Datafication refers to the practices through which children’s online actions are pervasively recorded, tracked, aggregated, analysed, and exploited by online services in ways including behavioural engineering and monetisation. Previous research has shown that not only do children care significantly about various aspects of datafication, but they demand a chance to take action. Through 10 co-design sessions with 53 children, we examined how children in the UK want to be supported to cope with the datafication practices. Our findings provide insights for creating age-appropriate support for children’s algorithmic literacy development, highlighting and unpacking the importance of no one-size-fitting-all designs to support children’s coping with datafication. We contribute a first understanding of how children aged 7–14 would like to be supported with datafication and what future data-driven digital experiences should be like for them, who demand a shift of the current data ecosystem towards a more humane-by-design and autonomy-supportive future.
Ge Wang 0004, Jun Zhao 0003, Max Van Kleek, Nigel Shadbolt
CHI2
2023 How Can We Design Privacy-Friendly Apps for Children? Using a Research through Design Process to Understand Developers' Needs and Challenges
abstract
Mobile apps used by children often make use of harmful techniques, such as data tracking and targeted advertising. Previous research has suggested that developers face several systemic challenges in designing apps that prioritise children's best interests. To understand how developers can be better supported, we used a Research through Design (RtD) method to explore what the future of privacy-friendly app development could look like. We performed an elicitation study with 20 children's app developers to understand their needs and requirements. We found a number of specific technical requirements from the participants about how they would like to be supported, such as having actionable transnational design guidelines and easy-to-use development libraries. However, participants were reluctant to adopt these design ideas in their development practices due to perceived financial risks associated with increased privacy in apps. To overcome this critical gap, participants formulated socio-technical requirements that extend to other stakeholders in the mobile industry, including parents and marketplaces. Our findings provide important immediate and long-term design opportunities for the HCI community, and indicate that support for changing app developers' practices must be designed in the context of their relationship with other stakeholders.
Anirudh Ekambaranathan, Jun Zhao 0003, Max Van Kleek
Proc. ACM Hum. Comput. Interact.2
2022 KOALA Hero: Inform Children of Privacy Risks of Mobile Apps
abstract
Children’s online activities are routinely tracked, aggregated, and exploited by online services, to manipulate children’s online behaviour or monetise. This contributes to the so-called datafied childhood. Unfortunately, such datafication remains largely invisible behind the services and is practically impossible to avoid. Existing approaches largely focus on direct online harms, and provide limited support to raise children’s awareness or understanding of how their data may be processed, transmitted across platforms, and used to affect their best interests. Through co-design workshops, we identified key barriers for children and families to cope with this type of data privacy risk. Our contribution is that instead of regarding children as passive users and needing protection, we draw on critical digital literacy theories and design a KOALA Hero app, which is aimed to enhance children’s cognitive, situated and critical thinking of datafication and online data privacy risks. KOALA Hero represents our first step towards facilitating children’s understanding of the invisible data privacy risks. We hope future empirical evaluations will further inform us regarding how our design approaches may affect the thinking process and behaviours of children and families.
Jun Zhao 0003, Blanche Duron, Ge Wang 0004
IDC1
2022 Poster: An Analysis of Privacy Features in 'Expert-Approved' Kids' Apps
abstract
During the course of the past decade, children have become avid consumers of digital media through mobile devices. The industry for children's mobile applications is booming and marketplaces offer categories of apps aimed specifically at children. In this study, we perform a mixed-methods privacy analysis of 137 'expert-approved' children's apps from the Google Play Store. Our findings show that these apps do not sufficiently support children to exercise their privacy rights, whilst simultaneously making use of libraries and data trackers which may collect and share sensitive user data.
Anirudh Ekambaranathan, Jun Zhao 0003, Max Van Kleek
CCS2
2022 Informing Age-Appropriate AI: Examining Principles and Practices of AI for Children
abstract
AI systems are becoming increasingly pervasive within children’s devices, apps, and services. However, it is not yet well-understood how risks and ethical considerations of AI relate to children. This paper makes three contributions to this area: first, it identifies ten areas of alignment between general AI frameworks and codes for age-appropriate design for children. Then, to understand how such principles relate to real application contexts, we conducted a landscape analysis of children’s AI systems, via a systematic literature review including 188 papers. This analysis revealed a wide assortment of applications, and that most systems’ designs addressed only a small subset of principles among those we identified. Finally, we synthesised our findings in a framework to inform a new “Code for Age-Appropriate AI”, which aims to provide timely input to emerging policies and standards, and inspire increased interactions between the AI and child-computer interaction communities.
Ge Wang 0004, Jun Zhao 0003, Max Van Kleek, Nigel Shadbolt
CHI2
2022 'Don't make assumptions about me!': Understanding Children's Perception of Datafication Online
abstract
Datafication, which is the process in which children's actions online are pervasively recorded, tracked, aggregated, analysed, and exploited by online services in multiple ways that include behavioural engineering, and monetisation, is becoming increasing common in the online world today. However, we know little about how children feel about such practices and how they perceive datafication. Through online interviews with 48 children aged 7-13 from UK schools, we examined how children perceive datafication practices, especially how such practices could make inference on them. We identified three key knowledge gaps in children's perceptions, including their lack of recognition of who were involved in the data processing and how, data being transmitted across platforms, and their data ownership. Through situating our findings under a critical algorithmic literacy framework, our findings provided some immediate indications regarding how we could better support children in the datafied society through more transparency and autonomy-supportive designs, as well as the need for a fundamental shift of the current data governance structure.
Ge Wang 0004, Jun Zhao 0003, Max Van Kleek, Nigel Shadbolt
Proc. ACM Hum. Comput. Interact.2
2021 "Money makes the world go around": Identifying Barriers to Better Privacy in Children's Apps From Developers' Perspectives
abstract
The industry for children’s apps is thriving at the cost of children’s privacy: these apps routinely disclose children’s data to multiple data trackers and ad networks. As children spend increasing time online, such exposure accumulates to long-term privacy risks. In this paper, we used a mixed-methods approach to investigate why this is happening and how developers might change their practices. We base our analysis against 5 leading data protection frameworks that set out requirements and recommendations for data collection in children’s apps. To understand developers’ perspectives and constraints, we conducted 134 surveys and 20 semi-structured interviews with popular Android children’s app developers. Our analysis revealed that developers largely respect children’s best interests; however, they have to make compromises due to limited monetisation options, perceived harmlessness of certain third-party libraries, and lack of availability of design guidelines. We identified concrete approaches and directions for future research to help overcome these barriers.
Anirudh Ekambaranathan, Jun Zhao 0003, Max Van Kleek
CHI2
2021 Protection or Punishment? Relating the Design Space of Parental Control Apps and Perceptions about Them to Support Parenting for Online Safety
abstract
Parental control apps, which are mobile apps that allow parents to monitor and restrict their children's activities online, are becoming increasingly adopted by parents as a means of safeguarding their children's online safety. However, it is not clear whether these apps are always beneficial or effective in what they aim to do; for instance, the overuse of restriction and surveillance has been found to undermine parent-child relationship and children's sense of autonomy. While previous research has categorised and taken inventory of key features of popular parental control apps, they have not systematically analysed the ways such features were designed or realised in such apps, or in particular how aspects of such designs might relate to parents and children's experiences with such apps. In this work, we investigate this gap, asking specifically: how might children's and parents' perceptions be related to how parental control features were designed? To investigate this question, we conducted an analysis of 58 top Android parental control apps designed for the purpose of promoting children's online safety, finding three major axes of variation in how key restriction and monitoring features were realised: granularity, feedback/transparency, and parent-child communications support. To relate these axes to perceived benefits and problems, we then analysed 3264 app reviews to identify references to aspects of the each of the axes above, to understand children's and parents' views of how such dimensions related to their experiences with these apps. Our findings led towards 1) an understanding of how parental control apps realise their functionalities differently along three axes of variation, 2) an analysis of exactly the ways that such variation influences children's and parents' perceptions, respectively of the usefulness or effectiveness of these apps, and finally 3) an identification of design recommendations and opportunities for future apps by contextualising our findings within existing digital parenting theories.
Ge Wang 0004, Jun Zhao 0003, Max Van Kleek, Nigel Shadbolt
Proc. ACM Hum. Comput. Interact.2
2020 "It's your private information. it's your life.": young people's views of personal data use by online technologies
abstract
Children and young people make extensive and varied use of digital and online technologies, yet issues about how their personal data may be collected and used by online platforms are rarely discussed. Additionally, despite calls to increase awareness, schools often do not cover these topics, instead focusing on online safety issues, such as being approached by strangers, cyberbullying or access to inappropriate content. This paper presents the results of one of the activities run as part of eleven workshops with 13-18 year olds, using co-designed activities to encourage critical thinking. Sets of 'data cards' were used to stimulate discussion about sharing and selling of personal data by online technology companies. Results highlight the desire and need for increased awareness about the potential uses of personal data amongst this age group, and the paper makes recommendations for embedding this into school curriculums as well as incorporating it into interaction design, to allow young people to make informed decisions about their online lives.
Liz Dowthwaite, Helen Creswick, Virginia Portillo, Jun Zhao 0003, Menisha Patel, Elvira Perez, Ansgar R. Koene, Marina Jirotka
IDC4
2020 'I Just Want to Hack Myself to Not Get Distracted': Evaluating Design Interventions for Self-Control on Facebook
abstract
Beyond being the world's largest social network, Facebook is for many also one of its greatest sources of digital distraction. For students, problematic use has been associated with negative effects on academic achievement and general wellbeing. To understand what strategies could help users regain control, we investigated how simple interventions to the Facebook UI affect behaviour and perceived control. We assigned 58 university students to one of three interventions: goal reminders, removed newsfeed, or white background (control). We logged use for 6 weeks, applied interventions in the middle weeks, and administered fortnightly surveys. Both goal reminders and removed newsfeed helped participants stay on task and avoid distraction. However, goal reminders were often annoying, and removing the newsfeed made some fear missing out on information. Our findings point to future interventions such as controls for adjusting types and amount of available information, and flexible blocking which matches individual definitions of 'distraction'.
Ulrik Lyngs, Kai Lukoff, Petr Slovák, William Seymour, Helena Webb, Marina Jirotka, Jun Zhao 0003, Max Van Kleek, Nigel Shadbolt
CHI7
2019 'I make up a silly name': Understanding Children's Perception of Privacy Risks Online
abstract
Children under 11 are often regarded as too young to comprehend the implications of online privacy. Perhaps as a result, little research has focused on younger kids' risk recognition and coping. Such knowledge is, however, critical for designing efficient safeguarding mechanisms for this age group. Through 12 focus group studies with 29 children aged 6-10 from UK schools, we examined how children described privacy risks related to their use of tablet computers and what information was used by them to identify threats. We found that children could identify and articulate certain privacy risks well, such as information oversharing or revealing real identities online; however, they had less awareness with respect to other risks, such as online tracking or game promotions. Our findings offer promising directions for supporting children's awareness of cyber risks and the ability to protect themselves online.
Jun Zhao 0003, Ge Wang 0004, Carys Dally, Petr Slovák, Julian Edbrooke-Childs, Max Van Kleek, Nigel Shadbolt
CHI1
2018 'It's Reducing a Human Being to a Percentage': Perceptions of Justice in Algorithmic Decisions
abstract
Data-driven decision-making consequential to individuals raises important questions of accountability and justice. Indeed, European law provides individuals limited rights to 'meaningful information about the logic' behind significant, autonomous decisions such as loan approvals, insurance quotes, and CV filtering. We undertake three experimental studies examining people's perceptions of justice in algorithmic decision-making under different scenarios and explanation styles. Dimensions of justice previously observed in response to human decision-making appear similarly engaged in response to algorithmic decisions. Qualitative analysis identified several concerns and heuristics involved in justice perceptions including arbitrariness, generalisation, and (in)dignity. Quantitative analysis indicates that explanation styles primarily matter to justice perceptions only when subjects are exposed to multiple different styles---under repeated exposure of one style, scenario effects obscure any explanation effects. Our results suggests there may be no 'best' approach to explaining algorithmic decisions, and that reflection on their automated nature both implicates and mitigates justice dimensions.
Reuben Binns, Max Van Kleek, Michael Veale, Ulrik Lyngs, Jun Zhao 0003, Nigel Shadbolt
CHI5
2018 X-Ray Refine: Supporting the Exploration and Refinement of Information Exposure Resulting from Smartphone Apps
abstract
Most smartphone apps collect and share information with various first and third parties; yet, such data collection practices remain largely unbeknownst to, and outside the control of, end-users. In this paper, we seek to understand the potential for tools to help people refine their exposure to third parties, resulting from their app usage. We designed an interactive, focus-plus-context display called X-Ray Refine (Refine) that uses models of over 1 million Android apps to visualise a person's exposure profile based on their durations of app use. To support exploration of mitigation strategies, emphRefine can simulate actions such as app usage reduction, removal, and substitution. A lab study of emphRefine found participants achieved a high-level understanding of their exposure, and identified data collection behaviours that violated both their expectations and privacy preferences. Participants also devised bespoke strategies to achieve privacy goals, identifying the key barriers to achieving them.
Max Van Kleek, Reuben Binns, Jun Zhao 0003, Adam Slack, Sauyon Lee, Dean Ottewell, Nigel Shadbolt
CHI3
2018 Measuring Third-party Tracker Power across Web and Mobile
abstract
Third-party networks collect vast amounts of data about users via websites and mobile applications. Consolidations among tracker companies can significantly increase their individual tracking capabilities, prompting scrutiny by competition regulators. Traditional measures of market share, based on revenue or sales, fail to represent the tracking capability of a tracker, especially if it spans both web and mobile. This article proposes a new approach to measure the concentration of tracking capability, based on the reach of a tracker on popular websites and apps. Our results reveal that tracker prominence and parent–subsidiary relationships have significant impact on accurately measuring concentration.
Reuben Binns, Jun Zhao 0003, Max Van Kleek, Nigel Shadbolt
ACM Trans. Internet Techn.2
2017 Better the Devil You Know: Exposing the Data Sharing Practices of Smartphone Apps
abstract
Most users of smartphone apps remain unaware of what data about them is being collected, by whom, and how these data are being used. In this mixed methods investigation, we examine the question of whether revealing key data collection practices of smartphone apps may help people make more informed privacy-related decisions. To investigate this question, we designed and prototyped a new class of privacy indicators, called Data Controller Indicators (DCIs), that expose previously hidden information flows out of the apps. Our lab study of DCIs suggests that such indicators do support people in making more confident and consistent choices, informed by a more diverse range of factors, including the number and nature of third-party companies that access users' data. Furthermore, personalised DCIs, which are contextualised against the other apps an individual already uses, enable them to reason effectively about the differential impacts on their overall information exposure.
Max Van Kleek, Ilaria Liccardi, Reuben Binns, Jun Zhao 0003, Daniel J. Weitzner, Nigel Shadbolt
CHI4
2015 ONCAPS: An Ontology-Based Car Purchase Guiding System
Jianfeng Du, Jun Zhao 0003, Jiayi Cheng, Qingchao Su, Jiacheng Liang
APWeb2
2015 Using a suite of ontologies for preserving workflow-centric research objects
abstract
Scientific workflows are a popular mechanism for specifying and automating data-driven in silico experiments. A significant aspect of their value lies in their potential to be reused. Once shared, workflows become useful building blocks that can be combined or modified for developing new experiments. However, previous studies have shown that storing workflow specifications alone is not sufficient to ensure that they can be successfully reused, without being able to understand what the workflows aim to achieve or to re-enact them. To gain an understanding of the workflow, and how it may be used and repurposed for their needs, scientists require access to additional resources such as annotations describing the workflow, datasets used and produced by the workflow, and provenance traces recording workflow executions. In this article, we present a novel approach to the preservation of scientific workflows through the application of research objects—aggregations of data and metadata that enrich the workflow specifications. Our approach is realised as a suite of ontologies that support the creation of workflow-centric research objects. Their design was guided by requirements elicited from previous empirical analyses of workflow decay and repair. The ontologies developed make use of and extend existing well known ontologies, namely the Object Reuse and Exchange (ORE) vocabulary, the Annotation Ontology (AO) and the W3C PROV ontology (PROVO). We illustrate the application of the ontologies for building Workflow Research Objects with a case-study that investigates Huntington’s disease, performed in collaboration with a team from the Leiden University Medial Centre (HG-LUMC). Finally we present a number of tools developed for creating and managing workflow-centric research objects.
Khalid Belhajjame, Jun Zhao 0003, Daniel Garijo, Matthew Gamble, Kristina M. Hettne, Raúl Palma, Eleni Mina, Óscar Corcho, José Manuél Gómez-Pérez, Sean Bechhofer, Graham Klyne, Carole A. Goble
J. Web Semant.2
2013 When History Matters - Assessing Reliability for the Reuse of Scientific Workflows
José Manuél Gómez-Pérez, Esteban García-Cuesta, Aleix Garrido, José Enrique Ruiz, Jun Zhao 0003, Graham Klyne
ISWC (2)5
2012 MIM: A Minimum Information Model vocabulary and framework for Scientific Linked Data
abstract
Linked Data holds great promise in the Life Sciences as a platform to enable an interoperable data commons, supporting new opportunities for discovery. Minimum Information Checklists have emerged within the Life Sciences as a means of standardising the reporting of experiments in an effort to increase the quality and reusability of the reported data. Existing tooling built around these checklists is aimed at supporting experimental scientists in the production of experiment reports that are compliant. It remains a challenge to quickly and easily assess an arbitrary set of data against these checklists. We present the MIM (Minimum Information Model) vocabulary and framework which aims to provide a practical, and scalable approach to describing and assessing Linked Data against minimum information checklists. The MIM framework aims to support three core activities: (1) publishing well described minimum information checklists in RDF as Linked Data; (2) publishing Linked Data against these checklists; and (3) validating existing “in the wild” Linked Data against a published checklist. We discuss the design considerations of the vocabulary and present its main classes. We demonstrate the utility of the framework with a checklist designed for the publishing of Chemical Structure Linked Data using data extracted from Wikipedia as an example.
Matthew Gamble, Carole A. Goble, Graham Klyne, Jun Zhao 0003
eScience4
2012 Why workflows break - Understanding and combating decay in Taverna workflows
abstract
Workflows provide a popular means for preserving scientific methods by explicitly encoding their process. However, some of them are subject to a decay in their ability to be re-executed or reproduce the same results over time, largely due to the volatility of the resources required for workflow executions. This paper provides an analysis of the root causes of workflow decay based on an empirical study of a collection of Taverna workflows from the myExperiment repository. Although our analysis was based on a specific type of workflow, the outcomes and methodology should be applicable to workflows from other systems, at least those whose executions also rely largely on accessing third-party resources. Based on our understanding about decay we recommend a minimal set of auxiliary resources to be preserved together with the workflows as an aggregation object and provide a software tool for end-users to create such aggregations and to assess their completeness.
Jun Zhao 0003, José Manuél Gómez-Pérez, Khalid Belhajjame, Graham Klyne, Esteban García-Cuesta, Aleix Garrido, Kristina M. Hettne, Marco Roos, David De Roure, Carole A. Goble
eScience1
2012 Translating standards into practice - One Semantic Web API for Gene Expression
Helena F. Deus, Eric Prud'hommeaux, Michael Miller 0001, Jun Zhao 0003, James Malone, Tomasz Adamusiak, Jamie P. McCusker, Sudeshna Das 0001, Philippe Rocca-Serra, Ronan Fox, M. Scott Marshall
J. Biomed. Informatics4
2012 Emerging practices for mapping and linking life sciences data using RDF - A case series
abstract
Members of the W3C Health Care and Life Sciences Interest Group (HCLS IG) have published a variety of genomic and drug-related data sets as Resource Description Framework (RDF) triples. This experience has helped the interest group define a general data workflow for mapping health care and life science (HCLS) data to RDF and linking it with other Linked Data sources. This paper presents the workflow along with four case studies that demonstrate the workflow and addresses many of the challenges that may be faced when creating new Linked Data resources. The first case study describes the creation of linked RDF data from microarray data sets while the second discusses a linked RDF data set created from a knowledge base of drug therapies and drug targets. The third case study describes the creation of an RDF index of biomedical concepts present in unstructured clinical reports and how this index was linked to a drug side-effect knowledge base. The final case study describes the initial development of a linked data set from a knowledge base of small molecules. This paper also provides a detailed set of recommended practices for creating and publishing Linked Data sources in the HCLS domain in such a way that they are discoverable and usable by people, software agents, and applications. These practices are based on the cumulative experience of the Linked Open Drug Data (LODD) task force of the HCLS IG. While no single set of recommendations can address all of the heterogeneous information needs that exist within the HCLS domains, practitioners wishing to create Linked Data should find the recommendations useful for identifying the tools, techniques, and practices employed by earlier developers. In addition to clarifying available methods for producing Linked Data, the recommendations for metadata should also make the discovery and consumption of Linked Data easier.
M. Scott Marshall, Richard D. Boyce, Helena F. Deus, Jun Zhao 0003, Egon L. Willighagen, Matthias Samwald, Elgar Pichler, Janos G. Hajagos, Eric Prud'hommeaux, Susie Stephens
J. Web Semant.4
2010 The Evolution of myExperiment
abstract
The myExperiment social website for sharing scientific workflows, designed according to Web 2.0 principles, has grown to be the largest public repository of its kind. It is distinctive for its focus on sharing methods, its researcher-centric design and its facility to aggregate content into sharable `research objects'. This evolution of myExperiment has occurred hand in hand with its users. myExperiment now supports Linked Data as a step toward our vision of the future research environment, which we categorise here as 3rd generation e-Research.
David De Roure, Carole A. Goble, Sergejs Aleksejevs, Sean Bechhofer, Jiten Bhagat, Don Cruickshank, Paul Fisher, Nandkumar Kollara, Danius T. Michaelides, Paolo Missier, David R. Newman, Marcus Ramsden, Marco Roos, Katy Wolstencroft, Ed Zaluska, Jun Zhao 0003
eScience16
2010 A Linked Data Approach to Sharing Workflows and Workflow Results
Marco Roos, Sean Bechhofer, Jun Zhao 0003, Paolo Missier, David R. Newman, David De Roure, M. Scott Marshall
ISoLA (1)3
2010 OpenFlyData: An exemplar data web integrating gene expression data on the fruit fly Drosophila melanogaster
Alistair J. Miles, Jun Zhao 0003, Graham Klyne, Helen White-Cooper, David M. Shotton
J. Biomed. Informatics2
2009 Linked data and provenance in biological data webs
abstract
The Web is now being used as a platform for publishing and linking life science data. The Web's linking architecture can be exploited to join heterogeneous data from multiple sources. However, as data are frequently being updated in a decentralized environment, provenance information becomes critical to providing reliable and trustworthy services to scientists. This article presents design patterns for representing and querying provenance information relating to mapping links between heterogeneous data from sources in the domain of functional genomics. We illustrate the use of named resource description framework (RDF) graphs at different levels of granularity to make provenance assertions about linked data, and demonstrate that these assertions are sufficient to support requirements including data currency, integrity, evidential support and historical queries.
Jun Zhao 0003, Alistair J. Miles, Graham Klyne, David M. Shotton
Briefings Bioinform.1
2009 A journey to Semantic Web query federation in the life sciences
abstract
BACKGROUND: As interest in adopting the Semantic Web in the biomedical domain continues to grow, Semantic Web technology has been evolving and maturing. A variety of technological approaches including triplestore technologies, SPARQL endpoints, Linked Data, and Vocabulary of Interlinked Datasets have emerged in recent years. In addition to the data warehouse construction, these technological approaches can be used to support dynamic query federation. As a community effort, the BioRDF task force, within the Semantic Web for Health Care and Life Sciences Interest Group, is exploring how these emerging approaches can be utilized to execute distributed queries across different neuroscience data sources. METHODS AND RESULTS: We have created two health care and life science knowledge bases. We have explored a variety of Semantic Web approaches to describe, map, and dynamically query multiple datasets. We have demonstrated several federation approaches that integrate diverse types of information about neurons and receptors that play an important role in basic, clinical, and translational neuroscience research. Particularly, we have created a prototype receptor explorer which uses OWL mappings to provide an integrated list of receptors and executes individual queries against different SPARQL endpoints. We have also employed the AIDA Toolkit, which is directed at groups of knowledge workers who cooperatively search, annotate, interpret, and enrich large collections of heterogeneous documents from diverse locations. We have explored a tool called "FeDeRate", which enables a global SPARQL query to be decomposed into subqueries against the remote databases offering either SPARQL or SQL query interfaces. Finally, we have explored how to use the vocabulary of interlinked Datasets (voiD) to create metadata for describing datasets exposed as Linked Data URIs or SPARQL endpoints. CONCLUSION: We have demonstrated the use of a set of novel and state-of-the-art Semantic Web technologies in support of a neuroscience query federation scenario. We have identified both the strengths and weaknesses of these technologies. While Semantic Web offers a global data model including the use of Uniform Resource Identifiers (URI's), the proliferation of semantically-equivalent URI's hinders large scale data integration. Our work helps direct research and tool development, which will be of benefit to this community.
Kei-Hoi Cheung, H. Robert Frost, M. Scott Marshall, Eric Prud'hommeaux, Matthias Samwald, Jun Zhao 0003, Adrian Paschke
BMC Bioinform.6
2008 Building a Semantic Web Image Repository for Biological Research Images
abstract
Images play a vital role in scientific studies. An image repository would become a costly and meaningless data graveyard without descriptive metadata. We adapted EPrints, a conventional repository software system, to create a biological research image repository for a local research group, in order to publish images with structured metadata with a minimum of development effort. However, in its native installation, this repository cannot easily be linked with information from third parties, and the user interface has limited flexibility. We address these two limitations by providing Semantic Web access to the contents of this image repository, causing the image metadata to become programmatically accessible through a SPARQL endpoint and enabling the images and their metadata to be presented in more flexible faceted browsers, jSpace and Exhibit. We show the feasibility of publishing image metadata on the Semantic Web using existing tools, and examine the inadequacies of the Semantic Web browsers in providing effective user interfaces. We highlight the importance of a loosely coupled software framework that provides a lightweight solution and enables us to switch between alternative components.
Jun Zhao 0003, Graham Klyne, David M. Shotton
ESWC1
2008 Special Issue: The First Provenance Challenge
abstract
Abstract The first Provenance Challenge was set up in order to provide a forum for the community to understand the capabilities of different provenance systems and the expressiveness of their provenance representations. To this end, a functional magnetic resonance imaging workflow was defined, which participants had to either simulate or run in order to produce some provenance representation, from which a set of identified queries had to be implemented and executed. Sixteen teams responded to the challenge, and submitted their inputs. In this paper, we present the challenge workflow and queries, and summarize the participants' contributions. Copyright © 2007 John Wiley & Sons, Ltd.
Luc Moreau 0001, Bertram Ludäscher, Ilkay Altintas, Roger S. Barga, Shawn Bowers, Steven P. Callahan, George Chin, Ben Clifford, Shirley Cohen, Sarah Cohen Boulakia, Susan B. Davidson, Ewa Deelman, Luciano A. Digiampietri, Ian T. Foster, Juliana Freire, James Frew, Joe Futrelle, Tara Gibson, Yolanda Gil, Carole A. Goble, Jennifer Golbeck, Paul Groth, David A. Holland, Jihie Kim, David Koop, Ales Krenek, Timothy M. McPhillips, Gaurang Mehta, Simon Miles, Dominic Metzger, Steve Munroe, James D. Myers, Beth Plale, Norbert Podhorszki, Varun Ratnakar, Emanuele Santos, Carlos Scheidegger, Karen Schuchardt, Margo I. Seltzer, Yogesh L. Simmhan, Cláudio T. Silva, Peter Slaughter, Eric G. Stephan, Robert Stevens 0001, Daniele Turi, Huy T. Vo, Michael Wilde, Jun Zhao 0003, Yong Zhao 0009
Concurr. Comput. Pract. Exp.49
2008 Mining Taverna's semantic web of provenance
abstract
Abstract Taverna is a workflow workbench developed as part of the UK's myGrid project. Taverna's provenance model captures both internal provenance locally generated in Taverna and external provenance gathered from third‐party data providers. This model also supports overlaying secondary provenance over the primary logs and lineage. This design is motivated by the particular properties of bioinformatics data and services used in Taverna. A Semantic Web of provenance, Ouzo, is built to combine the above different provenance by means of semantic annotations. This paper shows how Ouzo can be mined by a provenance usage component, Provenance Query and Answer (ProQA). ProQA supports provenance retrievals as well as provenance abstraction, aggregation, and semantic reasoning. ProQA is implemented as a suite APIs which can be deployed as provenance services to compose system provenance workflows that analyse experiment results using the provenance records. We show how these features of Taverna's provenance support us in answering the questions from the provenance challenge workshop and a set of additional provenance queries. Copyright © 2007 John Wiley & Sons, Ltd.
Jun Zhao 0003, Carole A. Goble, Robert Stevens 0001, Daniele Turi
Concurr. Comput. Pract. Exp.1
2007 Using provenance to manage knowledge of In Silico experiments
abstract
This article offers a briefing in one of the knowledge management issues of in silico experimentation in bioinformatics. Recording of the provenance of an experiment-what was done; where, how and why, etc. is an important aspect of scientific best practice that should be extended to in silico experimentation. We will do this in the context of eScience which has been part of the move of bioinformatics towards an industrial setting. Despite the computational nature of bioinformatics, these analyses are scientific and thus necessitate their own versions of typical scientific rigour. Just as recording who, what, why, when, where and how of an experiment is central to the scientific process in laboratory science, so it should be in silico science. The generation and recording of these aspects, or provenance, of an experiment are necessary knowledge management goals if we are to introduce scientific rigour into routine bioinformatics. In Silico experimental protocols should themselves be a form of managing the knowledge of how to perform bioinformatics analyses. Several systems now exist that offer support for the generation and collection of provenance information about how a particular in silico experiment was run, what results were generated, how they were generated, etc. In reviewing provenance support, we will review one of the important knowledge management issues in bioinformatics.
Robert Stevens 0001, Jun Zhao 0003, Carole A. Goble
Briefings Bioinform.2
2004 Using Semantic Web Technologies for Representing E-science Provenance
Jun Zhao 0003, Chris Wroe, Carole A. Goble, Robert Stevens 0001, Dennis Quan, Robert Mark Greenwood
ISWC1