Marc Cheong

dblp:16/7536 · also Marc C. Y. Cheong · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
6since 2021 · last 2025
0000-0002-0637-3436ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A Whole New World: Migrant Journeys Through Digital Information Landscapes
abstract
Migrants moving across national boundaries are experiencing a transition—one that requires them to readapt and reestablish themselves in new information environments. In experiencing a transition, migrants need information on the process of the move and on life in their new country. Navigating new information environments can prove difficult. Usually, that information will be structured and presented differently than ‘at home’: migrants need to adapt to a whole new digital environment, and often do so while retaining informational ties to home. Language, bureaucratic, and emotional barriers make it difficult for migrants to find the information they need in their new environment; and instead push them to rely on familiar information sources from their home countries that contain less specific and relevant information. However, migrants also bring their unique information experiences and habits to their new digital environments. The exposure to multiple digital environments can be a strength and better equip migrants to deal with different types of information. This paper examines both the barriers migrants face and the skills they use to navigate new information landscapes. It makes a case for considering migrants as a specific user group for many information tools and resources, when user testing, and makes suggestions for further research that can be done to support and learn from migrants. Migrants are an under-studied and under-considered group in the CHIIR community, and this paper is a call to action to address that gap.
Rashika Bahl, Shanton Chang, Dana McKay, George Buchanan 0001, Marc Cheong
CHIIR5
2024 Documenting Ethical Considerations in Open Source AI Models
abstract
Background: The development of AI-enabled software heavily depends on AI model documentation, such as model cards, due to different domain expertise between software engineers and model developers. From an ethical standpoint, AI model documentation conveys critical information on ethical considerations along with mitigation strategies for downstream developers to ensure the delivery of ethically compliant software. However, knowledge on such documentation practice remains scarce. Aims: The objective of our study is to investigate how developers document ethical aspects of open source AI models in practice, aiming at providing recommendations for future documentation endeavours. Method: We selected three sources of documentation on GitHub and Hugging Face, and developed a keyword set to identify ethics-related documents systematically. After filtering an initial set of 2,347 documents, we identified 265 relevant ones and performed thematic analysis to derive the themes of ethical considerations. Results: Six themes emerge, with the three largest ones being model behavioural risks, model use cases, and model risk mitigation. Conclusions: Our findings reveal that open source AI model documentation focuses on articulating ethical problem statements and use case restrictions. We further provide suggestions to various stakeholders for improving documentation practice regarding ethical considerations.
Mansooreh Zahedi, Christoph Treude, Sarita Rosenstock, Marc Cheong
ESEM5
2024 Nigerian Software Engineer or American Data Scientist? GitHub Profile Recruitment Bias in Large Language Models
abstract
Large Language Models (LLMs) have taken the world by storm, demonstrating their ability not only to automate tedious tasks, but also to show some degree of proficiency in completing software engineering tasks. A key concern with LLMs is their “black-box” nature, which obscures their internal workings and could lead to societal biases in their outputs. In the software engineering context, in this early results paper, we empirically explore how well LLMs can automate recruitment tasks for a geographically diverse software team. We use OpenAI's ChatGPT to conduct an initial set of experiments using GitHub User Profiles from four regions to recruit a six-person software development team, analyzing a total of 3,657 profiles over a five-year period (2019–2023). Results indicate that ChatGPT shows preference for some regions over others, even when swapping the location strings of two profiles (counterfactuals). Furthermore, ChatGPT was more likely to assign certain developer roles to users from a specific country, revealing an implicit bias. Overall, this study reveals insights into the inner workings of LLMs and has implications for mitigating such societal biases in these models.
Takashi Nakano, Kazumasa Shimari, Raula Gaikovina Kula, Christoph Treude, Marc Cheong, Ken-ichi Matsumoto
ICSME5
2022 Social Media Harms as a Trilemma: Asymmetry, Algorithms, and Audacious Design Choices
abstract
Social media has expanded in its use, and reach, since the inception of early social networks in the early 2000s. Increasingly, users turn to social media for keeping up to date with current affairs and information. However, social media is increasingly used to promote disinformation and cause harm. In this contribution, I argue that as information (eco)systems, social media sites are vulnerable from three aspects, each corresponding to the classical 3-tier architecture in information systems: asymmetric networks (data tier); algorithms powering the supposed personalisation for the user experience (application tier); and adverse or audacious design of the user experience and overall information ecosystem (presentation tier) - which can be summarised as the 3 A’s. Thus, the open question remains: how can we ‘fix’ social media? I will unpack suggestions from various allied disciplines- from philosophy to data ethics to social psychology - in untangling the 3A’s above.
Marc Cheong
ISTAS1
2022 Gender Bias in AI Recruitment Systems: A Sociological-and Data Science-based Case Study
abstract
This paper explores the extent to which gender bias is introduced in the deployment of automation for hiring practices. We use an interdisciplinary methodology to test our hypotheses: observing a human-led recruitment panel and building an explainable algorithmic prototype from the ground up, to quantify gender bias. The key findings of this study are threefold: identifying potential sources of human bias from a recruitment panel’s ranking of CVs; identifying sources of bias from a potential algorithmic pipeline which simulates human decision making; and recommending ways to mitigate bias from both aspects. Our research has provided an innovative research design that combines social science and data science to theorise how automation may introduce bias in hiring practices, and also pinpoint where it is introduced. It also furthers the current scholarship on gender bias in hiring practices by providing key empirical inferences on the factors contributing to bias.
Sheilla Njoto, Marc Cheong, Reeva Lederman, Aidan McLoughney, Leah Ruppanner, Anthony Wirth
ISTAS2
2022 Using public data to measure diversity in computer science research communities: A critical data governance perspective
abstract
Encouraging and supporting diversity and inclusion in computer science research communities is a critical issue for many reasons, including the ethical and robust design, delivery and publication of research that addresses real-world situations ranging from the use of digital tools in health to predictive policing to workplace hiring practices, just to name a few. One way to measure diversity is to apply analytical research methods to data sourced from the public domain for use in research. However, attempts to measure diversity using public data may themselves raise legal and ethical questions about the provenance of the data, research methods adopted, and treatment of diversity in the publication of results. This article interrogates the challenges of measuring diversity using public data, examining an illustrative case study framed around an academic research project at an Australian university using a public data set to identify gender representation in computer science communities. Employing a critical data governance perspective, we point to a range of ethical and legal concerns and recommend greater regulatory guardrails to better balance public interests in research and the privacy, data protection and other ethical interests of research subjects.
Rachelle Bosua, Marc Cheong, Karin Clark, Damian Clifford, Simon Coghlan, Chris Culnane, Kobi Leins, Megan Richardson 0001
Comput. Law Secur. Rev.2
2020 To Honor our Heroes: Analysis of the Obituaries of Australians Killed in Action in WWI and WWII
abstract
Obituaries represent a prominent way of expressing the human universal of grief. According to philosophers, obituaries are a ritualized way of evaluating both individuals who have passed away and the communities that helped to shape them. The basic idea is that you can tell what it takes to count as a good person of a particular type in a particular community by seeing how persons of that type are described and celebrated in their obituaries. Obituaries of those killed in conflict, in particular, are rich repositories of communal values, as they reflect the values and virtues that are admired and respected in individuals who are considered to be heroes in their communities. In this paper, we use natural language processing techniques to map the patterns of values and virtues attributed to Australian military personnel who were killed in action during World War I and World War II. Doing so reveals several clusters of values and virtues that tend to be attributed together. In addition, we use named entity recognition and geotagging the track the movements of these soldiers to various theatres of the wars, including North Africa, Europe, and the Pacific.
Marc Cheong, Mark Alfano
ICPR1
2012 Large-scale socio-demographic pattern discovery on microblog metadata
abstract
Microblogging services, such as Twitter, generate huge volumes of data reflecting the current zeitgeist. As such they are of enormous potential value to studies ranging from data mining to social anthropology. To realize the potential, this study investigates improvements of algorithms specifically tailored for the discovery of latent socio-demographic patterns in Twitter metadata. These newly improved hybrid algorithms improve on existing ones in terms of speed and scalability (from thousands of records to millions). Testing on a real-world Twitter data set (~7.4 million messages) reveals emergent patterns in global day-to-day Twitter activity. The results demonstrate novel insight when applied to real-world Twitter data, practical large-scale applications of the methods, and suggest potential areas of future research.
Marc Cheong, Sid Ray, David G. Green
ISDA1
2012 Interpreting the 2011 London riots from twitter metadata
abstract
Social media have rapidly become one of the principal venues for personal and public communication. This makes them rich sources of information about real-world events. As a case study, we used Twitter metadata to investigate social dimensions of the 2011 London riots. The results showed that Twitter-based commentary and participation in the London Riots are closely linked to the real-world manifestation of the riots (e.g. in terms of geographic presence). Twitter metadata on users and their messages during the riots can be used to generate useful inferences which allows us to gain a better insight into intents, information-sharing behavior, and demographics of both the rioters and observers of the riots. Pattern recognition approaches can be used to further reveal latent properties from the acquired inferences.
Marc Cheong, Sid Ray, David G. Green
ISDA1
2010 Twittering for Earth: A Study on the Impact of Microblogging Activism on Earth Hour 2009 in Australia
Marc Cheong, Vincent Cheng-Siong Lee
ACIIDS (2)1
2010 A Study on Detecting Patterns in Twitter Intra-topic User and Message Clustering
abstract
Timely detection of hidden patterns is the key for the analysis and estimating of driving determinants for mission critical decision making. This study applies Cheong and Lee's “context-aware” content analysis framework to extract latent properties from Twitter messages (tweets). In addition, we incorporate an unsupervised Self-organizing Feature Map (SOM) as a machine learning-based clustering tool that has not been investigated in the context of opinion mining and sentimental analysis using microblogging. Our experimental results reveal the detection of interesting patterns for topics of interest which are latent and cannot be easily detected from the observed tweets without the aid of machine learning tools.
Marc Cheong, Vincent Cheng-Siong Lee
ICPR1
2008 Textile Recognition Using Tchebichef Moments of Co-occurrence Matrices
Marc Cheong, Kar-Seng Loke
ICIC (1)1
2008 An approach to texture-based image recognition by deconstructing multispectral co-occurrence matrices using Tchebichef orthogonal polynomials
abstract
The existing use of summary statistics from co-occurrence matrices of images for texture recognition and classification has inadequacies when dealing with non-uniform and colored texture such as traditional `Batik¿ and `Songket¿ cloth motifs. This study uses the Tchebichef orthogonal polynomial as a way to preserve the shape information of cooccurrence matrices generated using the RGB multispectral method; allowing prominent features and shapes of the matrices to be preserved while discarding extraneous information. The decomposition of the six multispectral co-occurrence matrices yields a set of moment coefficients which can be used to quantify the difference between textures. The proposed method, when tested with a subset of the Vision Texture (VisTex) database and a collection of `Batik¿ and `Songket¿ motifs, yielded promising results of 99.5% and 95.28% classification rates respectively using the 3-nearest neighbor classifier.
Marc Cheong, Kar-Seng Loke
ICPR1