Wenchuan Mu

dblp:319/2941 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
6since 2021 · last 2025
0009-0007-2395-9731ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Get Global Guarantees: On the Probabilistic Nature of Perturbation Robustness
abstract
Robustness is a critical requirement for deploying machine learning models in safety-sensitive domains, where even imperceptible input perturbations can lead to hazardous outcomes. However, existing robustness assessment techniques prior to deployment often face a trade-off between computational feasibility and measurement precision, limiting their effectiveness in practice. To address these limitations, we provide a systematic comparative study of prevailing robustness definitions and their corresponding evaluation methodologies. Building on this analysis, we propose tower robustness, which is a novel and practical concept setting out from a global perspective. Further, we provide upper and lower bounds of tower robustness, based on hypothesis testing, for quantitative evaluation, enabling more rigorous and efficient pre-deployment assessments. Through empirical investigation, we demonstrate that our approach provides reliable robustness assessments. These findings advance the systematic understanding of robustness and contribute a practical framework for enhancing the safety of machine learning models in safety-critical applications.
Wenchuan Mu, Kwan Hui Lim 0001
CIKM1
2025 Bayesian Privacy Guarantee for User History in Sequential Recommendation Using Randomised Response
abstract
Sequential recommendation systems play an important role in delivering personalised user experiences, yet they rely heavily on detailed user history, raising serious privacy concerns. In this work, we introduce a novel framework that integrates a randomised response mechanism into sequential recommendation to provide strong privacy guarantees while preserving recommendation effectiveness. By obfuscating user history through controlled probabilistic item substitution based on semantic similarity, our approach ensures that released sequences protect individual behaviour with provable Bayesian posterior privacy. We further propose training strategies tailored for privacy-filtered data, including a frequency-based vocabulary expansion method inspired by subword tokenisation. Experiments on four real-world datasets demonstrate that our approach preserves recommendation quality under strong privacy constraints and outperforms existing baselines even without applying privacy filters.
Wenchuan Mu, Kwan Hui Lim 0001
CIKM1
2025 Data-Free Functional Projection of Large Language Models onto Social Media Tagging Domain
Wenchuan Mu, Kwan Hui Lim 0001
MMM (1)1
2024 Label-Free Topic-Focused Summarization Using Query Augmentation
abstract
In today’s data and information-rich world, summarization techniques are essential in harnessing vast text to extract key information and enhance decision-making and efficiency. In particular, topic-focused summarization is important due to its ability to tailor content to specific aspects of an extended text. However, this usually requires extensive labelled datasets and considerable computational power. This study introduces a novel method, Augmented-Query Summarization (AQS), for topic-focused summarization without the need for extensive labelled datasets, leveraging query augmentation and hierarchical clustering. This approach facilitates the transferability of machine learning models to the task of summarization, circumventing the need for topic-specific training. Through real-world tests, our method demonstrates the ability to generate relevant and accurate summaries, showing its potential as a cost-effective solution in data-rich environments. This innovation paves the way for broader application and accessibility in the field of topic-focused summarization technology, offering a scalable, efficient method for personalized content extraction.
Wenchuan Mu, Kwan Hui Lim 0001
IJCNN1
2023 Modelling Text Similarity: A Survey
abstract
Online social networking services such as Twitter and Instagram have become pervasive platforms for engaging in discussions on a wide array of topics. These platforms cater to both mainstream subjects, like music and movies, as well as more specialized areas, such as politics. With the growing volume of textual data generated on these platforms, the ability to define and identify similar texts becomes crucial for effective investigation and clustering. In this paper, we explore the challenges and significance of text similarity regression models in the context of online social networking services. We delve into the methods and techniques employed to define and find similarities among texts, enabling the extraction of meaningful patterns and insights. Specifically, we categorize text similarity regression models into four distinct types: set-theoretic, sequence-theoretic, real-vector, and end-to-end methods. This categorization is based on the mathematical formalisation of similarity used by each model. Ultimately, our survey aims to provide a comprehensive overview of the interlinkages between independently proposed methods for text similarity. By understanding the strengths and weaknesses of these methods, researchers can make informed decisions when designing novel approaches and algorithms. We hope this survey serves as a valuable resource for advancing the state-of-the-art in addressing the complex problem of text similarity.
Wenchuan Mu, Kwan Hui Lim 0001
ASONAM1
2023 Photozilla: An Image Dataset of Photography Styles and its Application to Visual Embedding and Style Detection
abstract
The widespread sharing of digital photography and images have led to the rapid development of various vision-related applications, such as photography style detection. Towards this effort, we introduce a photography style dataset termed Photozilla, which comprises over 990k images belonging to 10 different photographic styles. We used Photozilla to train 3 classification models for categorizing images into the relevant style and achieve an accuracy of ~96%. To better detect new photography styles that are constantly emerging, we also present a Siamese-based network that uses the trained classification models as the base architecture to adapt and classify unseen styles with only 25 training samples. Experiment results show an accuracy of over 68% in terms of identifying 10 additional distinct categories of photography styles. This dataset can be found at https://trisha025.github.io/Photozilla/.
Trisha Singhal, Junhua Liu 0002, Wenchuan Mu, Luciënne T. M. Blessing, Kwan Hui Lim 0001
ASONAM3