Xihui Chen

dblp:49/7565 · DBLP profile ↗
← Back
23ranked-venue papers
13as first author
11since 2021 · last 2026
0000-0002-8131-5092ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 4 first-author · 7 since 2021Security and privacy · 9 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Data augmentation for intelligent fault diagnosis based on feature pre-extraction mechanism and improved conditional denoising diffusion probabilistic models under data scarcity
Xihui Chen, Leyan Fan, Zihao Xing, Fengtao Wang
Eng. Appl. Artif. Intell.1
2026 Distilling knowledge from large language models: A concept bottleneck model for hate and counter speech recognition
abstract
The rapid increase in hate speech on social media has exposed an unprecedented impact on society, making automated methods for detecting such content important. Unlike prior black-box models, we propose a novel transparent method for automated hate and counter speech recognition, i.e., “Speech Concept Bottleneck Model” (SCBM), using adjectives as human-interpretable bottleneck concepts. SCBM leverages large language models (LLMs) to map input texts to an abstract adjective-based representation, which is then sent to a light-weight classifier for downstream tasks. Across five benchmark datasets spanning multiple languages and platforms (e.g., Twitter, Reddit, YouTube), SCBM achieves an average macro-F1 score of 0.69 which outperforms the most recently reported results from the literature on four out of five datasets. Aside from high recognition accuracy, SCBM provides a high level of both local and global interpretability. Furthermore, fusing our adjective-based concept representation with transformer embeddings, leads to a 1.8% performance increase on average across all datasets, showing that the proposed representation captures complementary information. Our results demonstrate that adjective-based concept representations can serve as compact, interpretable, and effective encodings for hate and counter speech recognition. With adapted adjectives, our method can also be applied to other NLP tasks.
Roberto Labadie, Djordje Slijepcevic, Xihui Chen, Adrian Jaques Böck, Andreas Babic, Liz Freimann, Christiane Atzmüller, Matthias Zeppelzauer
Inf. Process. Manag.3
2025 CounterHelp: Promoting Online Civil Courage Among Young People Through AI-Generated Counterspeech
abstract
We present CounterHelp, a mobile app designed to support adolescents and young people in generating customized counter speech in response to hateful comments on TikTok. CounterHelp allows users to specify counter strategies to combat hate speech. After users share a hateful comment with CounterHelp, it retrieves TikTok metadata to capture its context. Leveraging large language models, CounterHelp generates customised and context-sensitive counter speech in one of four predefined styles, particularly tailored for adolescents. A user experience lab study confirms the effectiveness and usability of CounterHelp and the generated counter speech.
Andreas Babic, Xihui Chen, Djordje Slijepcevic, Adrian Jaques Böck, Matthias Zeppelzauer
ACM Multimedia2
2025 "Double vaccinated, 5G boosted!": Learning Attitudes towards COVID-19 Vaccination from Social Media
abstract
The sudden onset of the recently concluded COVID-19 pandemic has driven substantial progress in various scientific fields. One notable example is the comprehension of public vaccination attitudes and the timely monitoring of their fluctuations through social media platforms. This approach can serve as a cost-effective means to supplement surveys in gathering public vaccine hesitancy levels. In this article, we propose a deep learning framework leveraging textual posts on social media to extract and track users’ vaccination stances in near real time. Compared to previous works, we integrate into the framework the recent posts of a user’s social network friends to collaboratively detect the user’s genuine attitude towards vaccination. Based on our annotated dataset from X (formerly known as Twitter), the models instantiated from our framework can increase the performance of attitude extraction by up to 23% compared to the state-of-the-art text-only models. Using this framework, we successfully confirm the feasibility of using social media to track the evolution of vaccination attitudes in real life. In addition, we illustrate the generality of our framework in extracting other public opinions such as political ideology. We further show one practical use of our framework by validating the possibility of forecasting a user’s vaccine hesitancy changes with information perceived from social media.
Ninghan Chen, Xihui Chen, Zhiqiang Zhong 0001, Jun Pang 0001
ACM Trans. Web2
2024 Not One Less: Exploring Interplay between User Profiles and Items in Untargeted Attacks against Federated Recommendation
abstract
Federated recommendation (FR) is a decentralised approach to training personalised recommender systems, protecting users' privacy by avoiding data collection. Despite its privacy advantages, FR remains vulnerable to poisoning attacks. We focus on untargeted poisoning attacks against FR which degrade the overall performance of recommender services, leading to a detrimental impact on user experience and service quality. In this paper, we propose a general framework to formalise untargeted attacks and identify the vital role played by the interplay between items and user profiles in determining FR's performance. We present an untargeted attack FRecAttack2 which exploits this interplay. Specifically, we develop various methods for sampling user profiles, which approximate user distributions with and without collusion among malicious users. Then we leverage a new measurement to identify items that can disrupt the original interplay with user profiles, based on the change velocity of items' recommendation scores during optimisation. Extensive experiments demonstrate the superiority of our attack, outperforming existing methods by up to 27.56%, and its stealthiness in evading mainstream defences. To counteract untargeted attacks, we present a defence GuardCQ to detect malicious users by quantifying their contribution to boost the right interplay between items and user profiles. Empirical results show that GuardCQ effectively mitigates the attack's impact on FR and enhances the robustness of FR against poisoning attacks.
Yurong Hao, Xihui Chen, Xiaoting Lyu, Jiqiang Liu, Yongsheng Zhu, Zhiguo Wan, Sjouke Mauw, Wei Wang 0012
CCS2
2024 A tale of two roles: exploring topic-specific susceptibility and influence in cascade prediction
abstract
Abstract We propose a new deep learning cascade prediction model CasSIM that can simultaneously achieve two most demanded objectives: popularity prediction and final adopter prediction. Compared to existing methods based on cascade representation, CasSIM simulates information diffusion processes by exploring users’ dual roles in information propagation with three basic factors: users’ susceptibilities, influences and message contents. With effective user profiling, we are the first to capture the topic-specific property of susceptibilities and influences. In addition, the use of graph neural networks allows CasSIM to capture the dynamics of susceptibilities and influences during information diffusion. We evaluate the effectiveness of CasSIM on three real-life datasets and the results show that CasSIM outperforms the state-of-the-art methods in popularity and final adopter prediction.
Ninghan Chen, Xihui Chen, Zhiqiang Zhong 0001, Jun Pang 0001
Data Min. Knowl. Discov.2
2024 Eyes on Federated Recommendation: Targeted Poisoning With Competition and Its Mitigation
abstract
Federated recommendation (FR) addresses privacy concerns in recommender systems by training a global model without requiring raw user data to leave individual devices. A server, known as the aggregator, integrates users’ local gradients and updates the global model parameters. However, FR is vulnerable to attacks where malicious users manipulate these updates, known as model poisoning attacks. In this work, we propose a new targeted attack calledStairClimbingto promote specific items through model poisoning, and a new defence mechanismCrossEU. StairClimbingadopts a new strategy resembling stair climbing to enable target items to beat competitive items and increase their popularity level by level. Compared to prior attacks,StairClimbingguarantees balanced effectiveness, efficiency and stealthiness simultaneously. Our defence mechanismCrossEUleverages two patterns regarding the lists of items updated by benign users between iterative epochs. Extensive experiments on six real-world datasets demonstrateStairClimbing’s superiority across all three desirable attack properties, even with a small proportion of malicious users (1%). In addition,CrossEUeffectively delays the impact of all tested attacks and even eliminates their damage entirely.
Yurong Hao, Xihui Chen, Wei Wang 0012, Jiqiang Liu, Tao Li 0022, Witold Pedrycz
IEEE Trans. Inf. Forensics Secur.2
2024 Bridging Performance of X (formerly known as Twitter) Users: A Predictor of Subjective Well-Being During the Pandemic
abstract
The outbreak of the COVID-19 pandemic triggered the perils of misinformation over social media. By amplifying the spreading speed and popularity of trustworthy information, influential social media users have been helping overcome the negative impacts of such flooding misinformation. In this article, we use the COVID-19 pandemic as a representative global health crisisand examine the impact of the COVID-19 pandemic on these influential users’ subjective well-being (SWB), one of the most important indicators of mental health. We leverage X (formerly known as Twitter) as a representative social media platform and conduct the analysis with our collection of 37,281,824 tweets spanning almost two years. To identify influential X users, we propose a new measurement called user bridging performance (UBM) to evaluate the speed and wideness gain of information transmission due to their sharing. With our tweet collection, we manage to reveal the more significant mental sufferings of influential users during the COVID-19 pandemic. According to this observation, through comprehensive hierarchical multiple regression analysis , we are the first to discover the strong relationship between individual social users’ subjective well-being and their bridging performance. We proceed to extend bridging performance from individuals to user subgroups. The new measurement allows us to conduct a subgroup analysis according to users’ multilingualism and confirm the bridging role of multilingual users in the COVID-19 information propagation. We also find that multilingual users not only suffer from a much lower SWB in the pandemic, but also experienced a more significant SWB drop.
Ninghan Chen, Xihui Chen, Zhiqiang Zhong 0001, Jun Pang 0001
ACM Trans. Web2
2022 The Burden of Being a Bridge: Analysing Subjective Well-Being of Twitter Users During the COVID-19 Pandemic
Ninghan Chen, Xihui Chen, Zhiqiang Zhong 0001, Jun Pang 0001
ECML/PKDD (2)2
2021 From #jobsearch to #mask: improving COVID-19 cascade prediction with spillover effects
abstract
An information outbreak occurs on social media along with the COVID-19 pandemic and leads to infodemic. Predicting the popularity of online content, known as cascade prediction, allows for not only catching in advance hot information that deserves attention, but also identifying false information that will widely spread and require quick response to mitigate its impact. Among the various information diffusion patterns leveraged in previous works, the spillover effect of the information exposed to users on their decision to participate in diffusing certain information is still not studied. In this paper, we focus on the diffusion of information related to COVID-19 preventive measures. Through our collected Twitter dataset, we validated the existence of this spillover effect. Building on the finding, we proposed extensions to three cascade prediction methods based on Graph Neural Networks (GNNs). Experiments conducted on our dataset demonstrated that the use of the identified spillover effect significantly improves the state-of-the-art GNNs methods in predicting the popularity of not only preventive measure messages, but also other COVID-19 related messages.
Ninghan Chen, Xihui Chen, Zhiqiang Zhong 0001, Jun Pang 0001
ASONAM2
2021 Pattern Recognition and Reconstruction: Detecting Malicious Deletions in Textual Communications
abstract
Digital forensic artifacts aim to provide evidence from digital sources for attributing blame to suspects, assessing their intents, corroborating their statements or alibis, etc. Textual data is a significant source of artifacts, which can take various forms, for instance in the form of communications. E-mails, memos, tweets, and text messages are all examples of textual communications. Complex statistical, linguistic and other scientific procedures can be manually applied to this data to uncover significant clues that point the way to factual information. While expert investigators can undertake this task, there is a possibility that critical information is missed or overlooked. The primary objective of this work is to aid investigators by partially automating the detection of suspicious e-mail deletions. Our approach consists in building a dynamic graph to represent the temporal evolution of communications, and then using a Variational Graph Autoencoder to detect possible e-mail deletions in this graph. Our model uses multiple types of features for representing node and edge attributes, some of which are based on metadata of the messages and the rest are extracted from the contents using natural language processing and text mining techniques. We use the autoencoder to detect missing edges, which we interpret as potential deletions; and to reconstruct their features, from which we emit hypotheses about the topics of deleted messages. We conducted an empirical evaluation of our model on the Enron e-mail dataset, which shows that our model is able to accurately detect a significant proportion of missing communications and to reconstruct the corresponding topic vectors.
Abiodun A. Solanke, Xihui Chen, Yunior Ramírez-Cruz
IEEE BigData2
2020 Active Re-identification Attacks on Periodically Released Dynamic Social Graphs
Xihui Chen, Ema Këpuska, Sjouke Mauw, Yunior Ramírez-Cruz
ESORICS (2)1
2020 Publishing Community-Preserving Attributed Social Graphs with a Differential Privacy Guarantee
abstract
Abstract We present a novel method for publishing differentially private synthetic attributed graphs. Our method allows, for the first time, to publish synthetic graphs simultaneously preserving structural properties, user attributes and the community structure of the original graph. Our proposal relies on CAGM, a new community-preserving generative model for attributed graphs. We equip CAGM with efficient methods for attributed graph sampling and parameter estimation. For the latter, we introduce differentially private computation methods, which allow us to release communitypreserving synthetic attributed social graphs with a strong formal privacy guarantee. Through comprehensive experiments, we show that our new model outperforms its most relevant counterparts in synthesising differentially private attributed social graphs that preserve the community structure of the original graph, as well as degree sequences and clustering coefficients.
Xihui Chen, Sjouke Mauw, Yunior Ramírez-Cruz
Proc. Priv. Enhancing Technol.1
2014 Measuring User Similarity with Trajectory Patterns: Principles and New Metrics
Xihui Chen, Ruipeng Lu, Xiaoxing Ma, Jun Pang 0001
APWeb1
2014 MinUS: Mining User Similarity with Trajectory Patterns
Xihui Chen, Piotr Kordy, Ruipeng Lu, Jun Pang 0001
ECML/PKDD (3)1
2014 Protecting query privacy in location-based services
Xihui Chen, Jun Pang 0001
GeoInformatica1
2014 Constructing and Comparing User Mobility Profiles
abstract
Nowadays, the accumulation of people's whereabouts due to location-based applications has made it possible to construct their mobility profiles. This access to users' mobility profiles subsequently brings benefits back to location-based applications. For instance, in on-line social networks, friends can be recommended not only based on the similarity between their registered information, for instance, hobbies and professions but also referring to the similarity between their mobility profiles. In this article, we propose a new approach to construct and compare users' mobility profiles. First, we improve and apply frequent sequential pattern mining technologies to extract the sequences of places that a user frequently visits and use them to model his mobility profile. Second, we present a new method to calculate the similarity between two users using their mobility profiles. More specifically, we identify the weaknesses of a similarity metric in the literature, and propose a new one which not only fixes the weaknesses but also provides more precise and effective similarity estimation. Third, we consider the semantics of spatio-temporal information contained in user mobility profiles and add them into the calculation of user similarity. It enables us to measure users' similarity from different perspectives. Two specific types of semantics are explored in this article: location semantics and temporal semantics . Last, we validate our approach by applying it to two real-life datasets collected by Microsoft Research Asia and Yonsei University, respectively. The results show that our approach outperforms the existing works from several aspects.
Xihui Chen, Jun Pang 0001, Ran Xue
ACM Trans. Web1
2013 Demonstrating a trust framework for evaluating GNSS signal integrity
abstract
Through real-life experiments, it has been proved that spoofing is a practical threat to applications using the free civil service provided by Global Navigation Satellite Systems (GNSS). In this paper, we demonstrate a prototype that can verify the integrity of GNSS civil signals. By integrity we intuitively mean that civil signals originate from a GNSS satellite without having been artificially interfered with. Our prototype provides interfaces that can incorporate existing spoofing detection methods whose results are then combined into an overall evaluation of the signal's integrity, which we call integrity level. Considering the various security requirements from different applications, integrity levels can be calculated in many ways determined by their users. We also present an application scenario that deploys our prototype and offers a public central service -- localisation assurance certification. Through experiments, we successfully show that our prototype is not only effective but also efficient in practice.
Xihui Chen, Carlo Harpes, Gabriele Lenzini, Miguel Martins, Sjouke Mauw, Jun Pang 0001
CCS1
2013 Exploring dependency for query privacy protection in location-based services
abstract
Location-based services have been enduring a fast development for almost fifteen years. Due to the lack of proper privacy protection, especially in the early stage of the development, an enormous amount of user request records have been collected. This exposes potential threats to users' privacy as new contextual information can be extracted from such records. In this paper, we study query dependency which can be derived from users' request history, and investigate its impact on users' query privacy. To achieve our goal, we present an approach to compute the probability for a user to issue a query, by taking into account both user's query dependency and observed requests. We propose new metrics incorporating query dependency for query privacy, and adapt spatial generalisation algorithms in the literature to generate requests satisfying users' privacy requirements expressed in the new metrics. Through experiments, we evaluate the impact of query dependency on query privacy and show that our proposed metrics and algorithms are effective and efficient for practical applications.
Xihui Chen, Jun Pang 0001
CODASPY1
2013 A Trust Framework for Evaluating GNSS Signal Integrity
abstract
Through real-life experiments, it has been proved, not only in theory but also in practice, that civil signals of Global Navigation Satellite Systems (GNSS) can be spoofed. Consequently, a number of spoofing detection techniques have been proposed to verify the integrity of GNSS signals. In this paper, we develop a novel trust framework based on subjective logic to evaluate the integrity of received GNSS civil signals. We formally define signal integrity for the first time in the framework and use it to precisely characterise different spoofing detection methods. Our framework captures the uncertainty during the inference of signal integrity which has been largely ignored or not explicitly specified in the literature. Our framework also gives rise to several natural ways to combine the outputs of various spoofing detection methods on signal integrity. We validate our framework through experiments using both real and simulated signals and the results show that our framework is effective.
Xihui Chen, Gabriele Lenzini, Miguel Martins, Sjouke Mauw, Jun Pang 0001
CSF1
2012 A Group Signature Based Electronic Toll Pricing System
abstract
With the prevalence of GNSS technologies, nowadays freely available for everyone, location-based vehicle services such as electronic tolling pricing systems and pay-as-you-drive services are rapidly growing. Because these systems collect and process travel records, if not carefully designed, they can threaten users' location privacy. Finding a secure and privacy-friendlysolution is a challenge for system designers. Besides location privacy, communication and computation overhead should be taken into account as well in order to make such systems widely adopted in practice. In this paper, we propose a new electronic toll pricing system based on group signatures. Our system preserves anonymity of users within groups, in addition to correctness and accountability. It also achieves a balance between privacy and overhead imposed upon user devices.
Xihui Chen, Gabriele Lenzini, Sjouke Mauw, Jun Pang 0001
ARES1
2012 Measuring query privacy in location-based services
abstract
The popularity of location-based services leads to serious concerns on user privacy. A common mechanism to protect users' location and query privacy is spatial generalisation. As more user information becomes available with the fast growth of Internet applications, e.g., social networks, attackers have the ability to construct users' personal profiles. This gives rise to new challenges and reconsideration of the existing privacy metrics, such as k-anonymity. In this paper, we propose new metrics to measure users' query privacy taking into account user profiles. Furthermore, we design spatial generalisation algorithms to compute regions satisfying users' privacy requirements expressed in these metrics. By experimental results, our metrics and algorithms are shown to be effective and efficient for practical usage.
Xihui Chen, Jun Pang 0001
CODASPY1
2009 Improving Automatic Verification of Security Protocols with XOR
Xihui Chen, Ton van Deursen, Jun Pang 0001
ICFEM1