EDBT 2026 Demo / reviewers in the wild / expert
Tingmin Wu
dblp:192/6021
· DBLP profile ↗
20ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0003-0626-3576ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 8 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SAFE: A Novel Approach For Software Vulnerability Detection from Enhancing The Capability of Large Language ModelsabstractSoftware vulnerabilities (SVs) have emerged as a prevalent and crucial concern for safety-critical systems. This has spurred significant advancements in utilizing AI-based methods, including machine learning and deep learning, for software vulnerability detection (SVD). While AI-based methods have shown promising performance in SVD, their effectiveness on real-world, complex, and diverse source code datasets remains limited in practice. To tackle this challenge, in this paper, we propose a novel framework that enhances the capability of large language models to learn and utilize semantic and syntactic relationships from source code data for SVD. As a result, our proposed SAFE approach can enable the acquisition of fundamental knowledge from source code data while adeptly utilizing crucial relationships, i.e., semantic and syntactic associations, to improve the effectiveness of solving the SVD problem. The rigorous and extensive experimental results on three real-world challenging datasets (i.e., Devign, ReVeal, and D2A) demonstrate the superiority of our approach over eight effective and state-of-the-art baselines. In summary, on average, our SAFE approach achieves higher performances from 4.79% to 11.57% for F1-measure and from 16.93% to 26.24% for Recall compared to the baseline methods across all the datasets used. Van Nguyen 0002, Surya Nepal, Xingliang Yuan, Tingmin Wu, Carsten Rudolph |
AsiaCCS | 4 |
| 2025 | AI2TALE: An Innovative Information Theory-based Approach for Learning to Localize Phishing AttacksabstractPhishing attacks remain a significant challenge for detection, explanation, and defense, despite over a decade of research on both technical and non-technical solutions. AI-based phishing detection methods are among the most effective approaches for defeating phishing attacks, providing predictions on the vulnerability label (i.e., phishing or benign) of data. However, they often lack intrinsic explainability, failing to identify the specific information that triggers the classification. To this end, we propose AI2TALE, an innovative deep learning-based approach for email (the most common phishing medium) phishing attack localization. Our method aims to not only predict the vulnerability label of the email data but also provide the capability to automatically learn and identify the most important and phishing-relevant information (i.e., sentences) in the phishing email data, offering useful and concise explanations for the identified vulnerability.
Extensive experiments on seven diverse real-world email datasets demonstrate the capability and effectiveness of our method in selecting crucial information, enabling accurate detection and offering useful and concise explanations (via the most important and phishing-relevant information triggering the classification) for the vulnerability of phishing emails. Notably, our approach outperforms state-of-the-art baselines by 1.5% to 3.5% on average in Label-Accuracy and Cognitive-True-Positive metrics under a weakly supervised setting, where only vulnerability labels are used without requiring ground truth phishing information. Van Nguyen 0002, Tingmin Wu, Xingliang Yuan, Marthie Grobler, Surya Nepal, Carsten Rudolph |
ICLR | 2 |
| 2025 | Semantics-Aware Cookie Purpose ComplianceabstractWebsites commonly display cookie banners to inform users about the use and purposes of cookies. However, they may still, whether intentionally or unintentionally (e.g., due to third-party libraries imported), mis-declare cookies that may be abused for tracking. In this work, we introduce COOVER (cookie value examiner) to assess the non-compliance between the website-declared purpose and the semantic-intended purpose of cookies (denoted as potential cookie purpose violation ). We advocate that the value of the cookie is a more reliable indicator of its semantic-intended purpose compared to other features such as expiration time. COOVER decomposes the cookie value into primitive segments representing minimal semantic units, and fine-tunes a GPT-3.5 model to automatically interpret their value-inferred semantics. Based on the interpretation, it classifies cookies into four GDPR-defined purposes. COOVER achieves an F1 score of 95%, significantly outperforming other methods. We employ COOVER to analyze Alexa Top 1k websites to understand the status quo of potential cookie purpose violation on the web. Remarkably, out of 15,339 cookies across these websites, only 3.1% quality as truly necessary cookies, while 44.1% of websites suffer from issues of potential purpose violation. Baiqi Chen, Jiawei Lyu, Tingmin Wu, Mohan Baruwal Chhetri, Guangdong Bai |
WWW | 3 |
| 2025 | CyberLLaMA: A fine-tuned large language model for cybersecurity named entity recognition
Hao Zhang 0181, Tingmin Wu, Tianqing Zhu, Sheng Wen, Yang Xiang 0001 |
Knowl. Based Syst. | 2 |
| 2024 | How COVID-19 impacts telehealth: an empirical study of telehealth services, users and the use of metaverseabstractSince the outbreak of the coronavirus 2019 (COVID-19) pandemic, telehealth services are regarded as a good approach to keep health workers and patients safe while simultaneously managing available resources.In this paper, we discuss the impact that COVID-19 has on telehealth services and on telehealth users' opinion of the service.We collected 245 Android telehealth apps, 144 iOS telehealth apps and 86 telehealth websites, and performed a systematic analysis on this dataset.In this analysis, we conducted a comparison analysis and relevant content analysis of the telehealth apps as well as their security risks.Apart from the mobile platforms, we also inspected the telehealth websites' features, particularly those related to the use of metaverse to improve current telehealth solutions.To further understand people's attitude towards telehealth services, we invited users to participate in a user study aimed at revealing what impact COVID-19 has on users' willingness to adopt telehealth services and revealing the gap between the telehealth service and its users.Our result shows that 27.1% new iOS apps and 27.4% new Android apps were released after the COVID-19 announcement, and a surge of updates were noted within 4 weeks after the COVID-19 announcement.We further found that COVID-19 is frequently mentioned in telehealth app reviews in the second and third quarter of 2020, and the most mentioned aspects related to COVID-19 include family, test result and vaccine.According to our user study, COVID-19 has a significant impact on the selection of telehealth services, especially for female participants, people aged 46-55, and students.The investigation also finds out that the use of metaverse will significantly improves the effectiveness of traditional telehealth solutions. Lihong Tang, Tingmin Wu, Xiao Chen 0002, Sheng Wen, Wei Zhou 0044, Xiaogang Zhu 0001, Yang Xiang 0001 |
Connect. Sci. | 2 |
| 2023 | Investigating Users' Understanding of Privacy Policies of Virtual Personal Assistant ApplicationsabstractThe increasingly popular virtual personal assistant (VPA) services, e.g., Amazon Alexa and Google Assistant, enable third-party developers to create and release VPA apps for end users to access through smart speakers. Given that VPA apps handle sensitive personal data, VPA service providers require developers to release a privacy policy document to declare their data handling practice. The privacy policies are regarded as legal or semi-legal documents, which are usually lengthy and complex for users to understand. In this work, we conducted a subjective study to investigate the level of users’ understanding of the privacy policies, targeting the VPA apps (i.e., skills) of Amazon Alexa, the most popular VPA service. Our study focused on technical terms, one of the greatest hurdles to users’ understanding. We found that 84.2% of our participants faced difficulty in understanding technical terms appeared in the skills’ privacy policies, even for participants with IT background. Additionally, 64.3% of them reported that explanations for the technical terms are generally lacking. To address this issue, we proposed two principles, i.e., domain-specificity principle and implication-oriented principle, to guide skill developers in creating easy-to-understand privacy policies. We evaluated their effectiveness by creating explanation sentences for 23 representative terms and examining users’ understanding through a second user study. Our results show that using explanation sentences based on these principles can significantly improve users’ understanding. Baiqi Chen, Tingmin Wu, Yanjun Zhang 0002, Mohan Baruwal Chhetri, Guangdong Bai |
AsiaCCS | 2 |
| 2023 | Dynalogue: A Transformer-Based Dialogue System with Dynamic AttentionabstractBusinesses face a range of cyber risks, both external threats and internal vulnerabilities that continue to evolve over time. As cyber attacks continue to increase in complexity and sophistication, more organisations will experience them. For this reason, it is important that organisations seek timely consultancy from cyber professionals so that they can respond to and recover from cyber attacks as quickly as possible. However, huge surges in cyber attacks have long left cyber professionals short of what is required to cover the security needs. This problem is getting worse when an increasing number of people choose to work from home during the pandemic because this situation usually yields extra communication cost. Rongjunchen Zhang, Tingmin Wu, Xiao Chen 0002, Sheng Wen, Surya Nepal, Cécile Paris, Yang Xiang 0001 |
WWW | 2 |
| 2023 | How Does Visualisation Help App Practitioners Analyse Android Apps?abstractBehaviour analysis is essential for the security verification of suspicious Android applications, but analysts are usually faced with a huge obstacle when conducting the app behaviour analysis. They are expected to have comprehensive knowledge of different IT fields and a strong awareness of cyber threats. However, training a new security analyst typically requires a significant amount of time and can be extremely costly. Although there are tools available to assist analysts in studying Android behaviour and security, the completion of this task still heavily relies on the experience of the analysts. To address this problem, we recognise visualisation as a promising method and conduct a series of controlled experiments to demonstrate its effectiveness in the context of Android app behaviour and security analysis. We accordingly develop a visualisation tool based on apps’ call graphs (CG) (namedVisualDroid) and conduct an experiment and a follow-up interview. Compared to existing solutions, the results suggest that the CG-based visualisation solution (VisualDroid) can lower the barriers to Android behaviour and security analysis. The user study reveals that the platform includes CG-based visualisation components leads to a statistically significant improvement in Android behaviour analysis and security awareness. More specifically, it improvesAPK Analyzer,JD-GUI,JD-GUI+FlowDroidby 71.4%, 35.7%, and 39.2% in terms of the effectiveness of behaviour analysis. Participants who useVisualDroidalso show improvements in the aspect of security awareness with an increase of 155% againstAPK Analyzer, 96% againstJD-GUI, and 59.3%JD-GUI+FlowDroid. Lihong Tang, Tingmin Wu, Xiao Chen 0002, Sheng Wen, Li Li 0029, Xin Xia 0001, Marthie Grobler, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2023 | SAM: Multi-turn Response Selection Based on Semantic Awareness MatchingabstractMulti-turn response selection is a key issue in retrieval-based chatbots and has attracted considerable attention in the NLP (Natural Language processing) field. So far, researchers have developed many solutions that can select appropriate responses for multi-turn conversations. However, these works are still suffering from the semantic mismatch problem when responses and context share similar words with different meanings. In this article, we propose a novel chatbot model based on Semantic Awareness Matching, called SAM. SAM can capture both similarity and semantic features in the context by a two-layer matching network. Appropriate responses are selected according to the matching probability made through the aggregation of the two feature types. In the evaluation, we pick 4 widely used datasets and compare SAM’s performance to that of 12 other models. Experiment results show that SAM achieves substantial improvements, with up to 1.5% R 10 @1 on Ubuntu Dialogue Corpus V2, 0.5% R 10 @1 on Douban Conversation Corpus, and 1.3% R 10 @1 on E-commerce Corpus. Rongjunchen Zhang, Tingmin Wu, Sheng Wen, Surya Nepal, Cécile Paris, Yang Xiang 0001 |
ACM Trans. Internet Techn. | 2 |
| 2022 | Email Summarization to Assist Users in Phishing IdentificationabstractCyber-phishing attacks recently became more precise, targeted, and tailored by training data to activate only in the presence of specific information or cues. They are adaptable to a much greater extent than traditional phishing detection. Hence, automated detection systems cannot always be 100% accurate, increasing the uncertainty around expected behavior when faced with a potential phishing email. On the other hand, human-centric defence approaches focus extensively on user training but face the difficulty of keeping users up to date with continuously emerging patterns. Therefore, advances in analyzing the content of an email in novel ways along with summarizing the most pertinent content to the recipients of emails is a prospective gateway to furthering how to combat these threats. Addressing this gap, this work leverages transformer-based machine learning to (i) analyze prospective psychological triggers, to (ii) detect possible malicious intent, and to (iii) create representative summaries of emails. We then amalgamate this information and present it to the user to allow them to (i) easily decide whether the email is "phishy" and (ii) self-learn advanced malicious patterns. Amir Kashapov, Tingmin Wu, Alsharif Abuadbba, Carsten Rudolph |
AsiaCCS | 2 |
| 2022 | Profiler: Distributed Model to Detect PhishingabstractMany Machine Learning (ML) based phishing detection algorithms are not adept to recognise "concept drift"; attackers introduce small changes in the statistical characteristics of their phishing attempts to successfully bypass detection. This leads to the classification problem of frequent false positives and false negatives, and a reliance on manual reporting of phishing by users. Profiler is a distributed phishing risk assessment tool that combines three email profiling dimensions: (1) threat level, (2) cognitive manipulation, and (3) email content type to detect email phishing. Unlike pure ML-based approaches, Profiler does not require large data sets to be effective and evaluations on real-world data sets show that it can be useful in conjunction with ML algorithms to mitigate the impact of concept drift. Mariya Shmalko, Alsharif Abuadbba, Raj Gaire 0001, Tingmin Wu, Hye-Young Paik, Surya Nepal |
ICDCS | 4 |
| 2022 | RAIDER: Reinforcement-Aided Spear Phishing Detector
Keelan Evans, Alsharif Abuadbba, Tingmin Wu, Kristen Moore, Ganna Pogrebna, Surya Nepal, Mike Johnstone |
NSS | 3 |
| 2022 | Analysis of Trending Topics and Text-based Channels of Information Delivery in CybersecurityabstractComputer users are generally faced with difficulties in making correct security decisions. While an increasingly fewer number of people are trying or willing to take formal security training, online sources including news, security blogs, and websites are continuously making security knowledge more accessible. Analysis of cybersecurity texts from this grey literature can provide insights into the trending topics and identify current security issues as well as how cyber attacks evolve over time. These in turn can support researchers and practitioners in predicting and preparing for these attacks. Comparing different sources may facilitate the learning process for normal users by creating the patterns of the security knowledge gained from different sources. Prior studies neither systematically analysed the wide range of digital sources nor provided any standardisation in analysing the trending topics from recent security texts. Moreover, existing topic modelling methods are not capable of identifying the cybersecurity concepts completely and the generated topics considerably overlap. To address this issue, we propose a semi-automated classification method to generate comprehensive security categories to analyse trending topics. We further compare the identified 16 security categories across different sources based on their popularity and impact. We have revealed several surprising findings as follows: (1) The impact reflected from cybersecurity texts strongly correlates with the monetary loss caused by cybercrimes, (2) security blogs have produced the context of cybersecurity most intensively, and (3) websites deliver security information without caring about timeliness much. Tingmin Wu, Wanlun Ma, Sheng Wen, Xin Xia 0001, Cécile Paris, Surya Nepal, Yang Xiang 0001 |
ACM Trans. Internet Techn. | 1 |
| 2020 | What risk? I don't understand. An Empirical Study on Users' Understanding of the Terms Used in Security TextsabstractUsers receive a multitude of security information in written articles, e.g., newspapers, security blogs, and training materials. However, prior research suggests that these delivery methods, including security awareness campaigns, mostly fail to increase people's knowledge about cyber threats. It seems that users find such information challenging to absorb and understand. Yet, to raise users' security awareness and understanding, it is essential to ensure the users comprehend the provided information so that they can apply the advice it contains in practice. We conducted a subjective study to measure the level of users' understanding of security texts. We find that 61% of the terms security experts used in their writings are hard for the public to understand, even for people with some IT backgrounds. We also observe that 88% of security texts have at least one such term. Moreover, we notice that existing dictionaries, including the online ones (e.g., Google Dictionary), cover no more than 35% of the terms found in security texts. To improve users' ability to understand security texts, we developed a framework to build a user-oriented security-centric dictionary from multiple sources. To evaluate the effectiveness of the dictionary, we developed a tool as a service to detect technical terms and explain their meanings to the user in pop-ups. The results of a subjective study to measure the tool's performance showed that it could increase users' ability to understand security articles by 30%. Tingmin Wu, Rongjunchen Zhang, Wanlun Ma, Sheng Wen, Xin Xia 0001, Cécile Paris, Surya Nepal, Yang Xiang 0001 |
AsiaCCS | 1 |
| 2019 | Every word is valuable: Studied influence of negative words that spread during election period in social mediaabstractSummary Studying the influence of negative words that spread during election period is an important work in social media. Most of current methods rely on sentiment analysis of tweets to determine the users' preference. However, sentiment analysis can only makes use of emotional words (ie, adverbs and adjectives), which only take 30 percent of the context in the Internet. According to our empirical analysis based on real datasets, the bias on word selection largely reduced the accuracy of the context in the Internet. In order to address this critical problem, we propose a new method that makes use of nouns with emotional context to determine the election preference of each user. By collecting the frequencies of words in context, we weigh the impact of each supportive/objective noun to strengthen the determination of users' preference. Final results will further be integrated to examine the effectiveness and efficiency of our proposed method. To indicate this idea, we collect and adopt real datasets (UK Prime Minister 2017 and US President Campaign 2016) in the experiments. All the experiment results suggested that our integrated method largely outperformed previous prediction methods. In particular, the prediction results were quite similar to the final results of the UK and US election. Meanwhile, for UK election, we found that the daily approval rate is closely related to the event happened everyday. Xiangyu Hu 0006, Lemin Li, Tingmin Wu, Xiaoxiang Ai, Sheng Wen |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | Catering to Your Concerns: Automatic Generation of Personalised Security-Centric Descriptions for Android AppsabstractAndroid users are increasingly concerned with the privacy of their data and security of their devices. To improve the security awareness of users, recent automatic techniques produce security-centric descriptions by performing program analysis. However, the generated text does not always address users’ concerns as they are generally too technical to be understood by ordinary users. Moreover, different users have varied linguistic preferences that do not match the text. Motivated by this challenge, we develop an innovative scheme to help users avoid malware and privacy-breaching apps by generating security descriptions that explain the privacy and security related aspects of an Android app in clear and understandable terms. We implement a prototype system, PERSCRIPTION, to generate personalised security-centric descriptions that automatically learn users’ security concerns and linguistic preferences to produce user-oriented descriptions. We evaluate our scheme through experiments and user studies. The results clearly demonstrate the improvement on readability and users’ security awareness of PERSCRIPTION’s descriptions compared to existing description generators. Tingmin Wu, Lihong Tang, Rongjunchen Zhang, Sheng Wen, Cécile Paris, Surya Nepal, Marthie Grobler, Yang Xiang 0001 |
ACM Trans. Cyber Phys. Syst. | 1 |
| 2019 | STC: Exposing Hidden Compromised Devices in Networked Sustainable Green Smart Computing Platforms by Partial ObservationabstractLarge-scale smart computing is generally more vulnerable to cyber attacks since their system devices are normally distributed as networked platforms and each device could be a target and get compromised. Due to resource constraints (i.e., Sustainable Computing demand) and cost-efficiency issues (i.e., Green Computing demand), we usually monitor only a few devices (i.e., partial observation) to ensure all operations across different platforms are under a secure environment. This leads to a critical problem for detecting compromised devices that are out of surveillance. To the best of our knowledge, this problem has not been solved so far. In this paper, we propose an unsupervised classifier based on source-tracing technique (STC in short) to expose hidden compromised devices with partial observation on the networked sustainable green smart computing platforms. STC mainly focuses on the cyber threats that can spread in the platform and compromise various system devices. To expose hidden compromised devices that are out-of-surveillance, STC first captures the spreading source by the reverse dissemination technique, and then relies on microscopic propagation modelling to probabilistically identify the most probable compromised devices. We carried out a series of experiments to validate the performance of our proposed method. The evaluations are based on three real networked platforms: Air Traffic Control system, AS-level Internet platform, and US Power Grid. The experiment results demonstrated that STC can accurately expose the hidden compromised devices in terms of following aspects: 1) Source-tracing (more than 80 percent runs got exact real source and 95 percent within two hops of real source); 2) Modelling (very close to the simulation results); 3) Exposing accuracy (almost all > 90 percent); and 4) Comparison to baseline (superiority to three supervised and two unsupervised classifiers). Derek Wang, Tingmin Wu, Sheng Wen, Xiaofeng Chen 0001, Yang Xiang 0001, Wanlei Zhou 0001 |
IEEE Trans. Sustain. Comput. | 2 |
| 2018 | Twitter spam detection: Survey of new approaches and comparative study
Tingmin Wu, Sheng Wen, Yang Xiang 0001, Wanlei Zhou 0001 |
Comput. Secur. | 1 |
| 2017 | How Spam Features Change in Twitter and the Impact to Machine Learning Based Detection
Tingmin Wu, Derek Wang, Sheng Wen, Yang Xiang 0001 |
ISPEC | 1 |
| 2017 | Detecting spamming activities in twitter based on deep-learning techniqueabstractSummary Twitter spam has long been a critical but difficult problem to be addressed. So far, researchers have developed a series of machine learning–based methods and blacklisting techniques to detect spamming activities on Twitter. According to our investigation, current methods and techniques have achieved the accuracy of around 87%. However, because of the problems of spam drift and information fabrication, these machine learning–based methods cannot efficiently detect spam activities in real‐life scenarios. Meanwhile, the blacklisting method also cannot catch up with the variations of spamming activities, as manually inspecting suspicious URLs is extremely timeconsuming. In this paper, we proposed a novel technique based on deep‐learning technique to address the above challenges. The syntax of each tweet will be learned through WordVector and trained by deep learning. We then constructed a binary classifier to differentiate spam and regular tweets. In experiments, we collected and labeled a 10‐day real tweet dataset as ground truth to evaluate our proposed method. We first went for empirical analysis with a series of comparisons to other methods: (1) performance of different classifiers, (2) other existing text‐based methods, and (3) nontext‐based detection techniques. According to the experiment results, our proposed method largely outperformed previous methods. We further conducted principle component analysis on typical methods to theoretically justify the outperformance of our method. We extracted all kinds of features via dimensionality reduction. It was found that our features were most distinct among all the detection methods. This well demonstrated the outperformance of our method. Tingmin Wu, Sheng Wen, Shigang Liu, Jun Zhang 0010, Yang Xiang 0001, Majed A. AlRubaian, Mohammad Mehedi Hassan |
Concurr. Comput. Pract. Exp. | 1 |