Guoai Xu

dblp:76/10013 · DBLP profile ↗
← Back
63ranked-venue papers
1as first author
41since 2021 · last 2026
0000-0002-9582-0698ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 21 · 13 since 2021Security and privacy · 17 · 1 first-author · 10 since 2021Computer networks · 13 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 IA-CLIP: A Single-Source Industrial Anomaly Detection Method for Multi-Target Domain Generalization
abstract
In industrial manufacturing, ensuring product quality is of paramount importance. A key component of this process is anomaly detection, which aims to promptly identify defective products to reduce operational losses. However, practical industrial environments are characterized by complexity, including limited availability of labeled data, a wide variety of defect categories, and frequent changes in these categories. Such factors pose significant challenges to the effective cross-domain generalization of anomaly detection methods. To address this limitation, IA-CLIP, a novel framework that enhances cross-domain generalization for industrial anomaly detection, is proposed. IA-CLIP integrates global and local prompts with contrastive learning to overcome the limitations of existing approaches. The proposed class-agnostic global-local semantic prompts enable the model to capture general patterns of normality and anomaly without relying on object-specific semantics. We further introduce a Similarity-aware Triplet Contrastive Learning strategy to facilitate complementary learning between global and local prompts, and an Adaptive Focal Contrastive Learning scheme to help the model focus more effectively on hard-to-identify anomalous regions. Extensive experiments on nine real-world target-domain datasets, covering 50 categories of industrial products, demonstrate that IA-CLIP achieves impressive cross-domain generalization performance in realistic industrial settings. Code and data will be released upon publication.
Guoai Xu, Jianping Yin
IEEE Trans Autom. Sci. Eng.2
2026 PVF-FD: Free-Rider Detection in Privacy-Preserving Vertical Federated Learning
abstract
Vertical federated learning (VFL) enables collaborative learning across different entities with disjoint data features for the same instances. In VFL, passive parties with partial data features extract embeddings from their data and forward them to the active party, which holds disjoint data features and labels, for aggregation and subsequent prediction. However, some passive parties may act as free-riders and submit valueless embeddings to deceive rewards, which undermines the fairness of collaborative training and increases communication overhead. Compared to horizontal federated learning (HFL), detecting free-riders in VFL is more challenging due to the distinct data features and heterogeneous embeddings each party produces. This makes it difficult to identify disguised embeddings of free-riders using anomaly detection methods typically employed in HFL. This paper proposes the first free-rider detection strategy in VFL using an unsupervised auxiliary task based on maximum mean discrepancy (MMD). It helps benign parties capture shared information from the active party, resulting in smaller MMD distances for benign embeddings compared to those of free-riders. Additionally, considering that embeddings may be exploited to infer local data, we introduce PVF-FD, a ciphertext-domain verifiable embedding learning scheme that enables the main and auxiliary tasks to be performed simultaneously in a privacy-preserving manner. We formally analyze the security of PVF-FD. Experimental results demonstrate that PVF-FD can effectively detect free-riders, reduce communication overhead, and maintain the performance of the main task.
Zhongyun Hua, Yifeng Zheng 0001, Guoai Xu, Xiaohua Jia
IEEE Trans. Dependable Secur. Comput.5
2026 Enabling Reliable and Anonymous Data Collection for Fog-Assisted Mobile Crowdsensing With Malicious User Detection
Zhongyun Hua, Yifeng Zheng 0001, Rushi Lan, Qing Liao 0001, Guoai Xu
IEEE Trans. Mob. Comput.6
2025 Multi-View Collaborative Learning Network for Speech Deepfake Detection
abstract
As deep learning techniques advance rapidly, deepfake speech synthesized through text-to-speech or voice conversion networks is becoming increasingly realistic, posing significant challenges for detection and raising potential threats to social security. This growing realism has prompted extensive research in speech deepfake detection. However, current detection methods primarily focus on extracting features from either the raw waveform or the spectrogram, often overlooking the valuable correspondences between these two modalities that could enhance the detection of previously unseen types of deepfakes. In this work, we propose a multi-view collaborative learning network for speech deepfake detection, which jointly learns robust speech representations from both raw waveforms and spectrograms. Specifically, we first design a Dual-Branch Contrastive Learning (DBCL) framework for learning different view features. DBCL consists of two branches that learn representations from the raw waveform or the spectrogram and utilizes contrastive learning to enhance inter- and inner-view correlations. Additionally, we introduce a Waveform-Spectrogram Fusion Module (WSFM) to exchange multi-view information for collaborative learning. In the feature learning process, WSFM converts features between views and merges them adaptively using waveform-spectrogram cross-attention. The final detection is conducted based on the concatenation of the waveform and spectrogram features. We conduct extensive experiments on four benchmark deepfake speech detection datasets, and the experimental results demonstrate that our method can achieve better detection performance than current state-of-the-art detection methods.
Zhongyun Hua, Rushi Lan, Yifang Guo, Yushu Zhang 0001, Guoai Xu
AAAI6
2025 Multi-view Leaderboard: Towards Evaluating the Code Intelligence of LLMs From Multiple Views
abstract
Large Language Models (LLMs) have shown remarkable performance in code intelligence tasks, prompting the development of various benchmarks and leaderboards to assess their effectiveness across diverse programming scenarios. However, existing leaderboards often rely on coarse-grained metrics and overlook performance variations across different types of tasks. In this paper, we introduce Multi-view Leaderboard, a comprehensive evaluation framework designed to assess the coding capabilities of LLMs from multiple views. Our leaderboard partitions widely-used datasets such as HumanEval, MBPP, and ComplexCodeEval into subsets based on factors like prompt length, problem complexity, and task type. It supports four popular code intelligence tasks including code generation, code completion, test case generation, and API recommendation. Additionally, our leaderboard presents results using ranking tables, line charts, radar charts, and heatmaps. Based on LLMs’ performance on different subsets, we provide model recommendations tailored to different real-world scenarios via a Sankey diagram. A user study involving 11 participants revealed that 90% valued the leaderboard’s practical usefulness for analyzing LLMs’ code intelligence from multiple perspectives. The Multi-view Leaderboard is available at https://huggingface.co/spaces/MVLLL/Multi-view-leaderboard. The demonstration video is available at https://youtu.be/J-zQiOYa1Y8
Zexun Zhan, Cuiyun Gao 0001, Yujia Chen 0004, Guoai Xu, Chun Yong Chong, Shan Gao 0009, Xin Xia 0001
APSEC5
2025 Decoding Secret Memorization in Code LLMs Through Token-Level Characterization
abstract
Code Large Language Models (LLMs) have demonstrated remarkable capabilities in generating, understanding, and manipulating programming code. However, their training process inadvertently leads to the memorization of sensitive information, posing severe privacy risks. Existing studies on memorization in LLMs primarily rely on prompt engineering techniques, which suffer from limitations such as widespread hallucination and inefficient extraction of the target sensitive information. In this paper, we present a novel approach to characterize real and fake secrets generated by Code LLMs based on token probabilities. We identify four key characteristics that differentiate genuine secrets from hallucinated ones, providing insights into distinguishing real and fake secrets. To overcome the limitations of existing works, we propose DeSec,a two-stage method that leverages token-level features derived from the identified characteristics to guide the token decoding process. DeSec consists of constructing an offline token scoring model using a proxy Code LLM and employing the scoring model to guide the decoding process by reassigning token likelihoods. Through extensive experiments on four state-of-the-art Code LLMs using a diverse dataset, we demonstrate the superior performance of DeSec in achieving a higher plausible rate and extracting more real secrets compared to existing baselines. Our findings highlight the effectiveness of our token-level approach in enabling an extensive assessment of the privacy leakage risks associated with Code LLMs.
Yuqing Nie, Chong Wang 0013, Kailong Wang 0001, Guoai Xu, Guosheng Xu 0001, Haoyu Wang 0001
ICSE4
2025 On the (In)Security of Non-resettable Device Identifiers in Custom Android Systems
abstract
User tracking is critical in the mobile ecosystem and relies on device identifiers to build user profiles. Early versions of Android allowed third-party apps to easily access non-resettable identifiers such as serial numbers and IMEI. As privacy concerns grew, Google has tightened identifier access in native Android. In response, stakeholders in custom Android systems introduced covert channels (e.g., system properties and settings) to maintain consistent and stable identifier access across systems and devices, which undoubtedly increases privacy risks. This paper examines the introduction of such channels through system customization and their vulnerability due to poor access control. We present IDRADAR, a scalable and accurate approach for identifying vulnerable properties and settings in custom Android systems. Applying our approach to 1,814 custom ROMs, we identified 8,192 system properties and 3,620 settings that store non-resettable device identifiers. Among these, 3,477 properties and 1,336 settings lack adequate access control and could be exploited by third-party apps to track users without permissions. Further validation on real devices demonstrates the effectiveness of our approach. Compared to state-of-the-art, IDRADAR offers improved scalability and analytical capabilities. Additionally, we investigate the root causes of the access control deficiencies and observe that such vulnerabilities frequently recur across devices from the same OEMs. We have reported our findings to the respective vendors and received positive confirmations. Our work underscores the need for greater scrutiny of covert access to device identifiers and better solutions to safeguard user privacy during system customizations.
Zikan Dong, Liu Wang 0002, Guoai Xu, Haoyu Wang 0001
ASE3
2025 Assuring Certified Database Utility in Privacy-Preserving Database Fingerprinting
Zhongyun Hua, Yifeng Zheng 0001, Tao Xiang 0001, Guoai Xu, Xingliang Yuan
USENIX Security Symposium5
2025 A Provably Secure Authentication Protocol Based on PUF and ECC for IoT Cloud-Edge Environments
abstract
The Internet of Things (IoT) cloud model provides an efficient scheme for rapid collection, storage, processing, and analysis of massive node data, and its application has gradually expanded to key areas such as healthcare and transportation. However, the security issues of open channel transmission in IoT still persist. Researchers have proposed a lot of solutions, but the forward secrecy, session key security, and other aspects have not been effectively solved. This paper proposes a provably secure authenticated key agreement scheme, which constructs a secure channel between endpoint, gateway, and cloud server (CS). Compared with other schemes, this scheme has three characteristics: (1) According to the different computing resources of devices, gateways and CSs, a segmented differential authentication and secret key negotiation protocol is designed by using cryptographic primitives with different computing overheads; (2) after verification with the ProVerif tool, rigorous proof with the real‐or‐random (ROR) model, and informal analysis, the protocol has been proven to be secure, effectively guarding against typical threats; and (3) compared with the five most recent schemes, it can be seen that the protocol is at least 35% superior to other schemes in endpoint computational overhead, and it meets 10 security objectives, making it very suitable for application scenarios where endpoint resources are limited.
Guosheng Xu 0001, Chenyu Wang 0002, Jinwen Xi, Guoai Xu
IET Inf. Secur.5
2025 Patch Feature Transformation: An Anomaly Detection Method with Succinct Feature Filtering
abstract
Anomaly detection is often approached as an out-of-distribution (OOD) detection task, where a feature distribution from normal samples is constructed, and deviations are flagged as anomalies. This approach is dependent on manual labeling, as subtle visual anomalies can be easily overlooked, resulting in the potential for bias in labeling and subsequent unsatisfactory detection results. Based on the issue, we propose Anomaly Detection with Succinct Feature Filtering (ADSFF) for unlabeled samples. Our method avoids sample labeling bias and provides a solution to the coexistence of anomalous and normal features in the feature space of unlabeled samples. ADSFF includes a data preprocessing module and a feature filtering module, where the data preprocessing module improves the visibility of subtle anomalies, while the feature filtering module screens the local features of the samples. In feature filtering, we found that feedforward neural networks do not lose feature information during the feature transformation process. Consequently, we utilized feedforward neural networks for feature filtering and achieved expected results. Furthermore, we investigate the impact of sample imbalance on the task of anomaly detection using unlabeled samples. This paper assesses the performance of ADSFF using the MVTec AD and BeanTech Anomaly Detection (BTAD) datasets. The results demonstrate that ADSFF achieves an average area under the curve (AUC) of 0.978 on the MVTec AD and an average AUC of 0.942 on the BTAD. ADSFF outperformed other methods on seven test datasets in MVTec AD, achieving the highest average accuracy on MVTec AD.
Guoai Xu, Jianping Yin
Int. J. Pattern Recognit. Artif. Intell.2
2025 Statistical Fault Attacks on ASCON Using Improved Square Euclidean Imbalance
abstract
In current environment of frequent information exchange between Internet of Things (IoT) devices, traditional cryptographic algorithms often fail to effectively play a role in resource-limited electronic devices, which highlights the importance of lightweight cryptographic algorithms. NIST has completed the standardization process of lightweight cryptographic algorithms, and ASCON has become the ultimate winner. It is foreseeable that its application demand will continue to grow in the future, so the security analysis for ASCON is of significant value. Among various cryptographic algorithm analysis methods, fault attack is an efficient choice. Currently, there are also some fault attack methods for ASCON, but they share a common issue: their time complexity is too high for practical application. In view of this, this article proposes a more efficient fault attack method for ASCON, including several key points: First, several new statistical scoring function are constructed based on square euclidean imbalance (SEI). Second, a single S-box fault model is proposed to reduce the complexity of injection. Lastly, the complete key is recovered by combining statistical ineffective fault attack (SIFA), statistical effective fault attack (SEFA), and statistical hybrid fault attack (SHFA). The time complexity is$2^{16.57}$,$2^{14.91}$, and$2^{14.52}$, respectively. Experimental results on a Python implementation of ASCON, the results show that our scheme significantly reduces the time complexity of attacks, which can provide warnings for cryptography design and applications, and pay more attention to avoiding such risks.
Guosheng Xu 0001, Yuque Zhang, Chenyu Wang 0002, Guoai Xu
IEEE Internet Things J.5
2025 Group-Capability-Based Access Control with Ring Signature
Shihong Zou, Guoai Xu, Jinwen Xi
J. Inf. Secur. Appl.3
2024 MalCertain: Enhancing Deep Neural Network Based Android Malware Detection by Tackling Prediction Uncertainty
abstract
The long-lasting Android malware threat has attracted significant research efforts in malware detection. In particular, by modeling malware detection as a classification problem, machine learning based approaches, especially deep neural network (DNN) based approaches, are increasingly being used for Android malware detection and have achieved significant improvements over other detection approaches such as signature-based approaches. However, as Android malware evolve rapidly and the presence of adversarial samples, DNN models trained on early constructed samples often yield poor decisions when used to detect newly emerging samples. Fundamentally, this phenomenon can be summarized as the uncertainly in the data (noise or randomness) and the weakness in the training process (insufficient training data). Overlooking these uncertainties poses risks in the model predictions. In this paper, we take the first step to estimate the prediction uncertainty of DNN models in malware detection and leverage these estimates to enhance Android malware detection techniques. Specifically, besides training a DNN model to predict malware, we employ several uncertainty estimation methods to train a Correction Model that determines whether a sample is correctly or incorrectly predicted by the DNN model. We then leverage the estimated uncertainty output by the Correction Model to correct the prediction results, improving the accuracy of the DNN model. Experimental results show that our proposed MalCertain effectively improves the accuracy of the underlying DNN models for Android malware detection by around 21% and significantly improves the detection effectiveness of adversarial Android malware samples by up to 94.38%. Our research sheds light on the promising direction that leverages prediction uncertainty to improve prediction-based software engineering tasks.
Guosheng Xu 0001, Liu Wang 0002, Xusheng Xiao, Xiapu Luo, Guoai Xu, Haoyu Wang 0001
ICSE6
2024 Same App, Different Behaviors: Uncovering Device-specific Behaviors in Android Apps
abstract
The Android ecosystem is significantly challenged by fragmentation, arising from diverse system versions, device specifications, and manufacturer customizations. The growing divergence among devices leads to marked variations in how a given app behaves across diverse devices. This is referred to as device-specific behaviors. Fragmentation not only complicates development processes but also impacts the overall industry by increasing maintenance costs and potentially harming user experience due to inconsistent app performance. In this work, we present the first large-scale empirical study of device-specific behaviors in real-world Android apps. We have designed a three-phase static analysis framework to accurately detect and understand the device-specific behaviors. Upon employing our tool on a dataset comprising more than 20,000 apps, we detected device-specific behaviors in 2,357 of them. By examining the distribution of device-specific behaviors, our analysis revealed that apps within the Chinese third-party app market exhibit more such behaviors compared to their counterparts in Google Play. Additionally, these behaviors are more likely to feature dominant brands that hold larger market shares. Reflecting this, we have classified these device-specific behaviors into 29 categories based on the functionalities implemented, providing a structured insight that is crucial for developers and stakeholders in the industry. Beyond the common behaviors, such as issue fixes and feature adaptations, we have observed 33 aggressive apps, including popular ones with millions of downloads. These apps abuse system properties of customized ROMs to obtain user-unresettable identifiers without requiring any permissions, posing significant privacy risks. Finally, we investigated the origins of device-specific behaviors, highlighting the significant challenges developers encounter in implementing them comprehensively. Our research aims to inform and equip industry practitioners with knowledge to enhance user experience and user privacy, marking a critical step toward addressing the less touched yet vital aspect of device-specific behaviors in the Android ecosystem.
Zikan Dong, Yanjie Zhao 0001, Tianming Liu 0002, Chao Wang 0097, Guosheng Xu 0001, Guoai Xu, Lin Zhang 0062, Haoyu Wang 0001
ASE6
2024 Exploring Covert Third-party Identifiers through External Storage in the Android New Era
Zikan Dong, Tianming Liu 0002, Jiapeng Deng, Haoyu Wang 0001, Li Li 0029, Guosheng Xu 0001, Guoai Xu
USENIX Security Symposium9
2024 APIGen: Generative API Method Recommendation
abstract
Automatic API method recommendation is an essential task of code intelligence, which aims to suggest suitable APIs for programming queries. Existing approaches can be categorized into two primary groups: retrieval-based and learning-based approaches. Although these approaches have achieved remarkable success, they still come with notable limitations. The retrieval-based approaches rely on the text representation capabilities of embedding models, while the learning-based approaches require extensive task-specific labeled data for training. To mitigate the limitations, we propose APIGen, a generative API recommendation approach through enhanced in-context learning (ICL). APIGen has a powerful representation capability and can make effective recommendations with only a few examples via I CL. To overcome the limitations of standard ICL in capturing task-specific knowledge, APIGen involves two main components: (1) Diverse Examples Selection. APIGen searches for similar posts to the programming queries from the lexical, syntactical, and semantic perspectives, providing more informative examples for ICL. (2) Guided API Recommendation. APIGen enables large language models (LLMs) to perform reasoning before generating API recommendations, where the reasoning involves fine-grained matching between the task intent behind the queries and the factual knowledge of the APIs. With the reasoning process, APIGen makes recommended APIs better meet the programming requirement of queries and also enhances the interpretability of results. We compare APIGen with four existing approaches on two publicly available benchmarks. Experiments show that APIGen outperforms the best baseline CLEAR by 105.8% in method-level API recommendation and 54.3 % in class-level API recommendation in terms of SuccessRate@l. Besides, APIGen achieves an average 49.87 % increase compared to the zero-shot performance of popular LLMs such as GPT-4 in method-level API recommendation regardina the SuccessRate@ 3 metric.
Yujia Chen 0004, Cuiyun Gao 0001, Muyijie Zhu, Qing Liao 0001, Yong Wang 0008, Guoai Xu
SANER6
2024 WalletRadar: towards automating the detection of vulnerabilities in browser-based cryptocurrency wallets
Pengcheng Xia 0001, Zhaowen Lin, Pengbo Duan, Ningyu He, Kailong Wang 0001, Tianming Liu 0002, Yinliang Yue, Guoai Xu, Haoyu Wang 0001
Autom. Softw. Eng.10
2024 A robust and effective 3-factor authentication protocol for smart factory in IIoT
Shihong Zou, Qiang Cao 0006, Ruichao Lu, Chenyu Wang 0002, Guoai Xu, Huanhuan Ma, Yingyi Cheng, Jinwen Xi
Comput. Commun.5
2024 A Physician's Privacy-Preserving Authentication and Key Agreement Protocol Based on Decentralized Identity for Medical Data Sharing in IoMT
abstract
As well known, Internet of medical things (IoMT) produces large amounts of medical data and promotes the medical data sharing which serves the data user (i.e., physicians) to boost the clinical treatment and medical research. To protect data user’s privacy and data security during the sharing of medical data, data user must have a self-sovereign decentralized identity (DID) and data access authority. In existing solutions, data user’s privacy protection and authenticated-key-agreement (AKA) for protecting data security are worked independently, which easily results in typical security attacks (e.g., phishing inquiry attacks, ephemeral secret leakage attacks) during data access and system computing overload. To solve the challenge, a new credential-embedded authentication and key agreement scheme (CAKA) is proposed, which can seamlessly combine DID-credentials into AKA. First, CAKA supports bilateral authentication by allowing a digital user to authenticate its service provider, which can enhance the security of unilateral scheme (such as CanDID, IEEE S&P, 2021) and prevent phishing query attacks. Second, for secure data session communication, the user’s DID-credentials are used as the kernel of the session key (SK) generation. In security analysis and performance metrics comparisons, the results indicate that CAKA holds a significant advantage, especially, the storage costs, communication costs and computation costs consumed in CAKA are at least 43% reduction, compared to alternatives. In simulation experiments of CAKA, the results show that decentralized identity authentication and session key agreement are both less than 15 ms, that means CAKA is a practical and promising solution to medical data sharing.
Shihong Zou, Qiang Cao 0006, Chonghui Huangqi, Anpeng Huang, Yanping Li 0001, Chenyu Wang 0002, Guoai Xu
IEEE Internet Things J.7
2023 WELID: A Weighted Ensemble Learning Method for Network Intrusion Detection
abstract
The requirements for intrusion detection technology are getting higher and higher, with the rapid expansion of network applications. There have been many studies on intrusion detection, however, the accuracy of these models is not high enough and time-consuming, making them unavailable. In this paper, we propose a novel weighted ensemble learning method for network intrusion detection (WELID). Firstly, data preprocessing and feature selection algorithms are used to filter out some redundant and unrelated features. Next, anomaly detection is performed on the dataset using different base classifiers, and a layered ten-fold cross-validation method is used to prevent program overfitting. Then, the best classifiers are selected for the use of a multi-classifier fusion algorithm based on probability-weighted voting. We compare the proposed model with lots of efficient classifiers and state-of-the-art models for intrusion detection. The results show that the proposed model is superior to these models in terms of accuracy and time consumption.
Yuanchen Gao, Guosheng Xu 0001, Guoai Xu
ISCC3
2023 Tree-IDS: An Incremental Intrusion Detection System for Connected Vehicles
abstract
The rapid development of Internet of Vehicles technology has led to the continuous upgrading of the functions of connected vehicles. While connected vehicles bring convenience to people’s life, there are also many security threats. Connected vehicles not only have intra-vehicle networks communication, but also communicate with the external network. The diversity of communication methods makes the attack surface wider, and some new attacks are constantly emerging. In order to ensure vehicle security, This paper focuses on the attacks that are vulnerable to vehicles, and proposes an incremental intrusion detection system, which can not only detect attacks, but also incrementally learn new types of attacks. Experimental results illustrate that the proposed system can incrementally learn new attacks and avoid catastrophic forgetting problems, and can detect various types of known attacks with 99.99% accuracy on the Car-Hacking Dataset and 99.37% accuracy on the CICIDS2017.
Zixiang Bi, Guosheng Xu 0001, Chenyu Wang 0002, Guoai Xu
LCN5
2022 Challenges in decentralized name management: the case of ENS
abstract
DNS has often been criticized for inherent design flaws, which make the system vulnerable to attack. Further, domain names are not fully controlled by users, meaning that they can easily be taken down by authorities and registrars. Due to this, there have been efforts to build a decentralized name service that gives greater control to domain owners. The Ethereum Name Service (ENS) is a major example. Yet, no existing work has systematically studied this emerging system, particularly regarding security and misbehavior. To address this gap, we present the first large-scale measurement study of ENS. Our findings suggest that ENS has shown growth during its four years' evolution. We identify several security issues, including traditional name system problems, as well as new issues introduced by the unique properties of ENS. We find that attackers are abusing the system with thousands of squatting ENS names, a number of scam blockchain addresses and indexing of malicious websites. We further develop a new record persistence attack, to find that 22,716 .eth names (3.7% of all names) are vulnerable to name hijacking. Our exploration suggests that our community should invest more effort into the detection and mitigation of issues in decentralized name services.
Pengcheng Xia 0001, Haoyu Wang 0001, Zhou Yu 0002, Xiapu Luo, Guoai Xu, Gareth Tyson
IMC6
2022 Privacy Analysis of Period Tracking Mobile Apps in the Post-Roe v. Wade Era
abstract
To help people manage their health, period tracking apps have become very popular in recent years. However, the U.S. Supreme Court overturned Roe v. Wade on June 24, 2022. Abortion will be banned in more and more states. Since the health data stored in the period tracking apps can be used to infer whether the user has had or is considering an abortion, mobile users are worrying that these apps may disclose their sensitive information, which can be used to prosecute users. Although period tracking apps have received attention from the research community, no existing work has performed a systematic privacy analysis of these apps, especially in the Post-Roe v. Wade era. To fill the void, this paper presents a comprehensive privacy analysis of popular period tracking apps. We first collect 35 popular period tracking apps from Google Play. Then, we analyze the sensitive user data collected by the period tracking apps using traffic analysis and static analysis. Further we inspect their privacy policies and check the consistency of the privacy policy with the app’s behavior. In addition, we analyze the app reviews to understand the users’ concerns about the period tracking apps. Our study reveals that some period tracking apps have indeed collected sensitive information and have the potential to share the data with third-party authorities. It is urgent for these apps to take action to protect user privacy, and mobile users should pay special attention to this kind of apps they used.
Zikan Dong, Liu Wang 0002, Guoai Xu, Haoyu Wang 0001
ASE4
2022 What did you pack in my app? a systematic analysis of commercial Android packers
abstract
Commercial Android packers have been widely used by developers as a way to protect their apps from being tampered with. However, app packer is usually provided as an online service developed by security vendors, and the packed apps are well protected. It is thus hard to know what exactly is packed in the app, and few existing studies in the community have systematically analyzed the behaviors of commercial app packers. In this paper, we propose PackDiff, a dynamic analysis system to inspect the fine-grained behaviors of commercial packers. By instrumenting the Android system, PackDiff records the runtime behaviors of Android apps (e.g., Linux system call invocations, Java API calls, Binder interactions, etc.), which are further processed to pinpoint the additional sensitive behaviors introduced by packers. By applying PackDiff to roughly 200 apps protected by seven commercial packers, we observe the disappointing facts of existing commercial packers. Most app packers have introduced unnecessary behaviors (e.g., accessing sensitive data), serious performance and compatibility issues, and they can even be abused to create evasive malware and repackaged apps, which contradicts with their design purposes.
Zikan Dong, Hongxuan Liu, Liu Wang 0002, Xiapu Luo, Yao Guo 0001, Guoai Xu, Xusheng Xiao, Haoyu Wang 0001
ESEC/SIGSOFT FSE6
2022 Demystifying the underground ecosystem of account registration bots
abstract
Member services are a core part of most online systems. For example, member services in online social networks and video platforms make it possible to serve users customized content or track their footprint for a recommendation. However, there is a dark side to membership that lurks behind influencer marketing, coupon harvesting, and spreading fake news. All these activities rely heavily on owning masses of fake accounts, and to create new accounts efficiently, malicious registrants use automated registration bots with anti-human verification services that can easily bypass a website’s security strategies.
Yuhao Gao, Guoai Xu, Li Li 0029, Xiapu Luo, Chenyu Wang 0002, Yulei Sui
ESEC/SIGSOFT FSE2
2022 Lie to Me: Abusing the Mobile Content Sharing Service for Fun and Profit
abstract
Online content sharing is a widely used feature in Android apps. In this paper, we observe a new Fake-Share attack that adversaries can abuse existing content sharing services to manipulate the displayed source of shared content to bypass the content review of targeted Online Social Apps (OSAs) and induce users to click on the shared fraudulent content. We show that seven popular content-sharing services (including WeChat, AliPay, and KakaoTalk) are vulnerable to such an attack. To detect this kind of attack and explore whether adversaries have leveraged it in the wild, we propose DeFash, a multi-granularity detection tool including static analysis and dynamic verification. The extensive in-the-lab and in-the-wild experiments demonstrate that DeFash is effective in detecting such attacks. We have identified 51 real-world apps involved in Fake-Share attacks. We have further harvested over 24K Sharing Identification Information (SIIs) that can be abused by attackers. It is hence urgent for our community to take actions to detect and mitigate this kind of attack.
Guosheng Xu 0001, Hao Zhou 0043, Shucen Liu, Yutian Tang, Li Li 0029, Xiapu Luo, Xusheng Xiao, Guoai Xu, Haoyu Wang 0001
WWW9
2022 Efficient privacy-preserving user authentication scheme with forward secrecy for industry 4.0
Chenyu Wang 0002, Ding Wang 0002, Guoai Xu, Debiao He
Sci. China Inf. Sci.3
2022 CrowdHB: A Decentralized Location Privacy-Preserving Crowdsensing System Based on a Hybrid Blockchain Network
abstract
With the advent of the Internet of Things (IoT), crowdsensing, as a new emerging application of the IoT that employs ubiquitous mobile users with smartphones for data collection and processing, has further deepened our knowledge. However, the problems of the current crowdsensing systems regarding system security, user privacy, and user payment (UP) raise serious privacy and security concerns, which affect participants’ adoption of the system. The Blockchain technology allows for nondeterministic multiple parties to interact with each other anonymously in a network that is not fully trusted. In this article, we propose a new decentralized crowdsensing system, calledCrowdHB. Unlike other blockchain-based crowdsensing systems,CrowdHBadopts a hybrid blockchain architecture and uses smart contracts to achieve location privacy preservation and ensure data quality while improving the system performance. Furthermore, to optimize task assignments to mobile users, we propose a location privacy-preserving optimization mechanism (LPPOM) and the approach of consistency optimization (ACO) to achieve a tradeoff between user privacy and system performance. The extensive experimental results show that the proposedCrowdHBoutperforms the other crowdsensing systems in terms of task success rate and performance for a large number of mobile users and tasks.
Shihong Zou, Jinwen Xi, Guoai Xu, Miao Zhang 0011, Yueming Lu
IEEE Internet Things J.3
2022 A new efficient hierarchical multi-secret sharing scheme based on linear homogeneous recurrence relations
Jiangtao Yuan, Jing Yang 0035, Chenyu Wang 0002, Xingxing Jia, Fang-Wei Fu 0001, Guoai Xu
Inf. Sci.6
2022 A novel model for voice command fingerprinting using deep learning
abstract
Smart speakers are becoming increasingly popular and permeate many aspects of human life. To improve the security of smart speakers, voice commands transmitted over a network are encrypted; however, user privacy issues related to smart speakers continue to emerge. In fact, attackers are still able to infer the content of a user’s specific voice commands from encrypted traffic through machine learning methods to obtain private information for advertising or to carry out malicious attacks. This traffic analysis attack is referred to as a voice command fingerprinting attack. In recent years, research on improving the accuracy of voice command fingerprinting attacks has become a hot topic and remains a challenging task. To improve the accuracy of voice command fingerprinting attacks, we design a new method in this paper. We use an adaptive and dilated residual network to process spatial features. In addition, we find that using temporal features helps improve fingerprinting attack accuracy, and therefore design an attention-based bidirectional gated recurrent unit. Then, we effectively combine the two models. Our method achieves an accuracy greater than 93.36% in a closed-world scenario, which exceeds those of other state-of-the-art methods (2020 WiSec Wang et al.). In a more realistic open-world setting, our model is still effective, obtaining a true-positive rate of 99.50% and a false-positive rate of 0.1% compared to Sirinam et al.’s rates of 90.66% and 0.1%, respectively. We also demonstrate that our model has good generalizability, as our model can also be applied to website fingerprinting and outperforms 2018 CCS Sirinam et al.
Jianghan Mao, Chenyu Wang 0002, Guoai Xu, Shoufeng Cao, Xuanwen Zhang, Zixiang Bi
J. Inf. Secur. Appl.4
2022 CrowdLBM: A lightweight blockchain-based model for mobile crowdsensing in the Internet of Things
Jinwen Xi, Shihong Zou, Guoai Xu, Yueming Lu
Pervasive Mob. Comput.3
2022 Practical and Provably Secure Three-Factor Authentication Protocol Based on Extended Chaotic-Maps for Mobile Lightweight Devices
abstract
Due to the limitations of symmetric-key techniques, authentication and key agreement (AKA) protocols based on public-key techniques have attracted much attention, providing secure access and communication mechanism for various application environments. Among these public-key techniques used for AKA protocols, chaotic-map is more effective than scalar multiplication and modular exponentiation, and it offers a list of desirable cryptographic properties such as un-predictability, un-repeatability, un-certainty, and higher efficiency than scalar multiplication and modular exponentiation. Furthermore, it is usually believed that three-factor AKA protocols can achieve a higher security level than single- and two-factor protocols. However, none of existing three-factor AKA protocols can meet all security requirements. One of the most prevalent problems is how to balance security and usability, and particularly how to achieve truly three-factor security while providing password change friendliness. To deal with this problem, in this article we put forward a provably secure three-factor AKA protocol based on extended chaotic-maps for mobile lightweight devices, by adopting the techniques of “Fuzzy-Verifiers” and “Honeywords”. We prove the security of the proposed protocol in the random oracle model, assuming the intractability of extended chaotic-maps Computational Diffie-Hellman problem. We also simulate the protocol by using the AVISPA tool. The security analysis and simulation results show that our protocol can meet all 13 evaluation criteria regarding security. We also assess the performance of our protocol by comparing with seven other related protocols. The evaluation results demonstrate that our protocol offers better balance between security and usability over state-of-the-art ones.
Shuming Qiu, Ding Wang 0002, Guoai Xu, Saru Kumari
IEEE Trans. Dependable Secur. Comput.3
2022 Understanding Node Capture Attacks in User Authentication Schemes for Wireless Sensor Networks
abstract
Despite decades of intensive research, it is still challenging to design a practical multi-factor user authentication scheme for wireless sensor networks (WSNs). This is because protocol designers are confronted with a long-standing “security versus efficiency” dilemma: sensor nodes are lightweight devices with limited storage and computation capabilities, while the security requirements are demanding as WSNs are generally deployed for sensitive applications. Hundreds of proposals have been proposed, yet most of them have been found to be problematic, and the same mistakes are repeated again and again. Two of the most common security failures are regarding smart card loss attacks and node capture attacks. The former has been extensively investigated in the literature, while little attention has been given to understanding the node capture attacks. To alleviate this undesirable situation, this article takes a substantial step towards systematically exploring node capture attacks against multi-factor user authentication schemes for WSNs. We first investigate the various causes and consequences of node capture attacks, and classify them into ten different types in terms of the attack targets, adversary’s capabilities and vulnerabilities exploited. Then, we elaborate on each type of attack through examining 11 typical vulnerable protocols, and suggest corresponding countermeasures. Finally, we conduct a large-scale comparative measurement of 61 representative user authentication schemes for WSNs under our extended evaluation criteria. We believe that such a systematic understanding of node capture attacks would help design secure user authentication schemes for WSNs.
Chenyu Wang 0002, Ding Wang 0002, Guoai Xu, Huaxiong Wang
IEEE Trans. Dependable Secur. Comput.4
2021 Intelligent Preprocessing Selection for Pavement Crack Detection based on Deep Reinforcement Learning
abstract
With the rapid increase of traffic, the pressure on road maintenance is gradually increasing.Pavement crack is a common problem in all kinds of pavement diseases.In the actual production process, pavement images have different kinds of noise influence.The proposed algorithm is to select optimal preprocessing methods for pavement images in various conditions to improve the accuracy of crack detection.The algorithm includes two parts, a crack detection network and an intelligent preprocessing decision system.The crack detection network identifies the cracks in road images.The intelligent preprocessing decision system selects the best preprocessing method for pavement images based on the deep reinforcement method.The experiment results indicate that the validity and effectiveness of our proposed method.
Guosheng Xu 0001, Guoai Xu, Jiankun Cao
SEKE3
2021 Demystifying Illegal Mobile Gambling Apps
abstract
Mobile gambling app, as a new type of online gambling service emerging in the mobile era, has become one of the most popular and lucrative underground businesses in the mobile app ecosystem. Since its born, mobile gambling app has received strict regulations from both government authorities and app markets. However, to the best of our knowledge, mobile gambling apps have not been investigated by our research community. In this paper, we take the first step to fill the void. Specifically, we first perform a 5-month dataset collection process to harvest illegal gambling apps in China, where mobile gambling apps are outlawed. We have collected 3,366 unique gambling apps with 5,344 different versions. We then characterize the gambling apps from various perspectives including app distribution channels, network infrastructure, malicious behaviors, abused third-party and payment services. Our work has revealed a number of covert distribution channels, the unique characteristics of gambling apps, and the abused fourth-party payment services. At last, we further propose a “guilt-by-association” expansion method to identify new suspicious gambling services, which help us further identify over 140K suspicious gambling domains and over 57K gambling app candidates. Our study demonstrates the urgency for detecting and regulating illegal gambling apps.
Yuhao Gao, Haoyu Wang 0001, Li Li 0029, Xiapu Luo, Guoai Xu, Xuanzhe Liu
WWW5
2021 Beyond the virus: a first look at coronavirus-themed Android malware
Liu Wang 0002, Haoyu Wang 0001, Pengcheng Xia 0001, Yuanchun Li 0003, Lei Wu 0012, Yajin Zhou, Xiapu Luo, Yulei Sui, Yao Guo 0001, Guoai Xu
Empir. Softw. Eng.11
2021 G2F: A Secure User Authentication for Rapid Smart Home IoT Management
abstract
Internet-of-Things (IoT) devices are widely deployed nowadays. A large number of smart home IoT devices are hosted on a cloud server for easy management. Users can use their accounts to initiate operations and management on IoT devices through a cloud server, such as updating firmware and configuring devices. However, the cloud account may be hacked resulting in adversarial attacks to the hosted IoT devices. As a consequence, an adversary may perform malicious operations through the cloud remotely to the hosted IoT devices without user awareness. Motivated by this, in this article we propose gateway-based 2 factor authentication (G2F), a secure user authentication framework dedicated for a gateway based on the universal 2nd factor (U2F) protocol to enhance the security of IoT devices management. In G2F, the user authentication on the gateway is completed utilizing a hardware token that interacts with the local gateway node to guarantee the token owner’s presence. Furthermore, G2F can grant multiple simultaneous operations on IoT devices through just one user authentication. We implement a prototype to further evaluate the performance of G2F. Based on our realization on the commercial IoT server, i.e., Alibaba Cloud, G2F demonstrates the ability to protect against malicious attacks with high authentication efficiency.
Chao Wang 0097, Hao Luo 0001, Fan Zhang 0010, Feng Lin 0004, Guoai Xu
IEEE Internet Things J.6
2021 Spoofing Speaker Verification System by Adversarial Examples Leveraging the Generalized Speaker Difference
abstract
Speaker verification system has gained great popularity in recent years, especially with the development of deep neural networks and Internet of Things. However, the security of speaker verification system based on deep neural networks has not been well investigated. In this paper, we propose an attack to spoof the state-of-the-art speaker verification system based on generalized end-to-end (GE2E) loss function for misclassifying illegal users into the authentic user. Specifically, we design a novel loss function to deploy a generator for generating effective adversarial examples with slight perturbation and then spoof the system with these adversarial examples to achieve our goals. The success rate of our attack can reach 82% when cosine similarity is adopted to deploy the deep-learning-based speaker verification system. Beyond that, our experiments also reported the signal-to-noise ratio at 76 dB, which proves that our attack has higher imperceptibility than previous works. In summary, the results show that our attack not only can spoof the state-of-the-art neural-network-based speaker verification system but also more importantly has the ability to hide from human hearing or machine discrimination.
Yijie Shen, Feng Lin 0004, Guoai Xu
Secur. Commun. Networks4
2021 An Efficient Compartmented Secret Sharing Scheme Based on Linear Homogeneous Recurrence Relations
abstract
Multipartite secret sharing schemes are those that have multipartite access structures. The set of the participants in those schemes is divided into several parts, and all the participants in the same part play the equivalent role. One type of such access structure is the compartmented access structure, and the other is the hierarchical access structure. We propose an efficient compartmented multisecret sharing scheme based on the linear homogeneous recurrence (LHR) relations. In the construction phase, the shared secrets are hidden in some terms of the linear homogeneous recurrence sequence. In the recovery phase, the shared secrets are obtained by solving those terms in which the shared secrets are hidden. When the global threshold is t , our scheme can reduce the computational complexity of the compartmented secret sharing schemes from the exponential time to polynomial time. The security of the proposed scheme is based on Shamir’s threshold scheme, i.e., our scheme is perfect and ideal. Moreover, it is efficient to share the multisecret and to change the shared secrets in the proposed scheme.
Guoai Xu, Jiangtao Yuan, Guosheng Xu 0001, Zhongkai Dang
Secur. Commun. Networks1
2021 DeepWukong: Statically Detecting Software Vulnerabilities Using Deep Graph Neural Network
abstract
Static bug detection has shown its effectiveness in detecting well-defined memory errors, e.g., memory leaks, buffer overflows, and null dereference. However, modern software systems have a wide variety of vulnerabilities. These vulnerabilities are extremely complicated with sophisticated programming logic, and these bugs are often caused by different bad programming practices, challenging existing bug detection solutions. It is hard and labor-intensive to develop precise and efficient static analysis solutions for different types of vulnerabilities, particularly for those that may not have a clear specification as the traditional well-defined vulnerabilities. This article presents D eep W ukong , a new deep-learning-based embedding approach to static detection of software vulnerabilities for C/C++ programs. Our approach makes a new attempt by leveraging advanced recent graph neural networks to embed code fragments in a compact and low-dimensional representation, producing a new code representation that preserves high-level programming logic (in the form of control- and data-flows) together with the natural language information of a program. Our evaluation studies the top 10 most common C/C++ vulnerabilities during the past 3 years. We have conducted our experiments using 105,428 real-world programs by comparing our approach with four well-known traditional static vulnerability detectors and three state-of-the-art deep-learning-based approaches. The experimental results demonstrate the effectiveness of our research and have shed light on the promising direction of combining program analysis with deep learning techniques to address the general static code analysis challenges.
Xiao Cheng 0002, Haoyu Wang 0001, Jiayi Hua, Guoai Xu, Yulei Sui
ACM Trans. Softw. Eng. Methodol.4
2021 Provably Secure ECC-Based Three-Factor Authentication Scheme for Mobile Cloud Computing with Offline Registration Centre
abstract
Mobile cloud computing (MCC) aims at solving the resource constrain problem of smart mobile devices. It has deeply affected the way modern humans live and work. In MCC, the authentication scheme is indispensable to prevent illegal attacks and privacy breaches. In this paper, we reveal that a recently proposed two‐factor authentication scheme for MCC has limitations like stolen‐verifier attack and denial of service attack. In addition, its single‐server architecture is not applicable to MCC. To enhance the security, we present a provably secure three‐factor authentication scheme using the elliptic curve cryptosystem (ECC). It has the merit that the user only needs to register once to access multiple servers with a pair of public and private key, and the registration center is offline in the authentication phase. Security analysis demonstrates that our scheme is immune to known attacks and provides user friendliness. Finally, performance comparisons indicate that our scheme has better security attributes and low computing and communication overheads, and it is more applicable to MCC.
Guoai Xu
Wirel. Commun. Mob. Comput.3
2020 Dissecting Mobile Offerwall Advertisements: An Explorative Study
abstract
Mobile advertising has become the most popular monetizing way in the Android app ecosystem. Offerwall, as a new form of mobile ads, has been widely adopted by apps, and a number of ad networks have provided such services. Although new to the ecosystem, offerwall ads have been criticized for being aggressive, and the contents disseminated are prone to security issues. However, to date, our community has not proposed any studies to dissect such issues related to offerwall ads. To this end, we present the first work to fill this gap. Specifically, we first develop a robust approach to identify apps that have embedded with offerwall ads. Then, we apply the tool to 10K apps and experimentally discover 312 offerwall apps. We go one step further to characterize them from several aspects, including security issues. Our observation reveals that offerwall ads could indeed be manipulated by hackers to fulfill malicious purposes.
Yangyu Hu, Li Li 0029, Guoai Xu, Zhihui Han, Haoyu Wang 0001
QRS6
2020 Sensitive Information Detection based on Convolution Neural Network and Bi-directional LSTM
abstract
Electronic documents can carry lots of information and are widely used in daily lives. It will cause substantial economic losses to individual users, enterprises, and governments when the documents containing sensitive information are leaked. How to detect sensitive information to prevent data leakage is still a challenge in the field of information security. This paper mainly focuses on the detection of unstructured documents containing sensitive information. Governments, military, and other institutions can actively mark whether the electronic documents contain sensitive information according to the detection results. We propose a reliable method to detect sensitive electronic documents automatically and compare it with other basic methods. The algorithm structure can extract the characteristics of the data more comprehensively to obtain better detection results. Our model outperformed the other models with 93.44 % accuracy. Our model can also reduce the time cost, which is beneficial for realistic production.
Guosheng Xu 0001, Guoai Xu
TrustCom3
2020 Mobile App Squatting
abstract
Domain squatting, the adversarial tactic where attackers register domain names that mimic popular ones, has been observed for decades. However, there has been growing anecdotal evidence that this style of attack has spread to other domains. In this paper, we explore the presence of squatting attacks in the mobile app ecosystem. In “App Squatting”, attackers release apps with identifiers (e.g., app name or package name) that are confusingly similar to those of popular apps or well-known Internet brands. This paper presents the first in-depth measurement study of app squatting showing its prevalence and implications. We first identify 11 common deformation approaches of app squatters and propose “AppCrazy”, a tool for automatically generating variations of app identifiers. We have applied AppCrazy to the top-500 most popular apps in Google Play, generating 224,322 deformation keywords which we then use to test for app squatters on popular markets. Through this, we confirm the scale of the problem, identifying 10,553 squatting apps (an average of over 20 squatting apps for each legitimate one). Our investigation reveals that more than 51% of the squatting apps are malicious, with some being extremely popular (up to 10 million downloads). Meanwhile, we also find that mobile app markets have not been successful in identifying and eliminating squatting apps. Our findings demonstrate the urgency to identify and prevent app squatting abuses. To this end, we have publicly released all the identified squatting apps, as well as our tool AppCrazy.
Yangyu Hu, Haoyu Wang 0001, Li Li 0029, Gareth Tyson, Ignacio Castro, Yao Guo 0001, Lei Wu 0012, Guoai Xu
WWW9
2020 Characterizing cryptocurrency exchange scams
Pengcheng Xia 0001, Haoyu Wang 0001, Ru Ji, Bingyu Gao, Lei Wu 0012, Xiapu Luo, Guoai Xu
Comput. Secur.8
2020 Game Theoretical Method for Anomaly-Based Intrusion Detection
abstract
In this paper, the game theoretical analysis method is presented to provide optimal strategies for anomaly-based intrusion detection systems (A-IDS). A two-stage game model is established to represent the interactions between the attackers and defenders. In the first stage, the players decide to do actions or keep silence, and in the second stage, attack intensity and detection threshold are considered as two important strategic variables for the attackers and defenders, respectively. The existence, uniqueness, and explicit computation of the Nash equilibrium are analyzed and obtained by considering six different scenarios, from which the optimal detection and attack actions are provided. Numerical examples are provided to validate our theoretical results.
Shengwei Xu, Guoai Xu, Yongfeng Yin, Miao Zhang 0011
Secur. Commun. Networks3
2020 CrowdBLPS: A Blockchain-Based Location-Privacy-Preserving Mobile Crowdsensing System
abstract
With the popularization of intelligent terminals, especially current trends, such as “Industrie 4.0” and the Internet of Things, mobile crowdsensing is becoming one of the promising applications built on smart devices in mobile networks. However, the existing mobile crowdsensing models are mostly based on a centralized platform, which is not fully trusted in reality and results in the existence of fraud and other security problems. Furthermore, the data quality collected through crowdsensing is varied, and the location privacy is difficult to guarantee, especially at the worker selection stage. To solve these two problems, an effective blockchain-based location-privacy-preserving crowdsensing model, CrowdBLPS, is proposed in this article. First, the idea of a blockchain is introduced into this model. The decentralized structure and the consensus approach are applied to realize the nonrepudiation and nontampering of information. Second, to improve the data sensing quality and protect worker privacy, a two-stage approach, including the preregistration stage and the final selection stage, is proposed. Finally, we further implement a prototype on the Ethereum public testing network, and the experimental results show the feasibility, availability, and reliability of CrowdBLPS.
Shihong Zou, Jinwen Xi, Honggang Wang 0001, Guoai Xu
IEEE Trans. Ind. Informatics4
2020 A Robust IoT-Based Three-Factor Authentication Scheme for Cloud Computing Resistant to Session Key Exposure
abstract
With the development of Internet of Things (IoT) technologies, Internet-enabled devices have been widely used in our daily lives. As a new service paradigm, cloud computing aims at solving the resource-constrained problem of Internet-enabled devices. It is playing an increasingly important role in resource sharing. Due to the complexity and openness of wireless networks, the authentication protocol is crucial for secure communication and user privacy protection. In this paper, we discuss the limitations of a recently introduced IoT-based authentication scheme for cloud computing. Furthermore, we present an enhanced three-factor authentication scheme using chaotic maps. The session key is established based on Chebyshev chaotic-based Diffie–Hellman key exchange. In addition, the session key involves a long-term secret. It ensures that our scheme is secure against all the possible session key exposure attacks. Besides, our scheme can effectively update user password locally. Burrows–Abadi–Needham logic proof confirms that our scheme provides mutual authentication and session key agreement. The formal analysis under random oracle model proves the semantic security of our scheme. The informal analysis shows that our scheme is immune to diverse attacks and has desired features such as three-factor secrecy. Finally, the performance comparisons demonstrate that our scheme provides optimal security features with an acceptable computation and communication overheads.
Guosheng Xu 0001, Guoai Xu, Yuejie Wang, Junhao Peng
Wirel. Commun. Mob. Comput.3
2019 Static Detection of Control-Flow-Related Vulnerabilities Using Graph Embedding
abstract
Static vulnerability detection has shown its effectiveness in detecting well-defined low-level memory errors. However, high-level control-flow related (CFR) vulnerabilities, such as insufficient control flow management (CWE-691), business logic errors (CWE-840), and program behavioral problems (CWE-438), which are often caused by a wide variety of bad programming practices, posing a great challenge for existing general static analysis solutions. This paper presents a new deep-learning-based graph embedding approach to accurate detection of CFR vulnerabilities. Our approach makes a new attempt by applying a recent graph convolutional network to embed code fragments in a compact and low-dimensional representation that preserves high-level control-flow information of a vulnerable program. We have conducted our experiments using 8,368 real-world vulnerable programs by comparing our approach with several traditional static vulnerability detectors and state-of-the-art machine-learning-based approaches. The experimental results show the effectiveness of our approach in terms of both accuracy and recall. Our research has shed light on the promising direction of combining program analysis with deep learning techniques to address the general static analysis challenges.
Xiao Cheng 0002, Haoyu Wang 0001, Jiayi Hua, Miao Zhang 0011, Guoai Xu, Yulei Sui
ICECCS5
2019 DaPanda: Detecting Aggressive Push Notifications in Android Apps
abstract
Mobile push notifications have been widely used in mobile platforms to deliver all sorts of information to app users. Although it offers great convenience for both app developers and mobile users, this feature was frequently reported to serve malicious and aggressive purposes, such as delivering annoying push notification advertisement. However, to the best of our knowledge, this problem has not been studied by our research community so far. To fill the void, this paper presents the first study to detect aggressive push notifications and further characterize them in the global mobile app ecosystem on a large scale. To this end, we first provide a taxonomy of mobile push notifications and identify the aggressive ones using a crowdsourcing-based method. Then we propose sc DaPanda, a novel hybrid approach, aiming at automatically detecting aggressive push notifications in Android apps. sc DaPanda leverages a guided testing approach to systematically trigger and record push notifications. By instrumenting the Android framework, sc DaPanda further collects all notification-relevant runtime information to flag the aggressive ones. Our experimental results show that sc DaPanda is capable of detecting different types of aggressive push notifications effectively in an automated way. By applying sc DaPanda to 20,000 Android apps from different app markets, it yields over 1,000 aggressive notifications, which have been further confirmed as true positives. Our in-depth analysis further reveals that aggressive notifications are prevalent across different markets and could be manifested in all the phases in the lifecycle of push notifications. It is hence urgent for our community to take actions to detect and mitigate apps involving aggressive push notifications.
Tianming Liu 0002, Haoyu Wang 0001, Li Li 0029, Guangdong Bai, Yao Guo 0001, Guoai Xu
ASE6
2019 Want to Earn a Few Extra Bucks? A First Look at Money-Making Apps
abstract
Have you ever thought of earning profits from the apps that you are using on your mobile device? It is actually achievable thanks to many so-called money-making apps, which pay app users to complete tasks such as installing another app or clicking an advertisement. To the best of our knowledge, no existing studies have investigated the characteristics of moneymaking apps. To this end, we conduct the first exploratory study to understand the features and implications of money-making apps. We first propose a semi-automated approach aiming to harvest money-making apps from Google Play and alternative app markets. Then we create a taxonomy to classify them into five categories and perform an empirical study from different aspects. Our study reveals several interesting observations: (1) moneymaking apps have become the target of malicious developers, as we found many of them expose mobile users to serious privacy and security risks. Roughly 26% of the studied apps are potentially malicious. (2) these apps have attracted millions of users, however, many users complain that they are cheated by these apps. We also revealed that ranking fraud techniques are widely used in these apps to promote the ranking of apps inside app markets. (3) these apps usually spread inappropriate and malicious contents, while unsuspicious users could get infected. Our study demonstrates the emergency for detecting and regulating this kind of apps and protect mobile users.
Yangyu Hu, Haoyu Wang 0001, Li Li 0029, Yao Guo 0001, Guoai Xu
SANER5
2019 A Secure and Efficient ECC-Based Anonymous Authentication Protocol
abstract
Nowadays, remote user authentication protocol plays a great role in ensuring the security of data transmission and protecting the privacy of users for various network services. In this study, we discover two recently introduced anonymous authentication schemes are not as secure as they claimed, by demonstrating they suffer from offline password guessing attack, desynchronization attack, session key disclosure attack, failure to achieve user anonymity, or forward secrecy. Besides, we reveal two environment-specific authentication schemes have weaknesses like impersonation attack. To eliminate the security vulnerabilities of existing schemes, we propose an improved authentication scheme based on elliptic curve cryptosystem. We use BAN logic and heuristic analysis to prove our scheme provides perfect security attributes and is resistant to known attacks. In addition, the security and performance comparison show that our scheme is superior with better security and low computation and communication cost.
Guoai Xu, Lize Gu
Secur. Commun. Networks2
2019 A Provably Secure Biometrics-Based Authentication Scheme for Multiserver Environment
abstract
With the rapid development of mobile services, multiserver authentication protocol with its high efficiency has emerged as an indispensable security mechanism for mobile services. Recently, Ali et al. introduced a biometric-based multiserver authentication scheme and claimed the scheme is resistant to various attacks. However, after a careful examination, we find that Ali et al.’s scheme is vulnerable to various security attacks, such as user impersonation attack, server impersonation attack, privileged insider attack, denial of service attack, fails to provide forward secrecy and three-factor secrecy. To overcome these weaknesses, we propose an improved biometric-based multiserver authentication scheme using elliptic curve cryptosystem. Formal security analysis under the random oracle model proves that our scheme is provably secure. Furthermore, BAN (Burrows-Abadi-Needham) logic analysis demonstrates our scheme achieves mutual authentication and session key agreement. In addition, the informal analysis proves that our scheme is secure against all current known attacks and achieves desirable features. Besides, the performance and security comparison shows that our scheme is superior to related schemes.
Guoai Xu, Chenyu Wang 0002, Junhao Peng
Secur. Commun. Networks2
2018 Towards Light-Weight Deep Learning Based Malware Detection
abstract
The explosive amount of malware continues threating the security of operating systems and networks. Traditional malware detection approaches fail to meet the requirements of detecting polymorphic and new samples. Existing neural network based detection approaches performs better, but consuming much more time in both feature extraction and training. In this paper, we propose a light-weight PC malware detection system which is based on deep convolutional neural network (CNN). The raw inputs of our system are sequences of grouped instructions, which were generated by our Instruction Analyzer in according to different functionalities of the instructions. The network will automatically learn features of malware from the grouped instruction sequences. The experiment results suggest that in a large dataset which contains roughly 70,000 samples, our detection system can achieve an overall accuracy of 95\%. The training time of our system with single convolutional layer was only about 10 hours, which is one order of magnitude less than traditional methods.
Zeliang Kan, Haoyu Wang 0001, Guoai Xu, Yao Guo 0001, Xiangqun Chen
COMPSAC (1)3
2018 Beyond Google Play: A Large-Scale Comparative Study of Chinese Android App Markets
Haoyu Wang 0001, Zhe Liu 0001, Jingyue Liang, Narseo Vallina-Rodriguez, Yao Guo 0001, Li Li 0029, Juan Tapiador, Jingcun Cao, Guoai Xu
Internet Measurement Conference9
2018 Why are Android apps removed from Google Play?: a large-scale empirical study
abstract
To ensure the quality and trustworthiness of the apps within its app market (i.e., Google Play), Google has released a series of policies to regulate app developers. As a result, policy-violating apps (e.g., malware, low-quality apps, etc.) have been removed by Google Play periodically. In reality, we have found that the number of removed apps are actually much more than what we have expected, as almost half of all the apps have been removed or replaced from Google Play during a two year period from 2015 to 2017. However, despite the significant number of removed apps, there are almost no study on the characterization of these removed apps. To this end, this paper takes the first step to understand why Android apps are removed from Google Play, aiming at observing promising insights for both market maintainers and app developers towards building a better app ecosystem. By leveraging two app sets crawled from Google Play in 2015 (over 1.5 million) and 2017 (over 2.1 million), we have identified a set of over 790K removed apps, which are then thoroughly investigated in various aspects. The experimental results have revealed various interesting findings, as well as insights for future research directions.
Haoyu Wang 0001, Li Li 0029, Yao Guo 0001, Guoai Xu
MSR5
2018 Re-checking App Behavior against App Description in the Context of Third-party Libraries
abstract
Recent research suggested promising approaches that identify potential malware by checking the inconsistence between app description and actual behavior of the app.However, state-of-the-art approaches have ignored the impact of thirdparty libraries (TPLs) when detecting outliers, which could affect the detection results greatly in two folds.On one hand, most Android apps would not list the functionality of TPLs in app description, which could cause false positives, as many apps that use TPLs will be identified as outliers.On the other hand, it is important to separate TPLs from custom code when analyzing the sensitive behaviors, otherwise the malicious behaviors of custom code will be obscured by TPLs.In this paper, we revisit the study of checking app behavior against app description in the context of TPLs.Experiment results on more than 400K Android apps suggest that more than 54% of apps are no longer identified as outliers after filtering TPLs, and we could identify roughly 50% of new outliers.Furthermore, removing the impact of TPLs could help to identify malware and pinpoint the malicious behavior of custom code.Out results shed a light on applying the TPL analysis to enhance a variety of mobile app analysis tasks.
Chengpeng Zhang, Haoyu Wang 0001, Yao Guo 0001, Guoai Xu
SEKE5
2018 FraudDroid: automated ad fraud detection for Android apps
abstract
Although mobile ad frauds have been widespread, state-of-the-art approaches in the literature have mainly focused on detecting the so-called static placement frauds, where only a single UI state is involved and can be identified based on static information such as the size or location of ad views. Other types of fraud exist that involve multiple UI states and are performed dynamically while users interact with the app. Such dynamic interaction frauds, although now widely spread in apps, have not yet been explored nor addressed in the literature. In this work, we investigate a wide range of mobile ad frauds to provide a comprehensive taxonomy to the research community. We then propose, FraudDroid, a novel hybrid approach to detect ad frauds in mobile Android apps. FraudDroid analyses apps dynamically to build UI state transition graphs and collects their associated runtime network traffics, which are then leveraged to check against a set of heuristic-based rules for identifying ad fraudulent behaviours. We show empirically that FraudDroid detects ad frauds with a high precision (∼ 93%) and recall (∼ 92%). Experimental results further show that FraudDroid is capable of detecting ad frauds across the spectrum of fraud types. By analysing 12,000 ad-supported Android apps, FraudDroid identified 335 cases of fraud associated with 20 ad networks that are further confirmed to be true positive results and are shared with our fellow researchers to promote advanced ad fraud detection.
Feng Dong 0008, Haoyu Wang 0001, Li Li 0029, Yao Guo 0001, Tegawendé F. Bissyandé, Tianming Liu 0002, Guoai Xu, Jacques Klein
ESEC/SIGSOFT FSE7
2018 A Secure and Anonymous Two-Factor Authentication Protocol in Multiserver Environment
abstract
With the great development of network technology, the multiserver system gets widely used in providing various of services. And the two-factor authentication protocols in multiserver system attract more and more attention. Recently, there are two new schemes for multiserver environment which claimed to be secure against the known attacks. However, after a scrutinization of these two schemes, we found that (1) their description of the adversary’s abilities is inaccurate; (2) their schemes suffer from many attacks. Thus, firstly, we corrected their description on the adversary capacities to introduce a widely accepted adversary model and then summarized fourteen security requirements of multiserver based on the works of pioneer contributors. Secondly, we revealed that one of the two schemes fails to preserve forward secrecy and user anonymity and cannot resist stolen-verifier attack and off-line dictionary attack and so forth and also demonstrated that another scheme fails to preserve forward secrecy and user anonymity and is not secure to insider attack and off-line dictionary attack, and so forth. Finally, we designed an enhanced scheme to overcome these identified weaknesses, proved its security via BAN logic and heuristic analysis, and then compared it with other relevant schemes. The comparison results showed the superiority of our scheme.
Chenyu Wang 0002, Guoai Xu, Wenting Li 0002
Secur. Commun. Networks2
2018 An Enhanced User Authentication Protocol Based on Elliptic Curve Cryptosystem in Cloud Computing Environment
abstract
With the popularity of cloud computing, information security issues in the cloud environment are becoming more and more prominent. As the first line of defense to ensure cloud computing security, user authentication has attracted extensive attention. Though considerable efforts have been paid for a secure and practical authentication scheme in cloud computing environment, most attempts ended in failure. The design of a secure and efficient user authentication scheme for cloud computing remains a challenge on the one hand and user’s smart card or mobile devices are of limited resource; on the other hand, with the combination of cloud computing and the Internet of Things, applications in cloud environments often need to meet various security requirements and are vulnerable to more attacks. In 2018, Amin et al. proposed an enhanced user authentication scheme in cloud computing, hoping to overcome the identified security flaws of two previous schemes. However, after a scrutinization of their scheme, we revealed that it still suffers from the same attacks (such as no user anonymity, no forward secrecy, and being vulnerable to offline dictionary attack) as the two schemes they compromised. Consequently, we take the scheme of Amin et al. (2018) as a study case, we discussed the inherent reason and the corresponding solutions to authentication schemes for cloud computing environment in detail. Next, we not only proposed an enhanced secure and efficient scheme, but also explained the design rationales for a secure cloud environment protocol. Finally, we applied BAN logic and heuristic analysis to show the security of the protocol and compared our scheme with related schemes. The results manifest the superiority of our scheme.
Chenyu Wang 0002, Guoai Xu, Ping Wang 0003
Wirel. Commun. Mob. Comput.5
2017 An Explorative Study of the Mobile App Ecosystem from App Developers' Perspective
abstract
With the prevalence of smartphones, app markets such as Apple App Store and Google Play has become the center stage in the mobile app ecosystem, with millions of apps developed by tens of thousands of app developers in each major market. This paper presents a study of the mobile app ecosystem from the perspective of app developers. Based on over one million Android apps and 320,000 developers from Google Play, we analyzed the Android app ecosystem from different aspects. Our analysis shows that while over half of the developers have released only one app in the market, many of them have released hundreds of apps. We classified developers into different groups based on the number of apps they have released, and compared their characteristics. Specially, we have analyzed the group of aggressive developers who have released more than 50 apps, trying to understand how and why they create so many apps. We also investigated the privacy behaviors of app developers, showing that some developers have a habit of producing apps with low privacy ratings. Our study shows that understanding the behavior of mobile developers can be helpful to not only other app developers, but also to app markets and mobile users.
Haoyu Wang 0001, Zhe Liu 0001, Yao Guo 0001, Xiangqun Chen, Miao Zhang 0011, Guoai Xu, Jason I. Hong
WWW6
2017 Cryptanalysis of Three Password-Based Remote User Authentication Schemes with Non-Tamper-Resistant Smart Card
abstract
Remote user authentication is the first step to guarantee the security of online services. Online services grow rapidly and numerous remote user authentication schemes were proposed with high capability and efficiency. Recently, there are three new improved remote user authentication schemes which claim to be resistant to various attacks. Unfortunately, according to our analysis, these schemes all fail to achieve some critical security goals. This paper demonstrates that they all suffer from offline dictionary attack or fail to achieve forward secrecy and user anonymity. It is worth mentioning that we divide offline dictionary attacks into two categories: (1) the ones using the verification from smart cards and (2) the ones using the verification from the open channel. The second is more complicated and intractable than the first type. Such distinction benefits the exploration of better design principles. We also discuss some practical solutions to the two kinds of attacks, respectively. Furthermore, we proposed a reference model to deal with the first kind of attack and proved its effectiveness by taking one of our cryptanalysis schemes as an example.
Chenyu Wang 0002, Guoai Xu
Secur. Commun. Networks2
2017 Differential Fault Attack on ITUbee Block Cipher
abstract
Differential Fault Attack (DFA) is a powerful cryptanalytic technique to retrieve secret keys by exploiting the faulty ciphertexts generated during encryption procedure. This article proposes a novel DFA attack that is effective on ITUbee, a software-oriented block cipher for resource-constrained devices. Different from other DFA, our attack makes use of not only faulty values, but also differences between fault-free intermediate values corresponding to 2 plaintexts, which combine traditional differential analysis with DFA. The possible injection positions with different number of faults are discussed. The most efficient attack takes 2 25 round function operations with 4 faults, which is achieved in a few seconds on a PC.
Shan Fu, Guoai Xu, Juan Pan, Zongyue Wang, An Wang 0001
ACM Trans. Embed. Comput. Syst.2