Xiaobo Ma 0001

dblp:18/10062-1 · DBLP profile ↗
← Back
51ranked-venue papers
11as first author
28since 2021 · last 2026
0000-0002-0934-5035ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 22 · 7 first-author · 13 since 2021Security and privacy · 19 · 3 first-author · 9 since 2021Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NEST: A Node-Interactive Generative Emulation Framework for Synthetic Traffic Generation
Jianfeng Li 0006, Jian Qu, Xiaobo Ma 0001
INFOCOM5
2026 $\mathbb {ABC}$ABC-$ {\mathbb{Channel}}$Channel: An Advanced Blockchain-Based Covert Channel
abstract
Establishing efficient and robust covert channels is crucial for secure communication within insecure network environments. With its inherent benefits of decentralization and anonymization, blockchain has gained considerable attention in developing covert channels. To guarantee a highly secure covert channel, channel negotiation should be contactlessbeforethe communication, carrier transaction features must be indistinguishable from normal transactionsduringthe communication, and communication identities must be untraceableafterthe communication. Such a full-lifecycle covert channel is indispensable to defend against a versatile adversary who intercepts two communicating parties comprehensively (e.g., on-chain and off-chain). Unfortunately, it has not been thoroughly investigated in the literature. We make the first effort to achieve a full-lifecycle covert channel, a novel blockchain-based covert channel namedABC-Channel. We tackle a series of challenges, such as off-chain contact dependency, increased masquerading difficulties as growing transaction volume, and time-evolving, communicable yet untraceable identities, to achieve contactless channel negotiation, indistinguishable transaction features, and untraceable communication identities, respectively. We develop a working prototype to validateABC-Channeland conduct extensive tests on the Bitcoin testnet. The experimental results demonstrate thatABC-Channelachieves substantially secure covert capabilities. In comparison to existing methods, it also exhibits state-of-the-art transmission efficiency.
Xiaobo Ma 0001, Pengyu Pan, Jianfeng Li 0006, Wei Wang 0012, Weizhi Meng 0001, Xiaohong Guan
IEEE Trans. Dependable Secur. Comput.1
2026 AirCloak: An App-Transparent Traffic Cloaking Middleware Against Wireless Fingerprinting Attack
Huafeng Bian, Jianfeng Li 0006, Haodan Luo, Xiaobo Ma 0001, Zhenhua Li 0001, Jigang Wang, Wei Wang 0012
IEEE Trans. Netw.5
2026 Enabling Entangled Cache Probing for Remotely Reconstructing DNS Query Dynamics
Jianfeng Li 0006, Wen Li 0007, Qinyu Liu, Xiaobo Ma 0001, Wei Wang 0012, Xiapu Luo, Xiaohong Guan
IEEE Trans. Netw.7
2025 Sylva: Tailoring Personalized Adversarial Defense in Pre-trained Models via Collaborative Fine-tuning
abstract
The growing adoption of large pre-trained models in edge computing has made deploying model inference on mobile clients both practical and popular. These devices are inherently vulnerable to direct adversarial attacks, which pose a substantial threat to the robustness and security of deployed models. Federated adversarial training (FAT) has emerged as an effective solution to enhance model robustness while preserving client privacy. However, FAT frequently produces a generalized global model, which struggles to address the diverse and heterogeneous data distributions across clients, resulting in insufficiently personalized performance, while also encountering substantial communication challenges during the training process. In this paper, we propose Sylva, a personalized collaborative adversarial training framework designed to deliver customized defense models for each client through a two-phase process. In Phase 1, Sylva employs LoRA for local adversarial fine-tuning, enabling clients to personalize model robustness while drastically reducing communication costs by uploading only LoRA parameters during federated aggregation. In Phase 2, a game-based layer selection strategy is introduced to enhance accuracy on benign data, further refining the personalized model. This approach ensures that each client receives a tailored defense model that balances robustness and accuracy effectively. Extensive experiments on benchmark datasets demonstrate that Sylva can achieve up to 50× improvements in communication efficiency compared to state-of-the-art algorithms, while achieving up to 29.5% and 50.4% enhancements in adversarial robustness and benign accuracy, respectively.
Tianyu Qi, Lei Xue 0001, Yufeng Zhan, Xiaobo Ma 0001
CCS4
2025 Cross-Environmental Website Fingerprinting
Jianfeng Li 0006, Xiaobo Ma 0001, Xiapu Luo, Xiaohong Guan
INFOCOM7
2025 SoFi: Spoofing OS Fingerprints Against Network Reconnaissance
abstract
Fingerprinting is a network reconnaissance technique utilized for gathering information about online computing systems, including operation systems and applications. Unfortunately, attackers typically leverage fingerprinting techniques to locate, enumerate, and subsequently target vulnerable systems, which is the first primary stage of a cyber attack. In this work, we explore the susceptibility of machine learning (ML)-based classifiers to misclassification, where a slight perturbation in the packet is included to spoof OS fingerprints. We propose SOFI (Spoof OS Fingerprints), an adversarial example generation algorithm under TCP/IP specification constraints, to create effective perturbations in a packet for deceiving an OS fingerprint. Specifically, SOFI has three major technical innovations: (1) it is the first to utilize adversarial examples to automatically perturb fingerprinting techniques; (2) it complies with constraints and integrity of network packets; (3) it achieves a high success rate in spoofing OS fingerprints. We validate the effectiveness of adversarial packets against active and passive OS fingerprints, verifying the transferability and robustness of SOFI. Comprehensive experimental results demonstrate that SOFI automatically identifies applicable and available OS fingerprint features, unlike existing tools relying on expert knowledge.
Haocong Li, Wei Wang 0012, Haining Wang 0001, Xiaobo Ma 0001, Shouling Ji, Qiang Li 0007
IEEE Trans. Inf. Forensics Secur.5
2025 Driving State-Aware Anomaly Detection for Autonomous Vehicles
abstract
With the increasing popularity of autonomous driving systems (ADS) in autonomous vehicles (AV), in recent years, there have been many attacks targeting AVs and ADSs. Meanwhile, recent studies have attempted to improve the safety and security of AVs from different perspectives, and they mainly focus on the spoofing attacks against the sensors and the injection attacks against the vehicle chassis and actuators. However, direct attacks on ADSs (i.e., communication hijacking and malicious codes) remain inadequately addressed, and even worse, such attacks can cause AVs to make unsafe driving decisions rapidly. In this paper, we introduceDSAD, a driving state-aware anomaly detection framework designed to enhance AV safety and security by identifying ADS attacks, such as communication hijacking and malicious codes, through chassis states. First,DSADmodels ADS operations (i.e., driving states) as a two-layer state machine, utilizing real-time chassis data to infer driving states and detect anomalies in ADS outputs. This reduces false positives and negatives by aligning detection with the diverse operational modes of AVs. To achieve this, we develop a prototype system,DSAD, incorporating a Detection Policy Update mechanism that dynamically adjusts detection policies based on the vehicle’s driving states, such as lane changing and obstacle avoidance. Second,DSADconsiders both collision avoidance and control stability, addressing potential conflicts through hard and soft requirements. Furthermore,DSADintegrates a fault handling module compatible with existing autonomous driving fault handling mechanisms, ensuring timely response to detected anomalies. We develop a prototype anomaly detection system calledDSADand deploy it on four ADSs. We evaluateDSADusing various attack scenarios, and the results show thatDSADcan identify over 90% of attacks on ADSs.
Lei Xue 0001, Xiapu Luo, Xiaobo Ma 0001, Guofei Gu
IEEE Trans. Inf. Forensics Secur.4
2025 Vehicular Intrusion Detection System for Controller Area Network: A Comprehensive Survey and Evaluation
abstract
The progress of automotive technologies has made cybersecurity a crucial focus, leading to various cyber attacks. These attacks primarily target the Controller Area Network (CAN) and specialized Electronic Control Units (ECUs). In order to mitigate these attacks and bolster the security of vehicular systems, numerous defense solutions have been proposed. These solutions aim to detect diverse forms of vehicular attacks. However, the practical implementation of these solutions still presents certain limitations and challenges. In light of these circumstances, this paper undertakes a thorough examination of existing vehicular attacks and defense strategies employed against the CAN and ECUs. The objective is to provide valuable insights and inform the future design of Vehicular Intrusion Detection Systems (VIDS). The findings of our investigation reveal that the examined VIDS primarily concentrate on particular categories of attacks, neglecting the broader spectrum of potential threats. Moreover, we provide a comprehensive overview of the significant challenges encountered in implementing a robust and feasible VIDS. Additionally, we put forth several defense recommendations based on our study findings, aiming to inform and guide the future design of VIDS in the context of vehicular security.
Lei Xue 0001, Sishan Wang, Xiapu Luo, Kaifa Zhao, Pengfei Jing, Xiaobo Ma 0001, Yajuan Tang, Haiying Zhou
IEEE Trans. Intell. Transp. Syst.7
2024 DNSScope: Fine-Grained DNS Cache Probing for Remote Network Activity Characterization
abstract
The domain name system (DNS) is indispensable to nearly every Internet service. It has been extensively utilized for network activity characterization in passive and active approaches. Compared to the passive approach, active DNS cache probing is privacy-preserving and low-cost, enabling worldwide characterization of remote network activities in different networks. Unfortunately, existing probing-based methods are too coarse-grained to characterize the time-varying features of network activities, substantially limiting their applications in time-sensitive tasks. In this paper, we advance DNSScope, a fine-grained DNS cache probing framework by tackling three challenges: sample sparsity, observational distortion, and cache entanglement. DNSScope synthesizes statistical learning and self-supervised transfer learning to achieve time-varying characterization. Extensive evaluations demonstrate that it can accurately estimate the time-varying DNS query arrival rates on recursive DNS resolvers. Its average mean absolute error is 0.124, as low as one-sixth that of the baseline methods.
Jianfeng Li 0006, Xiaobo Ma 0001, Jian Qu, Xiapu Luo, Xiaohong Guan
INFOCOM3
2024 FedComm: A Privacy-Enhanced and Efficient Authentication Protocol for Federated Learning in Vehicular Ad-Hoc Networks
abstract
In vehicular ad-hoc networks (VANET), federated learning enables vehicles to collaboratively train a global model for intelligent transportation without sharing their local data. However, due to dynamic network structure and unreliable wireless communication of VANET, various potential risks (e.g., identity privacy leakage, data privacy inference, model integrity compromise, and data manipulation) undermine the trustworthiness of intermediate model parameters necessary for building the global model. While existing cryptography techniques and differential privacy provide provable security paradigms, the practicality of secure federated learning in VANET is hindered in terms of training efficiency and model performance. Therefore, developing a secure and efficient federated learning in VANET remains a challenge. In this work, we propose a privacy-enhanced and efficient authentication protocol for federated learning in VANET, called FedComm. Unlike existing solutions, FedComm addresses the above challenge through user anonymity. First, FedComm enables vehicles to participate in training with unlinkable pseudonyms, ensuring both privacy preservation and efficient collaboration. Second, FedComm incorporates an efficient authentication protocol to guarantee the authenticity and integrity of model parameters originated from anonymous vehicles. Finally, FedComm accurately identifies and completely eliminates malicious vehicles in anonymous communication. Security analysis and verification with ProVerif demonstrate that FedComm enhances privacy and reliability of intermediate model parameters. Experimental results show that FedComm reduces the overhead of proof generation and verification by 67.38% and 67.39%, respectively, compared with the state-of-the-art authentication protocols used in federated learning.
Jiqiang Liu, Bin Wang 0062, Wei Wang 0012, Bin Wang 0066, Tao Li 0022, Xiaobo Ma 0001, Witold Pedrycz
IEEE Trans. Inf. Forensics Secur.7
2024 Automating Cloud Deployment for Real-Time Online Foundation Model Inference
abstract
Deep neural network (DNN) foundation models are currently exhibiting high prediction accuracy and strong adaptability to broad tasks with remarkably large model scales. They are increasingly becoming the backend support of DNN-driven real-time online services, e.g., Siri and Instagram. Such services require low-latency and cost-efficiency for quality-of-service and commercial competitiveness. When deployed in a cloud environment, these services call for an appropriate selection of cloud configurations (i.e., specific types of VM instances), as well as a considerate device placement plan that places the operations of the model to multiple GPUs via model parallelism for cost-efficiency. Currently, the deployment mainly relies on service providers’ manual efforts, which is not only onerous but also far from satisfactory oftentimes due to the huge joint search space of cloud configurations and device placement plans (for a same service, a poor deployment can incur significantly more costs by tens of times). In this paper, we attempt to efficiently automate the cloud deployment for real-time foundation model inference with minimum costs under the constraint of acceptably low latency. This attempt is enabled by 1) jointly leveraging the Bayesian Optimization and Deep Reinforcement Learning to adaptively unearth the (nearly) optimal cloud configuration and device placement with limited search time, and 2) enhancing the cost-efficiency of the deployment based on the probing-informed block multiplexing mechanism and Tensor Algebra SuperOptimizer. We implement a prototype system based on TensorFlow, conduct extensive experiments on top of Microsoft Azure, and demonstrate the generality and scalability of our solution. Results show that for lightweight DNN models and foundation models, our solution essentially saves inference costs by up to 15% and 47% with 57% and 38% lower search overheads respectively, compared with non-trivial baselines.
Yang Li 0092, Zhenhua Li 0001, Zhenhua Han, Quanlu Zhang, Xiaobo Ma 0001
IEEE/ACM Trans. Netw.5
2024 Robust App Fingerprinting Over the Air
abstract
Mobile apps have significantly transformed various aspects of modern life, leading to growing concerns about privacy risks. Despite widespread encrypted communication, app fingerprinting (AF) attacks threaten user privacy substantially. However, existing AF attacks, when targeted at wireless traffic, face four fundamental challenges, namely 1) sample inseparability; 2) app multiplexing; 3) signal attenuation; and 4) open-world recognition. In this paper, we advance a novel AF attack, dubbed PacketPrint, to recognize app user activities over the air in an open-world setting. We introduce two novel models, i.e., sequential XGBoost and hierarchical bag-of-words model, to tackle sample inseparability and enhance robustness against noise packets arising from app multiplexing. We also propose the environment-aware model enhancement to bolster PacketPrint’s robustness in handling packet loss at the sniffer caused by signal attenuation. We conduct extensive experiments to evaluate the proposed attack in a series of challenging scenarios, including 1) open-world setting; 2) simultaneous use of different apps; 3) severe packet loss at the sniffer; and 4) cross-dataset recognition. The experimental results show that PacketPrint can accurately recognize app user activities. It achieves the average F1-score 0.947 for open-world app recognition and the average F1-score 0.959 for in-app user action recognition.
Jianfeng Li 0006, Jian Qu, Shuohan Wu, Hao Zhou 0043, Xiaobo Ma 0001, Ting Wang 0006, Xiapu Luo, Xiaohong Guan
IEEE/ACM Trans. Netw.7
2024 Website Fingerprinting on Encrypted Proxies: A Flow-Context-Aware Approach and Countermeasures
abstract
Website fingerprinting (WFP) could infer which websites a user is accessing via an encrypted proxy by passively inspecting the traffic characteristics of accessing different websites between the user and the proxy. Designing WFP attacks is crucial for understanding potential vulnerabilities of encrypted proxies, which guides the design of defensive measures against WFP. In this paper, we design a novel WFP attack against (popular) encrypted proxies that relay connections between the user and the proxy individually (e.g., Shadowsocks, V2Ray), and accordingly implement lightweight countermeasures to effectively defend against the attack. The attack features flow-context-aware and is both accurate and immediately deployable, because it fully considers the obstacle (dubbed training-testing asymmetry) that fundamentally limits the practicability of WFP and addresses the obstacle with built-in spatial-temporal flow correlation mechanism. We implement the countermeasure as middleboxes installed on both the client and server sides of encrypted proxies, without altering any existing infrastructures for compatibility. The middleboxes can obfuscate a website’s flow regularities across different visits. Large-scale experiments in real-world scenarios demonstrate that the WFP attack can generally achieve a detection rate above 98.8% with a false positive rate below 0.2%. The countermeasure forces the attack’s false positive rate to be above 0.2 and true positive rate to be below 0.9 with just five persistent TCP connections while introducing very limited bandwidth overhead (e.g., 0.49%) and almost-zero additional network latency.
Xiaobo Ma 0001, Jian Qu, Mawei Shi, Bingyu An, Jianfeng Li 0006, Xiapu Luo, Junjie Zhang 0004, Zhenhua Li 0001, Xiaohong Guan
IEEE/ACM Trans. Netw.1
2024 On Smartly Scanning of the Internet of Things
abstract
Cyber search engines, such as Shodan and Censys, have gained popularity due to their strong capability of indexing the Internet of Things (IoT). They actively scan and fingerprint IoT devices for unearthing IP-device mapping. Because of the large address space of the Internet and the mapping’s mutative nature, efficiently tracking the evolution of IP-device mapping with a limited budget of scans is essential for building timely cyber search engines. An intuitive solution is to use reinforcement learning to schedule more scans to networks with high churn rates of IP-device mapping. However, such an intuitive solution has never been systematically studied. In this paper, we take the first step toward demystifying this problem based on our experiences in maintaining a global IoT scanning platform. Inspired by the measurement study of large-scale real-world IoT scan records, we land reinforcement learning onto a system capable of smartly scanning IoT devices in a principled way. We disclose key parameters affecting the effectiveness of different scanning strategies, and real-world experiments demonstrate that our system can scan up to around 40 times as many IP-device mapping mutations as random/sequential scanning.
Jian Qu, Xiaobo Ma 0001, Wenmao Liu, Hongqing Sang, Jianfeng Li 0006, Lei Xue 0001, Xiapu Luo, Zhenhua Li 0001, Xiaohong Guan
IEEE/ACM Trans. Netw.2
2023 An Input-Agnostic Hierarchical Deep Learning Framework for Traffic Fingerprinting
Jian Qu, Xiaobo Ma 0001, Jianfeng Li 0006, Xiapu Luo, Lei Xue 0001, Junjie Zhang 0004, Zhenhua Li 0001, Xiaohong Guan
USENIX Security Symposium2
2023 HANDOM: Heterogeneous Attention Network Model for Malicious Domain Detection
Qing Wang 0041, Cong Dong, Shijie Jian, Dan Du, Zhigang Lu 0002, Yinhao Qi, Dongxu Han, Xiaobo Ma 0001, Fei Wang 0014
Comput. Secur.8
2023 Who is DNS serving for? A human-software perspective of modeling DNS services
Jian Qu, Xiaobo Ma 0001, Wenmao Liu
Knowl. Based Syst.2
2023 ParaDefender: A Scenario-Driven Parallel System for Defending Metaverses
abstract
The metaverse, as an instance of cyber–physical–social systems (CPSS) that originates in cyber–physical systems (CPS), features growing complexity, and diversity in terms of functionalities, as well as the exponentially increasing demand in network bandwidth and computational resources, thereby leading to exaggerated security threats. However, compared with the extensive attention received by the metaverse, solutions defending against the threats have not kept pace. A major obstacle to such solutions is virtuality–reality-synthesized threats. Therefore, it is imperative to design new paradigms to defend the metaverse effectively. In this article, we advance a parallel system, dubbed ParaDefender, to defend the metaverse against emerging new threats effectively. Inspired by parallel intelligence, ParaDefender comprises artificial cyberspace, computational experiments, and parallel execution. The basic idea is to make artificial and real cyberspaces executed in parallel to mutually guide each other for enhanced security, wherein the parallel execution is scenario driven in the sense that the scenarios originate from all possible spatial–temporal combinations of security threats in the metaverse. We also demonstrate how to land ParaDefender onto real-world applications, including the Industrial Internet of Things (IIoT) security operation application in the industrial metaverse, and the social governance application.
Jinpeng Han, Manzhi Yang, Yuntao Wang 0004, Zhou Su 0001, Xiaobo Ma 0001
IEEE Trans. Syst. Man Cybern. Syst.9
2022 Landing Reinforcement Learning onto Smart Scanning of The Internet of Things
abstract
Cyber search engines, such as Shodan and Censys, have gained popularity due to their strong capability of indexing the Internet of Things (IoT). They actively scan and fingerprint IoT devices for unearthing IP-device mapping. Because of the large address space of the Internet and the mapping’s mutative nature, efficiently tracking the evolution of IP-device mapping with a limited budget of scans is essential for building timely cyber search engines. An intuitive solution is to use reinforcement learning to schedule more scans to networks with high churn rates of IP-device mapping. However, such an intuitive solution has never been systematically studied. In this paper, we take the first step toward demystifying this problem based on our experiences in maintaining a global IoT scanning platform. Inspired by the measurement study of large-scale real-world IoT scan records, we land reinforcement learning onto a system capable of smartly scanning IoT devices in a principled way. We disclose key parameters affecting the effectiveness of different scanning strategies, and find that our system would achieve growing advantages with the proliferation of IoT devices.
Jian Qu, Xiaobo Ma 0001, Wenmao Liu, Hongqing Sang, Jianfeng Li 0006, Lei Xue 0001, Xiapu Luo, Zhenhua Li 0001, Xiaohong Guan
INFOCOM2
2022 Packet-Level Open-World App Fingerprinting on Wireless Traffic
Jianfeng Li 0006, Shuohan Wu, Hao Zhou 0043, Xiapu Luo, Ting Wang 0006, Xiaobo Ma 0001
NDSS7
2022 FOAP: Fine-Grained Open-World Android App Fingerprinting
Jianfeng Li 0006, Hao Zhou 0043, Shuohan Wu, Xiapu Luo, Ting Wang 0006, Xian Zhan, Xiaobo Ma 0001
USENIX Security Symposium7
2022 Overlay-Based Android Malware Detection at Market Scales: Systematically Adapting to the New Technological Landscape
abstract
Androidoverlayenables one app to draw over other apps by creating an extraViewlayer atop the hostView, which nevertheless can be exploited by malicious apps (malware) to attack users. To combat this threat, prior countermeasures concentrate on restricting the capabilities of overlays at the OS level while sacrificing overlays’ usability; recently, the overlay mechanism has been substantially updated to prevent a variety of attacks, which however can still be evaded by considerable adversaries. To address these shortcomings, a more pragmatic approach is to enableearly detectionof overlay-based malware during the app market review process, so that all the capabilities of overlays can stay unchanged. For this purpose, in this paper we first conduct a large-scale comparative study of overlay characteristics in benign and malicious apps, and then implement the OverlayChecker system to automatically detect overlay-based malware for one of the world’s largest Android app stores. In particular, we have made systematic efforts in feature engineering, UI exploration, emulation architecture, and run-time environment, thus maintaining high detection accuracy (97 percent precision and 97 percent recall) and short per-app scan time ($\sim$1.7 minutes) with only two commodity servers, under an intensive workload of$\sim$10K newly submitted apps per day.
Liangyi Gong, Zhenhua Li 0001, Hongyi Wang 0009, Hao Lin 0005, Xiaobo Ma 0001, Yunhao Liu 0001
IEEE Trans. Mob. Comput.5
2022 Inferring Hidden IoT Devices and User Interactions via Spatial-Temporal Traffic Fingerprinting
abstract
With the popularization of Internet of Things (IoT) devices in smart home and industry fields, a huge number of IoT devices are connected to the Internet. However, what devices are connected to a network may not be known by the Internet Service Provider (ISP), since many IoT devices are placed within small networks (e.g., home networks) and are hidden behind network address translation (NAT). Without pinpointing IoT devices in a network, it is unlikely for the ISP to appropriately configure security policies and effectively manage the network. Additionally, inferring fine-grained user interactions of IoT devices is also an interesting yet unresolved problem. In this paper, we design an efficient and scalable system via spatial-temporal traffic fingerprinting from an ISP’s perspective in consideration of practical issues like learning-testing asymmetry. Our system can accurately identify typical IoT devices in a network, with the additional capability of identifying what devices are hidden behind NAT and the number of each type of device that share the same IP address. Our system can also detect user interactions and meanwhile identify their (concurrent) number through a multi-output regression model. Through extensive evaluation, we demonstrate that the system can generally identify IoT devices with an F1-Score above 0.999, and estimate the number of the same type of IoT device behind NAT with an average error below 5%. By studying 29 user interactions of 7 devices, we show that our system is promising in detecting user interactions.
Xiaobo Ma 0001, Jian Qu, Jianfeng Li 0006, John C. S. Lui, Zhenhua Li 0001, Wenmao Liu, Xiaohong Guan
IEEE/ACM Trans. Netw.1
2022 PackerGrind: An Adaptive Unpacking System for Android Apps
abstract
App developers are increasingly using packing services (or packers) to protect their code against being reverse engineered or modified. However, such packing techniques are also leveraged by the malicious developers to prevent the malware from being analyzed and detected by the static malware analysis and detection systems. Though there are already studies on unpacking packed Android apps, they usually leverage the manual reverse engineered packing behaviors to unpack apps packed by the specific packers and cannot be appified to the evolved and new packers. In this paper, we propose a novel unpacking approach with the capacity of adaptively unpacking the evolved and newly encountered packers. Also, we develop a new system, namedPackerGrind, based on this adaptive approach for unpacking Android packers. The evaluation with real packed apps demonstrates thatPackerGrindcan successfully reveal packers protection mechanisms, effectively handle their evolution and recover Dex files with low overhead.
Lei Xue 0001, Hao Zhou 0043, Xiapu Luo, Le Yu 0002, Dinghao Wu, Yajin Zhou, Xiaobo Ma 0001
IEEE Trans. Software Eng.7
2021 Context-aware Website Fingerprinting over Encrypted Proxies
abstract
Website fingerprinting (WFP) could infer which websites a user is accessing via an encrypted proxy by passively inspecting the traffic between the user and the proxy. The key to WFP is designing a classifier capable of distinguishing traffic characteristics of accessing different websites. However, when deployed in real-life networks, a well-trained classifier may face a significant obstacle of training-testing asymmetry, which fundamentally limits its practicability. Specifically, although pure traffic samples can be collected in a controlled (clean) testbed for training, the classifier may fail to extract such pure traffic samples as its input from raw complicated traffic for testing. In this paper, we are interested in encrypted proxies that relay connections between the user and the proxy individually (e.g., Shadowsocks), and design a context-aware system using built-in spatial-temporal flow correlation to address the obstacle. Extensive experiments demonstrate that our system does not only enable WFP against a popular type of encrypted proxies practical, but also achieves better performance than ideally training/testing pure samples.
Xiaobo Ma 0001, Mawei Shi, Bingyu An, Jianfeng Li 0006, Xiapu Luo, Junjie Zhang 0004, Xiaohong Guan
INFOCOM1
2021 Inaccurate Prediction Is Not Always Bad: Open-World Driver Recognition via Error Analysis
abstract
Driver identification is of fundamental importance in many vehicle-related applications, such as fleet monitoring and anti-theft system. The vast majority of existing methods work under the closed-world assumption, which may be unrealistic in practice. In this paper, we consider a more practical but challenging scenario, i.e., open-world driver recognition, and propose a systematic method dubbed DRIVERPRINT. To recognize the driver of interest, DRIVERPRINT takes advantage of the behavioral predictability of the driver himself, thereby no need to collect data from other drivers for model training. Specifically, DRIVERPRINT predicts the behavior-related traveling speed with a driver-specific predictor, compares the prediction error with a pre-trained error model and finally recognizes drivers via error analysis. Besides open-world setting, our method is also compatible with closed-world driver classification. Real-world experiments demonstrate our method achieves reasonable accuracy. The average F1-score for open-world driver recognition is up to 0.91, while that for closed-world driver classification is up to 0.973.
Jianfeng Li 0006, Kaifa Zhao, Yajuan Tang, Xiapu Luo, Xiaobo Ma 0001
VTC Spring5
2021 Systematically Landing Machine Learning onto Market-Scale Mobile Malware Detection
abstract
Despite being crucial to today's mobile ecosystem, app markets have meanwhile become a natural, convenient malware delivery channel as they actually “lend credibility” to malicious apps. In the past few years, machine learning (ML) techniques have been widely explored for automated, robust malware detection, but till now we have not seen an ML-based malware detection solution applied at market scales. To systematically understand the real-world challenges, we conduct a collaborative study with T-Market, a popular Android app market that offers us large-scale ground-truth data. Our study illustrates that the key to successfully developing such systems is multifold, including feature selection and encoding, feature engineering and exposure, app analysis speed and efficacy, developer and user engagement, as well as ML model evolution. Failure in any of the above aspects could lead to the “wooden barrel effect” of the whole system. This article presents our judicious design choices and first-hand deployment experiences in building a practical ML-powered malware detection system. It has been operational at T-Market, using a single commodity server to check ~12K apps every day, and has achieved an overall precision of 98.9 percent and recall of 98.1 percent with an average per-app scan time of 0.9 minutes.
Liangyi Gong, Hao Lin 0005, Zhenhua Li 0001, Feng Qian 0001, Yang Li 0092, Xiaobo Ma 0001, Yunhao Liu 0001
IEEE Trans. Parallel Distributed Syst.6
2020 Pinpointing Hidden IoT Devices via Spatial-temporal Traffic Fingerprinting
abstract
With the popularization of Internet of Things (IoT) devices in smart home and industry fields, a huge number of IoT devices are connected to the Internet. However, what devices are connected to a network may not be known by the Internet Service Provider (ISP), since many IoT devices are placed within small networks (e.g., home networks) and are hidden behind network address translation (NAT). Without pinpointing IoT devices in a network, it is unlikely for the ISP to appropriately configure security policies and effectively manage the network. In this paper, we design an efficient and scalable system via spatial-temporal traffic fingerprinting. Our system can accurately identify typical IoT devices in a network, with the additional capability of identifying what devices are hidden behind NAT and how many they are. Through extensive evaluation, we demonstrate that the system can generally identify IoT devices with an F-Score above 0.999, and estimate the number of the same type of IoT device behind NAT with an average error below 5%. We also perform small-scale (labor-intensive) experiments to show that our system is promising in detecting user-IoT interactions.
Xiaobo Ma 0001, Jian Qu, Jianfeng Li 0006, John C. S. Lui, Zhenhua Li 0001, Xiaohong Guan
INFOCOM1
2020 Taming energy cost of disk encryption software on data-intensive mobile devices
John C. S. Lui, Xiaobo Ma 0001, Jianfeng Li 0006
Future Gener. Comput. Syst.4
2020 Randomized Security Patrolling for Link Flooding Attack Detection
abstract
With the advancement of large-scale coordinated attacks, the adversary is shifting away from traditional distributed denial of service (DDoS) attacks against servers to sophisticated DDoS attacks against Internet infrastructures. Link flooding attacks (LFAs) are such powerful attacks against Internet links. Employing network measurement techniques, the defender could detect the link under attack. However, given the large number of Internet links, the defender can only monitor a subset of the links simultaneously, whereas any link might be attacked. Therefore, it remains challenging to practically deploy detection methods. This paper addresses this challenge from a game-theoretic perspective, and proposes a randomized approach (like security patrolling) to optimize LFA detection strategies. Specifically, we formulate the LFA detection problem as a Stackelberg security game, and design randomized detection strategies in consideration of the adversary's behavior, where best and quantal response models are leveraged to characterize the adversary's behavior. We employ a series of techniques to solve the nonlinear and nonconvex NP-hard optimization problems for finding the equilibrium. The experimental results demonstrate the necessity of handling LFAs from a game-theoretic perspective and the effectiveness of our solutions. We believe our study is a significant step forward in formally understanding LFA detection strategies.
Xiaobo Ma 0001, Bo An 0001, Mengchen Zhao, Xiapu Luo, Lei Xue 0001, Zhenhua Li 0001, Tony T. N. Miu, Xiaohong Guan
IEEE Trans. Dependable Secur. Comput.1
2019 DeepCG: Classifying Metamorphic Malware Through Deep Learning of Call Graphs
Xiaobo Ma 0001
SecureComm (1)2
2019 A Commit Messages-Based Bug Localization for Android Applications
abstract
Recently, there has been consistent growth in Android applications (apps). Under these circumstances, software maintenance for Android apps becomes an essential and important task. The core of software maintenance is to locate bugs in source files. Previous bug localization approaches mainly focus on open-source desktop software (e.g. Eclipse, Mozilla, GCC). Even though a few studies locate the bugs in the Android apps, they are dedicated to a special app named ZXing, without developing a general method to locate the bugs in Android apps by taking into account the unique characteristics of Android apps’ bug reports. Such characteristics include fewer number of historical bug reports, insufficient detailed description, etc. These characteristics hinder existing localization approaches from being directly delivered to Android apps, because lack of enough information degrades the performance of those localization approaches relying on historical bug reports. Commit messages include more informative data which can provide the details of reported bugs. Therefore, in this paper, we propose a novel information retrieval-based approach which utilizes commit messages to locate new bugs in Android apps. This approach not only considers the structured textual similarity between the given bug and the candidate source files, but also computes the unstructured textual similarities between the new bug and the commit messages linked to the corresponding source files. According to the experimental results on 10 popular open-source Android apps managed by GitHub, our approach outperforms the state-of-the-art bug localization methods that include BugLocator, BLUiR, and two-phase model.
Tao Zhang 0001, Xiapu Luo, Xiaobo Ma 0001
Int. J. Softw. Eng. Knowl. Eng.4
2019 Protecting internet infrastructure against link flooding attacks: A techno-economic perspective
Xiaobo Ma 0001, Jianfeng Li 0006, Yajuan Tang, Bo An 0001, Xiaohong Guan
Inf. Sci.1
2018 Can We Learn what People are Doing from Raw DNS Queries?
abstract
Domain Name System (DNS) is one of the pillars of today's Internet. Due to its appealing properties such as low data volume, wide-ranging applications and encryption free, DNS traffic has been extensively utilized for network monitoring. Most existing studies of DNS traffic, however, focus on domain name reputation. Little attention has been paid to understanding and profiling what people are doing from DNS traffic, a fundamental problem in the areas including Internet demographics and network behavior analysis. Consequently, simple questions like “How to determine whether a DNS query for www.google.com means searching or any other behaviors?” cannot be answered by existing studies. In this paper, we take the first step to identify user activities from raw DNS queries. We advance a multiscale hierarchical framework to tackle two practical challenges, i.e., behavior ambiguity and behavior polymorphism. Under this framework, a series of novel methods, such as pattern upward mapping and multi-scale random forest classifier, are proposed to characterize and identify user activities of interest. Evaluation using both synthetic and real-world DNS traces demonstrates the effectiveness of our method.
Jianfeng Li 0006, Xiaobo Ma 0001, Xiapu Luo, Junjie Zhang 0004, Wei Li 0029, Xiaohong Guan
INFOCOM2
2018 Shoot at a Pigeon and Kill a Crow: On Strike Precision of Link Flooding Attacks
Xiaobo Ma 0001, Jianfeng Li 0006, Lei Xue 0001
NSS2
2018 Revisiting Website Fingerprinting Attacks in Real-World Scenarios: A Case Study of Shadowsocks
Yankang Zhao, Xiaobo Ma 0001, Jianfeng Li 0006, Shui Yu 0001, Wei Li 0172
NSS2
2018 Exploiting Proximity-Based Mobile Apps for Large-Scale Location Privacy Probing
abstract
Proximity-based apps have been changing the way people interact with each other in the physical world. To help people extend their social networks, proximity-based nearby-stranger (NS) apps that encourage people to make friends with nearby strangers have gained popularity recently. As another typical type of proximity-based apps, some ridesharing (RS) apps allowing drivers to search nearby passengers and get their ridesharing requests also become popular due to their contribution to economy and emission reduction. In this paper, we concentrate on the location privacy of proximity-based mobile apps. By analyzing the communication mechanism, we find that many apps of this type are vulnerable to large-scale location spoofing attack (LLSA). We accordingly propose three approaches to performing LLSA. To evaluate the threat of LLSA posed to proximity-based mobile apps, we perform real-world case studies against an NS app named Weibo and an RS app called Didi. The results show that our approaches can effectively and automatically collect a huge volume of users’ locations or travel records, thereby demonstrating the severity of LLSA. We apply the LLSA approaches against nine popular proximity-based apps with millions of installations to evaluate the defense strength. We finally suggest possible countermeasures for the proposed attacks.
Xiapu Luo, Xiaobo Ma 0001, Yankang Zhao, Zeming Yang, Man Ho Au, Xinliang Qiu
Secur. Commun. Networks3
2018 LinkScope: Toward Detecting Target Link Flooding Attacks
abstract
A new class of target link flooding attacks (LFAs) can cut off the Internet connections of a target area without being detected, because they employ legitimate flows to congest selected links. Although new mechanisms for defending against LFA have been proposed, the deployment issues limit their usage, since they require either additional modules to enhance routers or using the software-defined network to replace the traditional routers. In this paper, we propose a novel framework that employs both the end-to-end and hop-by-hop network measurement techniques to capture the abnormal path performance degradation for detecting LFA and then locate the target links or areas whenever possible, and develop a prototype of the framework named LinkScope. Although using network measurement to capture network anomaly is not new, we tackle a number of challenging issues, such as conducting large-scale Internet path monitoring via non-cooperative measurement so that users do not need to install LinkScope on every host, profiling the performance of asymmetric Internet paths and detecting LFA. The extensive evaluation in a testbed and the Internet shows that with limited bandwidth and computational overhead, LinkScope can achieve timely detection and diagnosis of LFA with high detection rate and low false positive rate.
Lei Xue 0001, Xiaobo Ma 0001, Xiapu Luo, Edmond W. W. Chan, TungNgai Miu, Guofei Gu
IEEE Trans. Inf. Forensics Secur.2
2017 Is what you measure what you expect? Factors affecting smartphone-based mobile network measurement
abstract
Many apps have been developed to measure the performance of mobile networks. Unfortunately, their measurement results may not be what users expect, because the results could be biased by various factors and the apps' descriptions may confuse users. Although a few recent studies pointed out several factors, they missed other important factors and lacked of finegrained analysis on the factors and measurement apps. Moreover, none has studied whether or not the descriptions of such apps will mislead users. In this paper, we conduct the first systematic study of the factors that could bias the result from measurement apps and their descriptions. We identify new factors, revisit known factors, and propose a novel approach with new tools to discover these factors in proprietary apps. We also develop a new measurement app named MobiScope for demonstrating how to mitigate the negative effects of these factors. Furthermore, we construct enhanced descriptions for measurement apps to provide users more information about what is measured. The extensive experimental results illustrate the negative effects of various factors, the improvement in performance measurement brought by MobiScope, and the clarity of the enhanced descriptions.
Lei Xue 0001, Xiaobo Ma 0001, Xiapu Luo, Le Yu 0002, Shuai Wang 0012, Ting Chen 0002
INFOCOM2
2017 AutoFlowLeaker: Circumventing Web Censorship through Automation Services
abstract
By hiding messages inside existing network protocols, anti-censorship tools could empower censored users to visit blocked websites. However, existing solutions generally suffer from two limitations. First, they usually need the support of ISP or the deployment of many customized hosts to conceal the communication between censored users and blocked websites. Second, their manipulations of normal network traffic may result in detectable features, which could be captured by the censorship system. In this paper, to tackle these limitations, we propose a novel framework that exploits the publicly available automation services and the plenty of web services and contents to circumvent web censorship, and realize it in a practical tool named AutoFlowLeaker. Moreover, we conduct extensive experiments to evaluate AutoFlowLeaker, and the results show that it has promising performance and can effectively evade realworld web censorship.
Shengtuo Hu, Xiaobo Ma 0001, Muhui Jiang, Xiapu Luo, Man Ho Au
SRDS2
2017 Mining repeating pattern in packet arrivals: Metrics, models, and applications
Jianfeng Li 0006, Xiaobo Ma 0001, Junjie Zhang 0004, Pinghui Wang, Xiaohong Guan
Inf. Sci.2
2016 I Know Where You All Are! Exploiting Mobile Social Apps for Large-Scale Location Privacy Probing
Xiapu Luo, Xiaobo Ma 0001, Xinliang Qiu, Man Ho Au
ACISP (1)4
2015 Modeling repeating behaviors in packet arrivals: Detection and measurement
abstract
With the growing stickiness of the Internet, numerous automated programs running in terminal facilities (e.g., laptops) tend to keep closely connected to the Internet by repetitively interacting with remote services. It is of fundamental importance to study such repeating behaviors of automated programs in areas like traffic engineering and network monitoring. This paper focuses on repeating behaviors in packet arrivals that are of interest, aiming at a hierarchical characterization of packet arrivals, detection methods and quantitative metrics. To this end, we present a structure-oriented characterization of packet arrivals, which reflects the temporal structure of repeating behaviors at different scales. Based on such characterization, a repeating behavior detection method is proposed by leveraging online-learning prediction, and two novel metrics of repeating behaviors are proposed from different aspects. In addition, a denoising method is developed to enhance the noise-tolerant capability of detection and measurement in face of noises. Experimental results based on real-world traces demonstrate the effectiveness of our proposed approaches in automated program behavior detection and behavioral botnet analysis.
Jianfeng Li 0006, Xiaobo Ma 0001, Junjie Zhang 0004, Xiaohong Guan
INFOCOM3
2015 Accurate DNS query characteristics estimation via active probing
Xiaobo Ma 0001, Junjie Zhang 0004, Zhenhua Li 0001, Jianfeng Li 0006, Xiaohong Guan, John C. S. Lui, Don Towsley
J. Netw. Comput. Appl.1
2014 MIGDroid: Detecting APP-Repackaging Android malware via method invocation graph
abstract
With the increasing popularity of Android platform, Android malware, especially APP-Repackaging malware wherein the malicious code is injected into legitimate Android applications, is spreading rapidly. This paper proposes a new system named MIGDroid, which leverages method invocation graph based static analysis to detect APP-Repackaging Android malware. The method invocation graph reflects the “interaction” connections between different methods. Such graph can be naturally exploited to detect APP-Repackaging malware because the connections between injected malicious code and legitimate applications are expected to be weak. Specifically, MIGDroid first constructs method invocation graph on the smali code level, and then divides the method invocation graph into weakly connected sub-graphs. To determine which sub-graph corresponds to the injected malicious code, the threat score is calculated for each sub-graph based on the invoked sensitive APIs, and the subgraphs with higher scores will be more likely to be malicious. Experiment results based on 1,260 Android malware samples in the real world demonstrate the specialty of our system in detecting APP-Repackaging Android malware, thereby well complementing existing static analysis systems (e.g., Androguard) that do not focus on APP-Repackaging Android malware.
Xiaobo Ma 0001, Wenyu Zhou
ICCCN3
2014 DNSRadar: Outsourcing Malicious Domain Detection Based on Distributed Cache-Footprints
abstract
As the domain name system (DNS) plays a critical role in malicious services and number of networks, especially small enterprise networks and home networks that are generally and poorly managed, grows rapidly, it is highly desired to outsource the malicious domain detection service to a thirdparty system that can aggregate information from multiple vantage points to perform detection. To this end, we propose DNSRadar, a system that explores the coexistence of domain cache-footprints distributed in all networks that participate in the outsourcing service. Bootstrapping from a list of prelabeled malicious domains, DNSRadar leverages link analysis techniques to infer maliciousness likelihood of unknown domains based on coexistence information. As DNSRadar only uses the existence of an unknown domain in a network for detection, privacy concerns have been drastically reduced. Both MapReduce and lightweight matrix analysis techniques are employed to implement DNSRadar, making scalability as a built-in feature. Taking advantage of a large number of open recursive DNS servers, we have performed extensive evaluation at scale. Experimental results have demonstrated that DNSRadar can efficiently detect ~90% malicious domains given a low false positive rate of 1%. Of all these detected malicious domains, ~30% are on average 6 days earlier than public DNS reputation services, indicating DNSRadar's great early detection capability.
Xiaobo Ma 0001, Junjie Zhang 0004, Jianfeng Li 0006, Jue Tian, Xiaohong Guan
IEEE Trans. Inf. Forensics Secur.1
2013 Boosting practicality of DNS cache probing: A general estimator based on Bayesian forecasting
abstract
It is an important task in Internet demography and security monitoring to accurately measure the user population of an application in a network. In previous works, the Domain Name System (DNS) cache probing technique was proposed to estimate λ¯, the average DNS querying rate for a domain name associated with the given application. One can readily obtain the user population given another empirical parameter from DNS traces, i.e., average number of DNS queries per user. The previous estimator for λ¯ was based on the assumption that the DNS query arrivals can be described by a homogeneous Poisson process. In this paper, we verify this assumption by measuring real DNS traces and find that it is over-simplified. In fact, the DNS query arrivals exhibit non-stationary property dominated by a diurnal pattern in general, thereby making the previous estimator underestimate λ¯. Then, an asymptotically unbiased estimator is proposed using the Bayesian forecasting. The proposed estimator is more general as compared with the previous one because it can accurately estimate λ¯ when the DNS query arrivals can be described by either homogeneous or non-homogeneous Poisson processes. The proposed estimator meets the minimum mean squared error principle, and the experimental results show that it significantly outperforms the previous one. The DNS cache probing technique offers promising applications because it is low-cost, less invasive and privacy preserving. Our work greatly boosts the practicability of this technique.
Jianfeng Li 0006, Xiaobo Ma 0001, Xiaohong Guan
ICC2
2012 Cloud-based push-styled mobile botnets: a case study of exploiting the cloud to device messaging service
abstract
Given the popularity of smartphones and mobile devices, mobile botnets are becoming an emerging threat to users and network operators. We propose a new form of cloud-based push-styled mobile botnets that exploits today's push notification services as a means of command dissemination. To motivate its practicality, we present a new command and control (C&C) channel using Google's Cloud to Device Messaging (C2DM) service, and develop a C2DM botnet specifically for the Android platform. We present strategies to enhance its scalability to large botnet coverage and its resilience against service disruption. We prototype a C2DM botnet, and perform evaluation to show that the C2DM botnet is stealthy in generating heartbeat and command traffic, resource-efficient in bandwidth and power consumptions, and controllable in quickly delivering a command to all bots. We also discuss how one may deploy a C2DM botnet, and demonstrate its feasibility in launching an SMS-Spam-and-Click attack. Lastly, we discuss how to generalize the design to other platforms, such as iOS or Window-based systems, and recommend possible defense methods. Given the wide adoption of push notification services, we believe that this type of mobile botnets requires special attention from our community.
Patrick P. C. Lee, John C. S. Lui, Xiaohong Guan, Xiaobo Ma 0001
ACSAC5
2012 Towards active measurement for DNS query behavior of botnets
abstract
Domain names play an increasingly important role for the botnet activities. Traditionally, DNS traces from several local DNS servers are used passively to measure the DNS query behavior. However, since botnets are a wide-scale threat and usually reside in geographically dispersed networks, the vantage point of several local DNS servers is sometimes too small to help us understand the DNS query behavior (e.g., whether queried or not, average query rate) of botnets. In this paper, we actively measure the DNS query behavior of botnets in geographically dispersed networks via the DNS cache probing technique. We first analytically characterize how multiple domain names are queried by botnets in different networks under certain circumstances. Then, we actively measure real botnet samples in the wild to gain insight into how multiple domain names are queried by botnets in 480 geographically dispersed networks globally, and show that our analytical characterization well describes the DNS query behavior of the botnet samples. The active measurement technique can help to acquire extensive DNS query information in different networks and thus potentially facilitate various DNS-related research and applications.
Xiaobo Ma 0001, Jianfeng Li 0006, Xiaohong Guan
GLOBECOM1
2010 A Novel IRC Botnet Detection Method Based on Packet Size Sequence
abstract
Botnets have become a serious threat to Internet and are often deployed to control a large pool of zombies and perform notorious activities such as DDoS, information theft and spam sending. In this paper, a new method is developed for detecting IRC botnets by analyzing the characteristic of packet size sequence of the TCP conversation between IRC zombies and their command and control (C&C) servers. In comparison with IRC chat, the TCP conversations within IRC botnets show a nature of approximate periodicity defined as quasi-periodicity in this paper. A simple yet effective detection method is presented to detect IRC botnets by measuring the quasi-periodicity degree and packet average size of IRC conversations based on ukkonen algorithm. We evaluated our method using real-world IRC botnet traces captured from honeynet. The results show that our method can detect real-world IRC botnets from IRC traffic with high accuracy and has a low false positive rate.
Xiaobo Ma 0001, Xiaohong Guan, Yun Guo
ICC1