Qiang Li 0008

dblp:72/872-8 · DBLP profile ↗
← Back
43ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0001-7510-4718ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 11 · 3 since 2021Systems, architecture and hardware · 8 · 3 since 2021Computer networks · 8 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Theory of computation · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Selective Constraint Learning for Unsupervised Cross-Domain Image Retrieval
abstract
Unsupervised cross-domain image retrieval aims to retrieve semantically consistent images across domains with significant domain gaps, which poses substantial challenges under the absence of annotations in both domains. Existing approaches primarily rely on internally derived supervision signals for representation learning and cross-domain alignment. However, such internally induced supervision tends to impose an upper bound on achievable retrieval performance, as it lacks stable semantic references to support reliable category-level correspondence across domains. To address these limitations, we propose Selective Constraint Learning (SCL), a framework that introduces external semantic guidance as a stable prior for unsupervised cross-domain image retrieval. Leveraging a pre-trained foundation model, SCL constructs a dual-scope constraint bank to capture high-confidence positive and negative semantic relations within and across domains. Based on this, we design a generic constraint loss to jointly facilitate intra-domain compactness and inter-domain alignment. In addition, prototypical geometry regularization is designed to enhance in-domain structural stability through prototype-centered pull-and-push forces. Extensive experiments on multiple benchmarks demonstrate that SCL consistently outperforms state-of-the-art methods.
Wensi Fang, Xiaodan Zhang 0006, Xiaoyu Lian, Qiang Li 0008, Shuai Lü 0001
SIGIR4
2026 FL-LLM: Leveraging large language models for efficient device selection in federated learning
Qiang Li 0008, He Li 0001
Comput. Networks2
2026 FedPCAN: Robust mitigation of dynamic black-box backdoor attacks in federated learning
abstract
The decentralized structure of federated learning exposes it to backdoor threats, especially dynamic black-box attacks. These attacks employ multiple evasion mechanisms within malicious updates, rendering many existing defenses ineffective. To address this threat, this paper proposes a three-stage defense algorithm. First, it performs neuron parameter analysis on PCA-reduced gradients to identify and remove decoy models injected by malicious clients. Second, a novel backdoor detection method is designed, where the central server generates random inputs to compare neuron activation differences across clients, thereby distinguishing between benign and backdoor gradients. Finally, Gaussian noise is added to disrupt the accuracy of the attacker’s feedback indicators. Compared to other defense algorithms, the proposed algorithm defends against each evasion module of dynamic black-box attacks while rendering the attacker’s feedback indicators ineffective. We evaluate the performance under different attack scenarios on four datasets, and the results show that the proposed algorithm effectively defends against various backdoor attacks and outperforms state-of-the-art methods.
Qiang Li 0008, He Li 0001
Knowl. Based Syst.2
2025 Efficient Sharing of Energy Consumption Data: A Privacy-Preserving Threshold Aggregation Approach
abstract
Energy consumption data collected by smart meters (SMs) is increasingly used by various subscribers in the smart grid for load management, energy monitoring, and policy planning. To protect user privacy, edge-assisted privacy-preserving data aggregation (PPDA) techniques are commonly employed. However, existing methods face several challenges: 1) limited scalability, 2) strict trust requirements, and 3) the risk of revealing unique consumption patterns to data collectors. To address these challenges, we propose a privacy-preserving threshold aggregation method that is easily scalable and facilitates efficient energy data sharing under limited trust assumptions. Specifically, we design VFP-NTRU, a quantum-resistant homomorphic proxy re-encryption scheme with fault tolerance and re-encryption verification. In VFP-NTRU, SMs can encrypt data with a public key without the need for prior negotiation of decryption keys with multiple subscribers. Additionally, we develop a privacy threshold collection protocol that uses a verifiable oblivious pseudorandom function to provide privacy guarantees similar to k-anonymity for SM data collection. We further introduce an energy consumption model to determine optimal collection strategies, improving system responsiveness. We provide correctness analysis and prove the security of our scheme. Experimental results demonstrate that our approach outperforms existing PPDA methods, making it particularly suitable for resource-constrained SMs and central servers managing large-scale energy data.
Guohao Li 0004, Jiale Lian, Siyi Liu 0005, Li Yang 0005, Yantao Zhong, Qiang Li 0008
IEEE Internet Things J.7
2025 A Secret Sharing-Inspired Robust Distributed Backdoor Attack to Federated Learning
abstract
Federated Learning (FL) is vulnerable to backdoor attacks—especially distributed backdoor attacks (DBA) that are more persistent and stealthy than centralized backdoor attacks. However, we observe that the attack effectiveness of DBA can be largely reduced when encountering rebels, i.e., the agents promising to perform the attack but do not do so. To robustify DBAs, we present SSRDBA , a secret sharing-inspired robust DBA to FL. To be specific, given a same global trigger as DBA, SSRDBA carefully divides it into different shares based on secret sharing and exploits these shares to poison local data on malicious devices, respectively. SSRDBA enjoys several merits, e.g., only partial malicious agents guarantee the reconstruction of the global trigger. Extensive experimental results show that SSRDBA is more robust to rebels than DBA and can evade the state-of-the-art FL defenses mainly for centralized backdoor attacks. To mitigate SSRDBA , we further design a novel defense mechanism, termed NFDR, which shows great potential against SSRDBA on certain independent identically distributed datasets.
Yuxin Yang 0003, Qiang Li 0008, Yuede Ji, Binghui Wang
ACM Trans. Priv. Secur.2
2024 Distributed Backdoor Attacks on Federated Graph Learning and Certified Defenses
abstract
Federated graph learning (FedGL) is an emerging federated learning (FL) framework that extends FL to learn graph data from diverse sources without accessing the data. FL for non-graph data has shown to be vulnerable to backdoor attacks, which inject a shared backdoor trigger into the training data such that the trained backdoored FL model can predict the testing data containing the trigger as the attacker desires. However, FedGL against backdoor attacks is largely unexplored, and no effective defense exists.
Yuxin Yang 0003, Qiang Li 0008, Jinyuan Jia 0001, Yuan Hong 0001, Binghui Wang
CCS2
2024 Breaking State-of-the-Art Poisoning Defenses to Federated Learning: An Optimization-Based Attack Framework
abstract
Federated Learning (FL) is a novel client-server distributed learning framework that can protect data privacy. However, recent works show that FL is vulnerable to poisoning attacks. Many defenses with robust aggregators (AGRs) are proposed to mitigate the issue, but they are all broken by advanced attacks. Very recently, some renewed robust AGRs are designed, typically with novel clipping or/and filtering strategies, and they show promising defense performance against the advanced poisoning attacks. In this paper, we show that these novel robust AGRs are also vulnerable to carefully designed poisoning attacks. Specifically, we observe that breaking these robust AGRs reduces to bypassing the clipping or/and filtering of malicious clients, and propose an optimization-based attack framework to leverage this observation. Under the framework, we then design the customized attack against each robust AGR. Extensive experiments on multiple datasets and threat models verify our proposed optimizationbased attack can break the SOTA AGRs. We hence call for novel defenses against poisoning attacks to FL. Code is available at: https: //github.com/Yuxin104/BreakSTOAPoisoningDefenses.
Yuxin Yang 0003, Qiang Li 0008, Chenfei Nie, Yuan Hong 0001, Binghui Wang
CIKM2
2024 FedGMark: Certifiably Robust Watermarking for Federated Graph Learning
abstract
Federated graph learning (FedGL) is an emerging learning paradigm to collaboratively train graph data from various clients. However, during the development and deployment of FedGL models, they are susceptible to illegal copying and model theft. Backdoor-based watermarking is a well-known method for mitigating these attacks, as it offers ownership verification to the model owner. We take the first step to protect the ownership of FedGL models via backdoor-based watermarking. Existing techniques have challenges in achieving the goal: 1) they either cannot be directly applied or yield unsatisfactory performance; 2) they are vulnerable to watermark removal attacks; and 3) they lack of formal guarantees. To address all the challenges, we propose FedGMark, the first certified robust backdoor-based watermarking for FedGL. FedGMark leverages the unique graph structure and client information in FedGL to learn customized and diverse watermarks. It also designs a novel GL architecture that facilitates defending against both the empirical and theoretically worst-case watermark removal attacks. Extensive experiments validate the promising empirical and provable watermarking performance of FedGMark. Source code is available at: https://github.com/Yuxin104/FedGMark.
Yuxin Yang 0003, Qiang Li 0008, Yuan Hong 0001, Binghui Wang
NeurIPS2
2024 DeCoCDR: Deployable Cloud-Device Collaboration for Cross-Domain Recommendation
abstract
Cross-domain recommendation (CDR) is a widely used methodology in recommender systems to combat data sparsity.It leverages user data across different domains or platforms for providing personalized recommendations.Traditional CDR assumes user preferences and behavior data can be shared freely among cloud and users, which is now impractical due to strict restrictions of data privacy.In this paper, we propose a Deployment-friendly Cloud-Device Collaboration framework for Cross-Domain Recommendation (De-CoCDR).It splits CDR into a two-stage recommendation model through cloud-device collaborations, i.e., item-recall on cloud and item re-ranking on device.This design enables effective CDR while preserving data privacy for both the cloud and the device.Extensive offline and online experiments are conducted to validate the effectiveness of DeCoCDR.In offline experiments, DeCoCDR outperformed the state-of-the-arts in three large datasets.While in real-world deployment, DeCoCDR improved the conversion rate by 45.3% compared with the baseline.
Yi Zhang 0178, Zimu Zhou, Qiang Li 0008
SIGIR4
2024 EPIDL: Towards efficient and privacy-preserving inference in deep learning
abstract
Summary Deep learning has shown its great potential in real‐world applications. However, users(clients) who want to use deep learning applications need to send their data to the deep learning service provider (server), which can make the client's data leak to the server, resulting in serious privacy concerns. To address this issue, we propose a protocol named EPIDL to perform efficient and secure inference tasks on neural networks. This protocol enables the client and server to complete inference tasks by performing secure multi‐party computation (MPC) and the client's private data is kept secret from the server. The work in EPIDL can be summarized as follows: First, we optimized the convolution operation and matrix multiplication, such that the total communication can be reduced; Second, we proposed a new method for truncation following secure multiplication based on oblivious transfer and garbled circuits, which will not fail and can be executed together with the ReLU activation function; Finally, we replace complex activation function with MPC‐friendly approximation function. We implement our work in C++ and accelerate the local matrix computation with CUDA support. We evaluate the efficiency of EPIDL in privacy‐preserving deep learning inference tasks, such as the time to execute a secure inference on the MNIST dataset in the LeNet model is about 0.14 s. Compared with the state‐ofthe‐art work, our work is 1.8–98 faster over LAN and WAN, respectively. The experimental results show that our EPIDL is efficient and privacy‐preserving.
Chenfei Nie, Mianxiong Dong, Kaoru Ota, Qiang Li 0008
Concurr. Comput. Pract. Exp.5
2023 Injecting Revenue-awareness into Cold-start Recommendation: The Case of Online Insurance
abstract
In online insurance, one of the central challenges is the cold-starting of new insurance products, which means there are no previous samples to refer to. Previous studies have mainly focused on improving the prediction accuracy of new items, but they have failed to consider the revenue generated by existing items. As a result, the total revenue may suffer a loss even if new items get conversions. To address this issue, in this paper we propose RACRec, a Revenue-Aware Cold-start Recommendation framework for online insurance. Unlike previous works, RACRec uses a revenue-based objective function to ensure profitability when cold-starting new items. With this dedicated objective, RACRec orchestrates the cold-start and warm-start model by predicting their expected revenue from each user, thereby preserving the total profit. In order to accurately predict revenue, a double-ranking scheme is designed for the warm-start model to mitigate position bias, while an item embedding alignment algorithm is proposed for the cold-start model to learn the revenue of new items from similar old items. Furthermore, a reinforced orchestrate update scheme is designed to eliminate the impact of sparse conversions and continuously update revenue estimation. Extensive offline and online experiments have been conducted to validate the effectiveness of RACRec.
Yi Zhang 0178, Helen He Chang, Qiang Li 0008
ICPADS4
2022 POSGen: Personalized Opening Sentence Generation for Online Insurance Sales
abstract
The insurance industry is shifting their sales mode from offline to online, in expectation to reach massive potential customers in the digitization era. Due to the complexity and the nature of insurance products, a cost-effective online sales solution is to exploit chatbot AI to raise customers’ attention and pass those with interests to human agents for further sales. For high response and conversion rates of customers, it is crucial for the chatbot to initiate a conversation with personalized opening sentences, which are generated with user-specific topic selection and ordering. Such personalized opening sentence generation is challenging because (i) there are limited historical samples for conversation topic recommendation in online insurance sales and (ii) existing text generation schemes often fail to support customized topic ordering based on user preferences. We design POSGen, a personalized opening sentence generation scheme dedicated for online insurance sales. It transfers user embeddings learned from auxiliary online user behaviours to enhance conversation topic recommendation, and exploits a context management unit to arrange the recommended topics in user-specific ordering for opening sentence generation. POSGen is deployed on a real-world online insurance platform. It achieves 2.33x total insurance premium improvement through a two-month global test.
Yi Zhang 0178, Zimu Zhou, Qiang Li 0008
IEEE Big Data5
2022 LEGO: A hybrid toolkit for efficient 2PC-based privacy-preserving machine learning
Qianjun Wei, Qiang Li 0008
Comput. Secur.4
2022 BTS: A Blockchain-Based Trust System to Deter Malicious Data Reporting in Intelligent Internet of Things
abstract
Recent developments in collection, computation and communication have expanded the way of data reporting in intelligent Internet of Things (IoT). However, diversity and complexity of data sources also impose new trust challenge in data collection process since untrust reporters tend to report false or even malicious data, which highlights the need to develop a novel methodology to solve such challenge. Thus, based on this domain, inspired by deterrence theory, this article proposes a blockchain-based trust system with assistant of drones to deter malicious data reporting in intelligent IoT. Specifically, to deter malicious data reporting, based on the blockchain technology, the data sensed by fully trusted drones is public published on blockchain showing participants the data standards, named as malicious deterrence scheme. This scheme provides a barrier for malicious reporters to arbitrarily publish false data to blockchain, since the false data can be easily detected while they cannot deny. Second, to further reduce malicious data reporting, a strict penalty mechanism is proposed to punish malicious reporters who have reported false data to blockchain to reduce the malicious data reporting in the following task through punishment. Third, note that the sensing of data standard generates additional costs, therefore, a drone flight route scheme based on a simper deep reinforcement learning with multihead attention mechanism (MA-DRL) is designed to reduce the flight distance for drones. Finally, extensive experiments demonstrate efficiency of our proposed system in terms of reducing malicious data reporting in advance as well as reducing drone flight distance.
Ting Li 0009, Wei Liu 0077, Anfeng Liu, Mianxiong Dong, Kaoru Ota, Naixue Xiong, Qiang Li 0008
IEEE Internet Things J.7
2022 BPT: A Blockchain-Based Privacy Information Preserving System for Trust Data Collection Over Distributed Mobile-Edge Network
abstract
Contemporarily, the fast development of computing, communication, and storage technology has revolutionized the way that various data-based applications reach massive data from underlying sensor networks. However, such a process also raises two challenging but critical issues: 1) trustworthy and 2) privacy issue for data collectors. Therefore, this article proposes a novel system, which is designed over the distributed mobile-edge network to sufficiently exploit advantages of blockchain and differential privacy (DP) to collect trustworthy data and protect privacy for data collectors. First, to improve trustworthiness of data collections, a new consensus mechanism is proposed for blockchain-based data collection structure, which comprehensively incorporates trustworthy, collection contribution, and throughput together to prefer data collectors for the next block. Second, with the assistance of fully trusted devices, a verifiable trustworthy evaluation strategy is designed to accurately compute the trustworthiness for data collectors. Third, we enforce DP on the data stored in a global blockchain maintained by the cloud server to protect privacy for data collectors without influencing data availability. Finally, both theoretical analyses and experimental results prove that the proposed system comprehensively improves performance of data collections in distributed network without adding any additional cost for the cloud server, compared to other schemes.
Ting Li 0009, Wei Liu 0077, Shangsheng Xie, Mianxiong Dong, Kaoru Ota, Naixue Xiong, Qiang Li 0008
IEEE Internet Things J.7
2021 RevMan: Revenue-aware Multi-task Online Insurance Recommendation
abstract
Online insurance is a new type of e-commerce with exponential growth. An effective recommendation model that maximizes the total revenue of insurance products listed in multiple customized sales scenarios is crucial for the success of online insurance business. Prior recommendation models are ineffective because they fail to characterize the complex relatedness of insurance products in multiple sales scenarios and maximize the overall conversion rate rather than the total revenue. Even worse, it is impractical to collect training data online for total revenue maximization due to the business logic of online insurance. We propose RevMan, a Revenue-aware Multi-task Network for online insurance recommendation. RevMan adopts an adaptive attention mechanism to allow effective feature sharing among complex insurance products and sales scenarios. It also designs an efficient offline learning mechanism to learn the rank that maximizes the expected total revenue, by reusing training data and model for conversion rate maximization. Extensive offline and online evaluations show that RevMan outperforms the state-of-the-art recommendation systems for e-commerce.
Yi Zhang 0178, Gengwei Hong, Zimu Zhou, Qiang Li 0008
AAAI6
2021 Discovering unknown advanced persistent threat using shared features mined by neural networks
Longkang Shang, Dong Guo 0002, Yuede Ji, Qiang Li 0008
Comput. Networks4
2021 STMTO: A smart and trust multi-UAV task offloading system
Jialin Guo, Guosheng Huang, Qiang Li 0008, Naixue Xiong, Shaobo Zhang 0001, Tian Wang 0001
Inf. Sci.3
2021 Identifying compromised hosts under APT using DNS request sequences
Qiang Li 0008, Guangzhe Xuan, Dong Guo 0002
J. Parallel Distributed Comput.2
2021 Privacy-preserving two-parties logistic regression on vertically partitioned data using asynchronous gradient sharing
Qianjun Wei, Qiang Li 0008, Zhengqiang Ge
Peer-to-Peer Netw. Appl.2
2021 Incentive Mechanism for Mobile Devices in Dynamic Crowd Sensing System
abstract
Mobile crowdsensing (MCS) has gained much attention due to the proliferation of smart devices equipped with powerful sensors. Large-scale users are the foundation of MCS, so designing incentive mechanisms to motivate users to participate in MCS is necessary. Existing works on incentive mechanisms usually assume a scenario where a group of tasks arrive at the platform at the same time and are immediately assigned to users. We argue that a more realistic MCS scenario can delay a task, which is called the assignment duration time, to wait for appropriate users. In this scenario, we focus on proposing a truthful incentive mechanism to reduce the overall social cost. Due to the uncertainty of coming users, the problems of selecting the appropriate users and calculating the payment for each recruited user (winner) are more complicated. To overcome these challenges, we design a dynamic truthful incentive mechanism (DTIM) including winner selection and payment decision processes. The former uniformly recruits users before the assignment deadline of tasks and dynamically readjusts the recruiting frequency of other tasks to select winners iteratively, which achieves an approximation ratio. Furthermore, the latter determines truthful payment for each winner to encourage user participation as well as avoid being deceived, which achieves truthfulness, individual rationality, and computational efficiency. Finally, massive simulations based on a real dataset roma/taxi validate the DTIM, which can effectively reduce the overall social cost and make a truthful payment for each winner.
Hengzhi Wang, Yongjian Yang 0001, En Wang, Liang Wang 0017, Qiang Li 0008, Zhiyong Yu 0001
IEEE Trans. Hum. Mach. Syst.5
2020 Detecting mobile advanced persistent threats based on large-scale DNS logs
Zongyuan Xiang, Dong Guo 0002, Qiang Li 0008
Comput. Secur.3
2019 QuickSquad: A new single-machine graph computing framework for detecting fake accounts in large-scale social networks
Xinyang Jiang, Qiang Li 0008, Mianxiong Dong, Jun Wu 0001, Dong Guo 0002
Peer-to-Peer Netw. Appl.2
2019 Stopping the Cyberattack in the Early Stage: Assessing the Security Risks of Social Network Users
abstract
Online social networks have become an essential part of our daily life. While we are enjoying the benefits from the social networks, we are inevitably exposed to the security threats, especially the serious Advanced Persistent Threat (APT) attack. The attackers can launch targeted cyberattacks on a user by analyzing its personal information and social behaviors. Due to the wide variety of social engineering techniques and undetectable zero-day exploits being used by attackers, the detection techniques of intrusion are increasingly difficult. Motivated by the fact that the attackers usually penetrate the social network to either propagate malwares or collect sensitive information, we propose a method to assess the security risk of the user being attacked so that we can take defensive measures such as security education, training, and awareness before users are attacked. In this paper, we propose a novel user analysis model to find potential victims by analyzing a large number of users’ personal information and social behaviors in social networks. For each user, we extract three kinds of features, i.e., statistical features, social-graph features, and semantic features. These features will become the input of our user analysis model, and the security risk score will be calculated. The users with high security risk score will be alarmed so that the risk of being attacked can be reduced. We have implemented an effective user analysis model and evaluated it on a real-world dataset collected from a social network, namely, Sina Weibo (Weibo). The results show that our model can effectively assess the risk of users’ activities in social networks with a high area under the ROC curve of 0.9607.
Qiang Li 0008, Yuede Ji, Dong Guo 0002, Xiangyu Meng 0002
Secur. Commun. Networks2
2018 Combating the evolving spammers in online social networks
Dong Guo 0002, Qiang Li 0008
Comput. Secur.4
2018 SIoTFog: Byzantine-resilient IoT fog networking
abstract
The current boom in the Internet of Things (IoT) is changing daily life in many ways, from wearable devices to connected vehicles and smart cities. We used to regard fog computing as an extension of cloud computing, but it is now becoming an ideal solution to transmit and process large-scale geo-distributed big data. We propose a Byzantine fault-tolerant networking method and two resource allocation strategies for IoT fog computing. We aim to build a secure fog network, called “SIoTFog,” to tolerate the Byzantine faults and improve the efficiency of transmitting and processing IoT big data. We consider two cases, with a single Byzantine fault and with multiple faults, to compare the performances when facing different degrees of risk. We choose latency, number of forwarding hops in the transmission, and device use rates as the metrics. The simulation results show that our methods help achieve an efficient and reliable fog network.
Jianwen Xu, Kaoru Ota, Mianxiong Dong, Anfeng Liu, Qiang Li 0008
Frontiers Inf. Technol. Electron. Eng.5
2018 Towards fast and lightweight spam account detection in mobile social networks through fog computing
Qiang Li 0008, Dong Guo 0002
Peer-to-Peer Netw. Appl.2
2017 Discovering hidden suspicious accounts in online social networks
Qiang Li 0008, Dong Guo 0002
Inf. Sci.3
2016 Combating the evasion mechanisms of social bots
Yuede Ji, Xinyang Jiang, Qiang Li 0008
Comput. Secur.5
2016 Understanding a prospective approach to designing malicious social bots
abstract
Abstract The security implications of social bots are evident in consideration of the fact that data sharing and propagation functionality are well integrated with social media sites. Existing social bots primarily use Really Simple Syndication and OSN (online social network) application program interface to communicate with OSN servers. Researchers have profiled their behaviors well and have proposed various mechanisms to defend against them. We predict that a web test automation rootkit (WTAR) is a prospective approach for designing malicious social bots. In this paper, we first present the principles of designing WTAR‐based social bots. Second, we implement three WTAR‐based bot prototypes on Facebook, Twitter, and Weibo. Third, we validate this new threat by analyzing behaviors of the prototypes in a lab environment and on the Internet, and analyzing reports from widely‐used antivirus software. Our analyses show that WTAR‐based social bots have the following features: (i) they do not connect to OSN directly, and therefore produce few network flows; (ii) they can log in to OSNs easily and perform a variety of social activities; (iii) they can mimic the behaviors of a human user on an OSN. Finally, we propose several possible mechanisms in order to defend against WTAR‐based social bots. Copyright © 2016 John Wiley & Sons, Ltd.
Guangyan Zhang, Jie Wu 0001, Qiang Li 0008
Secur. Commun. Networks4
2015 Leveraging Behavior Diversity to Detect Spammers in Online Social Networks
Qiang Li 0008, Dong Guo 0002
ICA3PP (3)3
2015 BotCatch: leveraging signature and behavior for bot detection
abstract
Abstract The goal of bot detection is to discover malicious bot processes by signature comparison or behavior analysis. Existing approaches have several drawbacks, such as requiring a lot of prior knowledge, low detection accuracy, and high false alarm rate. In this paper, we propose a multi‐feedback approach, BotCatch, to detect bots effectively and efficiently on a host by leverage of a combination of signature and behavior. First, BotCatch assigns suspicious files to signature‐analysis and behavior‐analysis modules, which generate each detection result. Second, BotCatch correlates signature and behavior results to generate the final detection result through correlation engine. Third, BotCatch feeds back signature, behavior, and correlation results to dynamically adjust detecting modules through multi‐feedback engine. We evaluated the performance of BotCatch with 636 bot and 150 benign samples. Our results indicate that BotCatch achieves an accuracy of 97.1%and an F‐measure value of 0.982 simultaneously, which is better than existing approaches without feedbacks. BotCatch, due to the multi‐feedback mechanism, has the ability to gradually get more robust and accurate as the number of samples increases. The final stage even reaches an accuracy of 98.5%and F‐measure value of 0.991. Copyright © 2014 John Wiley & Sons, Ltd.
Yuede Ji, Qiang Li 0008, Dong Guo 0002
Secur. Commun. Networks2
2014 Towards social botnet behavior detecting in the end host
abstract
Social botnet utilizing online social network (OSN) as Command and Control channel (C&C) has caused enormous threats to Internet security. Server-side detection approaches mainly target on suspicious accounts, which cannot identify the specific bot hosts or processes. Host-side approaches target on suspicious process behaviors which are not robust enough to face the challenges of frequent variants and novel social bots. In this paper, we propose a novel social bot behavior detecting approach in the end host. Because social bot binaries or source codes are not easy to collect, we first design a novel social botnet, named wbbot, based on Sina Weibo. We analyze it from two aspects, wbbot architecture and wbbot behaviors. Second, we analyze the host behaviors of existing social botnets which come from public websites, other researchers, and our implementations. We identify six critical phases: infection, pre-defined host behaviors, establishment of C&C, receive the commands of botmaster, execution of social bot commands, and return the results. Third, we present our detection system which consists of three components: host behavior monitor, host behavior analyzer, and detection approach. We present behavior tree-based approach to detect social bot. After constructing the suspicious behavior tree, we match it with the template library to generate detection result. Finally, we collect real-world social botnet traces to evaluate the performance. We would like to share them for academic research. The results indicate that our system has an acceptable false positive rate of 29.6% and remarkable false negative rate of 4.5%. However, compared with other detection tools, our detection result is still remarkable.
Yuede Ji, Xinyang Jiang, Qiang Li 0008
ICPADS4
2014 A Mulitiprocess Mechanism of Evading Behavior-Based Bot Detection Approaches
Yuede Ji, Dewei Zhu, Qiang Li 0008, Dong Guo 0002
ISPEC4
2013 BotInfer: A Bot Inference Approach by Correlating Host and Network Information
Qiang Li 0008, Yuede Ji, Dong Guo 0002
NPC2
2009 Reconstruction of Worm Propagation Path by Causality
abstract
Fast and accurate online tracing of network worm during its propagation is essential for worm containment and reducing the loss. Though worm is randomly spread, there exists implicit causality between adjacent infected nodes. Using this causality can help to enhance the performance of worm tracing algorithm. Bayesian Network can be a very good probability description of the current results and prior conditions. Based on the analysis of causality, we present an improved online tracing algorithm -- Bayesian Network Correlation Algorithm to acquire worm propagation path, and analyze and verify its accuracy and performance through simulation experiments. Experiment result indicates that the detection accuracy of Bayesian Network Correlation Algorithm has risen by 10% compared to our previous work, this improved algorithm is more suitable for online detection.
Qiang Li 0008, Dong Guo 0002
NAS2
2008 Online Accumulation: Reconstruction of Worm Propagation Path
Qiang Li 0008, Dong Guo 0002
NPC2
2007 Online Tracing Scanning Worm with Sliding Window
Qiang Li 0008
Inscrypt2
2006 Succinct Text Indexes on Large Alphabet
Meng Zhang 0006, Jijun Tang, Dong Guo 0002, Liang Hu 0001, Qiang Li 0008
TAMC5
2005 Constructing Correlations in Attack Connection Chains Using Active Perturbation
Qiang Li 0008, Jiubin Ju
AAIM1
2005 Weighted Directed Word Graph
Meng Zhang 0006, Liang Hu 0001, Qiang Li 0008, Jiubin Ju
CPM3
2005 Fast Two Phrases PPM for IP Traceback
abstract
In probabilistic packet marking(PPM) for IP tracebacking, the number of packets that are needed to reconstruct the attacking paths depends on the precision of the attacking paths. In this paper, a Fast Two Phrases(FTP) PPM for IP tracebacking is proposed, which depends on the division of Autonomous System(AS) and two algorithms are used to reconstruct the attacking paths. It can reconstruct the exactly attacking paths between AS when it has received tens of packets, and reconstruct the attacking paths in AS after receiving more packets. This method can reduce the number of packets that are needed to reconstruct the attacking paths to the lowest while reducing the complexity of packet marking and reconstructing.
Qiang Li 0008, Qinyuan Feng, Liang Hu 0001, Jiubin Ju
PDCAT1
2005 Simulating and Improving Probabilistic Packet Marking Schemes Using Ns2
abstract
Simulation environments and approaches for evaluating real-time of IP traceback in different network scenarios and attacking patterns are very important. A comparison among some of the most promising PPM (Probabilistic Packet Marking) schemes is presented with several metrics, including the received packet number required for reconstructing the attacking path, computation complexity and false positive etc. We constructe a simulation environment via extending ns2, setting attacking topology and traffic, which can be used to evaluate and compare the effectiveness of different PPM schemes. The simulation approach also can be used to test the performing effects of different PPM schemes in large-scale DDoS attacks. Based on the simulation and evaluation results, several improvable aspects of PPM are proposed, which can increase real-time of IP traceback efficiently.
Qiang Li 0008, Hongzi Zhu, Meng Zhang 0006, Jiubin Ju
PDCAT1