VLDB 2026 Research / reviewers in the wild / expert
Hui Lu 0005
dblp:65/4062-5
· DBLP profile ↗
25ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0002-4120-7716ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Security and privacy · 3 · 3 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | C2graph: A Compression-Collaboration Algorithm for CPU-GPU Hybrid Weighted Graph Traversals
Ning Wang 0026, Huaibei Li, Shen Su, Yu Gu 0002, Ge Yu 0001, Zhigang Wang 0001, Dawei Zhao 0001, Hui Lu 0005, Zhihong Tian 0001 |
ICDE | 8 |
| 2026 | A Trustworthy Federated Learning Framework for Joint Privacy and Byzantine ResilienceabstractFederated Learning (FL) has emerged as a promising paradigm for collaborative model training without direct data sharing. However, existing frameworks struggle to simultaneously ensure strong privacy protection and robustness against adversarial participants. Although numerous robust aggregation algorithms can mitigate Byzantine attacks and various privacy-preserving approaches have been developed, achieving both objectives within a unified and efficient framework remains challenging. To address this compatibility gap, we propose a generic privacy-preserving FL framework that integrates robust aggregation with cryptographic privacy guarantees through secure two-party computation under a dual-server architecture. In our design, two non-colluding servers collaboratively execute the aggregation process, ensuring that sensitive client updates remain confidential while enabling the application of diverse robust aggregation strategies. The framework is flexible and supports a wide range of aggregation algorithms, such as Multi-Krum, Median, and Mean-based methods, under a consistent security model. We implement the proposed system and conduct extensive experiments on multiple benchmark datasets. Experimental results demonstrate that our framework preserves the effectiveness of existing robust aggregation algorithms while maintaining acceptable runtime and communication overhead compared with standard FL baselines. These findings confirm that the proposed approach provides a practical and balanced solution to the dual challenges of privacy protection and Byzantine robustness, offering a versatile foundation for secure and trustworthy FL in adversarial environments. Jiangang Shu, Zhiping Hu, Fuyi Wang, Yuyu He 0001, Hui Lu 0005, Zhihong Tian 0001 |
IEEE Internet Things J. | 5 |
| 2026 | Resilient Load Frequency Control for Multi-Region Power Systems Against False Data Injection and Asynchronous Leakage DelaysabstractThe deceptive nature of false data injection (FDI) attacks and the hysteresis effect of asynchronous leakage delay (ALD) can weaken the resilient load frequency control (RLFC) system’s dynamic response and stability. Therefore, this paper investigates the system dynamic fluctuations caused by the coupling of FDI and ALD. First, a dynamic attack model is established. An attack modulator (AM) is introduced to dynamically adjust the attack process to accurately simulate attacks on critical control paths. Second, we systematically characterize the asynchronous leakage delay caused by multi-source communication link switching. This delay behavior is embedded into the control system. This enables the system to more comprehensively reflect feedback delay characteristics under complex communication environments. Next, a robust control method is designed to enhance the system’s response to FDI and ALD under these complex situations, which can effectively ensure the stability of the system. Then, the improved Lyapunov-Krasovskii functional is used to analyze the system performance. Finally, simulations of a dual-region power system are conducted. The simulation results verify the effectiveness of the proposed method in terms of disturbance immunity, control stability, and engineering adaptability. Hailing Sun, Jun Wang 0128, Kaibo Shi, Lanfeng Hua, Yanbin Sun, Hui Lu 0005 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | Accelerating Heterogeneous Tensor Parallelism via Flexible Workload ControlabstractTransformer-based foundation models are becoming deeper and larger. For fast training, their billions of parameters (tensors) are split onto parallel tasks running on many modern yet expensive accelerators. To amortize the huge hardware investments, it is cost-effective to share the aggregated resources among multi-tenants. However, resource contention yields the heavy straggling problem. Existing works feature contributions for the traditional data parallelism. They cannot work well for the new tensor parallelism, due to the dependency among split tensors. This paper is the first attempt on accelerating heterogeneous tensor parallelism. We summarize specific challenges, including the very frequent synchronizations and the heavy tensor computation workloads. Our solution is to resize dimensions of parameters on demand, to quickly and dynamically balance workloads. The accuracy loss is reduced through priority resizing. We also migrate workloads between tasks, without any loss of accuracy. The most efficient communication primitives are selected and then scheduled in a non-redundant manner, to reduce the runtime latency. Our final hybrid solution is built on top of resizing and migration. By studying the tradeoff between accuracy and efficiency, it can smartly hit the “sweet spot”. Extensive experiments validate the effectiveness of our proposals. Zhigang Wang 0001, Ning Wang 0026, Chuanfei Xu, Yu Gu 0002, Hui Lu 0005, Dawei Zhao 0001, Zhihong Tian 0001 |
IEEE Trans. Big Data | 6 |
| 2026 | D3-Guard: An Adaptive Decision Fusion Framework for Cloud-Native Threat DetectionabstractTraditional threat detection systems fail in dynamic cloud-native environments, primarily due to their inability to resolve the inherent trade-off between high-precision detection of known attacks and high-recall discovery of unknown threats. We proposeD3-Guard, a novel framework whose core innovation lies in a dynamic decision fusion engine grounded in Bayesian decision-theoretic principles. The engine computes a similarity score from incoming requests to serve as a real-time contextual signal for the Bayesian framework. It then uses this signal directly as a dynamic weight to adaptively integrate the outputs of two complementary expert models: a high-precision LSTM-Attention for known patterns and a high-recall Isolation Forest for unknown threats. This synergy, reinforced by an adaptive update layer to mitigate concept drift, achieves F1-scores of 99.39% and 96.97% on the CSIC 2010 and ATRDF 2023 datasets, respectively. Performance tests in a production-like environment further validate its practical viability, demonstrating that this high accuracy is achieved with a justifiable overhead. An ablation study highlights the pivotal role of our fusion engine: disabling it results in a sharp decline in F1-score on the ATRDF 2023 dataset, from 96.97% to 11.47%, demonstrating its critical contribution to robust, dual-model collaboration. Weiyong Zhang, Hao Liu 0058, Hui Lu 0005, Zhihong Tian 0001 |
IEEE Trans. Cloud Comput. | 4 |
| 2026 | Weakly Supervised Learning Meets VIS-NIR Re-Identification: Progressive Cross-Modality Alignment for Representation Learning
Zefeng Lu, Yunfei He, Meijun Chen, Hui Lu 0005, Yuyu He 0001, Zhihong Tian 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | Solo-Adaptive Reliable Synchronization of Linear Multiagent Systems Under Composite Communication Link FaultsabstractThis article investigates reliable leader-follower synchronization of linear multi-agent systems (MASs) under composite communication faults. Firstly, a composite double-Bernoulli link model is proposed to characterize inter-layer and intra layer communications. It captures communication mismatch by allowing independent Bernoulli link states with distinct stochastic reliabilities across the two interaction layers. Secondly, a passive predictor-reset mechanism is embedded into a modified distributed measurement to address packet losses. This mechanism efficiently leverages locally predicted surrogate values to enable uninterrupted distributed feedback computation and resets these surrogates upon data reception. Furthermore, a reliable solo adaptive protocol is developed, where each follower updates only one scalar adaptive gain together with a reliability-oriented baseline parameter. In addition, a Lyapunov-Krasovskii func tional (LKF) with an inverse gain weighting is constructed to establish the mean-square ultimate boundedness of the closed loop synchronization error. This result provides an explicit reliability guarantee under composite link faults, including low probability link activations and packet losses. Finally, numerical simulations demonstrate the effectiveness and resilience of the proposed scheme. Kaibo Shi, Yue Yu 0013, Shiping Wen 0001, Yanbin Sun, Hui Lu 0005 |
IEEE Trans. Reliab. | 6 |
| 2025 | Context-Aware Phishing-Resistant Authentication for Federated Identity in Internet of Things PlatformsabstractThe Internet of Things (IoT) has seen widespread adoption across various industries, enabling cross-domain collaboration among devices. As IoT platforms that allows users to manage and control these IoT devices become more prevalent, the security of IoT platforms has come under increasing scrutiny. Cryptographic authentication (CA) has long been central to authentication but poses significant security risks, such as account takeovers due to credential disclosure. Unfortunately, CA remains the most popular authentication scheme for IoT platforms, making them vulnerable to phishing attacks. Current enhancement methods, including challenge-response protocols and multifactor authentication, are still susceptible to phishing and face usability issues, particularly in cross-domain collaboration systems where seamless access is crucial. This article introduces Context-Auth, a context-aware authentication scheme designed to defend against phishing attacks without requiring tamper-resistant hardware. By incorporating contextual features of authentication and access behaviors into access credentials, Context-Auth detects inconsistencies and strengthens security. Our scheme addresses cross-domain verification challenges in federated identity scenarios, ensuring high usability across various devices without additional requirement from users. We conducted comprehensive evaluations and security analyses to ensure that Context-Auth remains secure against potential threats. The experimental results demonstrate that Context-Auth effectively defends against phishing attacks launched by real-world phishing kits. Jiageng Yang, Binxing Fang, Hui Lu 0005, Zhihong Tian 0001 |
IEEE Internet Things J. | 3 |
| 2025 | Enhancing Container Security Through Phase-Based System Call FilteringabstractContainer technology in cloud computing has improved resource utilization and deployment efficiency, but it also introduces new security risks. Excessive privileges in containerized environments can allow attackers to exploit insufficiently restricted system calls, potentially leading to container escapes and other attacks. Some system calls are only necessary during the initialization phase of a container, and allowing them during runtime can increase the risk of exploitation. This paper proposes a Phase-based System Call Filtering (PSF) method to minimize system call permissions during the runtime of cloud containers. The PSF method builds a comprehensive whitelist of system calls during the initialization phase to cover all necessary calls for containerized applications. During runtime, a refined, phase-specific system call whitelist is enforced, dynamically adjusting privileges based on the functions encapsulated within the container. Additionally, we introduce a container phase recognition algorithm to distinguish between initialization and runtime phases, supporting the generation of phase-specific system call lists. Experimental results show that the proposed method enhances runtime system call restrictions, minimizes privileges, and improves the overall security of cloud containers. Hui Lu 0005, Yinnan Yao, Binxing Fang, Yuan Liu 0002, Zhihong Tian 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2025 | BPFGuard: Multi-Granularity Container Runtime Mandatory Access ControlabstractThe adoption of container-based cloud computing services has been prevalent, especially with the introduction of Kubernetes, which enables the automated deployment, scaling, and administration of applications in containers, hence boosting the popularity of containers. As a result, researchers have placed greater emphasis on container runtime security, notably investigating the efficacy of traditional techniques such as Capabilities, Seccomp, and Linux security modules in guaranteeing container security. However, due to the limitations imposed by the container environment, the results have been unsatisfactory. In addition, eBPF-based solutions face the problem of being unable to quickly load policies and affect real-time operations when faced with newer kernel vulnerabilities. This paper investigates the limitations of existing container security mechanisms. Additionally, it examines the specific constraints of these mechanisms in Kubernetes environments. The paper classifies container monitoring and obligatory access control into three distinct categories: system call access control, LSM hook access control, and kernel function access control. Therefore, we propose a technique for regulating container access with a variety of granularity levels. This technique is executed using eBPF and is tightly integrated with Kubernetes to collect relevant meta-information. In addition, we suggest implementing a consolidated routing method and employing function tail call chaining to overcome the limitation of eBPF in enforcing mandatory access control for containers. Lastly, we conducted a series of experiment to verify the effectiveness of the system's security using CVE-2022-0492 and to benchmark the system that had BPFGuard enabled. The results indicate that the average performance loss increased merely by 2.16%, demonstrating that there are no adverse effects on the container services. This suggests that greater security can be achieved at a minimal cost. Hui Lu 0005, Xiaojiang Du, Dawei Hu, Shen Su, Zhihong Tian 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2024 | Active and Passive Attack Detection Methods for Malicious Encrypted TrafficabstractThe encryption of traffic data offers a means of protecting the security of data and the private information of the public. However, this same technology also presents a potential avenue for attackers to conceal malicious activities. Attackers hide malicious behaviour in encrypted traffic data to bypass detection by firewalls or early intrusion detection systems (IDS). In order to cope with malicious encrypted traffic, traffic attack detection is classified into active and passive detection depending on the way it is handled. Active detection is mainly based on searchable traffic detection and traffic plaintext data parsing. Measures such as analysing controllable transmission protocols and trusted execution environments are used to ensure both attack detection efficiency and privacy security. Passive detection focuses on feature construction of encrypted traffic. The characterisation of traffic data from multiple perspectives, including channel and context, is achieved through the utilisation of machine learning or deep learning models, thereby facilitating the generation of accurate prediction outcomes. This paper offers an overview of existing detection methods from two distinct vantage points: active attack detection and passive attack detection. Finally, the paper presents a summary of the strengths and weaknesses of existing methods and suggests potential avenues for future research. Hui Lu 0005, Houlin Zhou, Chengcong Zheng, Zhihong Tian 0001, Xiaojiang Du |
WiMob | 2 |
| 2023 | Smart Contract Firewall: Protecting the on-Chain Smart Contract ProjectsabstractThe burgeoning landscape of blockchain technology has made the security of deployed smart contracts an imperative concern. While existing security measures excel in pre-deployment testing, they fall short in protecting smart contracts once they are deployed, leaving them susceptible to malicious attacks. In this paper, we propose a novel Smart Contract Firewall framework designed to bridge this security gap. Functioning as a dynamic gateway, the framework employs real-time transaction inspection through adaptable filtering rules, enabling the identification and rollback of malicious transactions as they occur. Our empirical analysis demonstrates the framework's efficacy in mitigating a majority of existing vulnerabilities in the deployed smart contracts. Although the added layer of security comes at a cost, we prove that the increased gas expenses could be limited to 30 % -50 % for most transactions. This trade-off, we argue, is a small price to pay for significantly enhanced security. Shen Su, Yue Xue, Liansheng Lin, Hui Lu 0005, Jing Qiu 0002, Yanbin Sun, Yuan Liu 0002, Zhihong Tian 0001 |
GLOBECOM | 5 |
| 2023 | Improving Precision of Detecting Deserialization Vulnerabilities with Bytecode AnalysisabstractTraditional static taint analysis based on bytecode analysis such as GadgetInspector to detect deserialization vulnerabilities always faced precision problems. For example, missing the fact that taints flowing to members in called methods, type confusion, and chaotic inheritance relationships when detecting deserialization vulnerabilities, which would lead to many error results. To alleviate these problems, this paper considers three measures of improving precision of detecting deserialization vulnerabilities, including cross-function members data flow tracking, local variables and arguments types inference, and call chain subject inference based on inheritance relationships. Weicheng Li, Hui Lu 0005, Yanbin Sun, Shen Su, Jing Qiu 0002, Zhihong Tian 0001 |
IWQoS | 2 |
| 2023 | Do Not Trust the Clouds Easily: The Insecurity of Content Security Policy Based on Object StorageabstractThe content security policy (CSP) is a World Wide Web Consortium (W3C) standard, designed to prevent and mitigate security vulnerabilities, such as cross-site scripting (XSS) attacks, data injection attacks, and clickjacking attacks on websites. In this article, we present a newly discovered front-end Web attack that uses the current object storage services vulnerability of cloud vendors to bypass CSP. We selected the object storage services from two cloud vendors with the most users, i.e., Google and Amazon, to conduct systematic and large-scale research and analysis. Three cyberspace search engines are used to retrieve data, from which we analyze the consequence and damage range of this security breach. We focus on reporting four key aspects of this security breach: 1) how to use object storage services to bypass CSP; 2) analysis on the existence of such vulnerability in real-world websites; 3) analysis on the existing security vulnerabilities in current object storage services; and 4) the new strategy on object storage services that we propose to use to eliminate the discovered security threat. Yangzixing Lv, Wei Shi 0001, Weiyong Zhang, Hui Lu 0005, Zhihong Tian 0001 |
IEEE Internet Things J. | 4 |
| 2022 | A novel flow-vector generation approach for malicious traffic detection
Jian Hou 0009, Fang'ai Liu, Hui Lu 0005, Zhiyuan Tan 0001, Xuqiang Zhuang, Zhihong Tian 0001 |
J. Parallel Distributed Comput. | 3 |
| 2021 | Research on Intelligent Detection of Command Level Stack Pollution for Binary Program Analysis
Hui Lu 0005, Chengjie Jin, Xiaohan Helu, Man Zhang 0004, Yanbin Sun, Zhihong Tian 0001 |
Mob. Networks Appl. | 1 |
| 2021 | Controlled Sharing Mechanism of Data Based on the Consortium BlockchainabstractIn the process of sharing data, the costless replication of electric energy data leads to the problem of uncontrolled data and the difficulty of third-party access verification. This paper proposes a controlled sharing mechanism of data based on the consortium blockchain. The data flow range is controlled by the data isolation mechanism between channels provided by the consortium blockchain by constructing a data storage consortium chain to achieve trusted data storage, combining attribute-based encryption to complete data access control and meet the demands for granular data accessibility control and secure sharing; the data flow transfer ledger is built to record the original data life cycle management and effectively record the data transfer process of each data controller. Taking the application scenario of electric energy data sharing as an example, the scheme is designed and simulated on the Linux system and Hyperledger Fabric. Experimental results have verified that the mechanism can effectively control the scope of access to electrical energy data and realize the control of the data by the data owner. Songqi Wu, Yundan Yang, Fenghui Duan, Hui Lu 0005, Yueming Lu |
Secur. Commun. Networks | 5 |
| 2021 | StFuzzer: Contribution-Aware Coverage-Guided Fuzzing for Smart DevicesabstractThe root cause of the insecurity for smart devices is the potential vulnerabilities in smart devices. There are many approaches to find the potential bugs in smart devices. Fuzzing is the most effective vulnerability finding technique, especially the coverage-guided fuzzing. The coverage-guided fuzzing identifies the high-quality seeds according to the corresponding code coverage triggered by these seeds. Existing coverage-guided fuzzers consider that the higher the code coverage of seeds, the greater the probability of triggering potential bugs. However, in real-world applications running on smart devices or the operation system of the smart device, the logic of these programs is very complex. Basic blocks of these programs play a different role in the process of application exploration. This observation is ignored by existing seed selection strategies, which reduces the efficiency of bug discovery on smart devices. In this paper, we propose a contribution-aware coverage-guided fuzzing, which estimates the contributions of basic blocks for the process of smart device exploration. According to the control flow of the target on any smart device and the runtime information during the fuzzing process, we propose the static contribution of a basic block and the dynamic contribution built on the execution frequency of each block. The contribution-aware optimization approach does not require any prior knowledge of the target device, which ensures our optimization adapting gray-box fuzzing and white-box fuzzing. We designed and implemented a contribution-aware coverage-guided fuzzer for smart devices, called StFuzzer. We evaluated StFuzzer on four real-world applications that are often applied on smart devices to demonstrate the efficiency of our contribution-aware optimization. The result of our trials shows that the contribution-aware approach significantly improves the capability of bug discovery and obtains better execution speed than state-of-the-art fuzzers. Jiageng Yang, Xinguo Zhang, Hui Lu 0005, Muhammad Shafiq 0003, Zhihong Tian 0001 |
Secur. Commun. Networks | 3 |
| 2021 | Adaptive Feature Selection and Construction for Day-Ahead Load Forecasting Use Deep Learning MethodabstractAs a networked cyber physical system, smart grid plays an important role in improving energy efficiency, which has higher prediction accuracy requirement for load forecasting to realize intelligent dispatching. The existing works of load forecasting have considered the external factors to improve prediction accuracy, but ignored their varying and different impacts in different time. In order to solve this problem, firstly, a new method is designed to select external meteorological factors and construct new features to deal with the different impacts of external factors in different seasons. Secondly, a fuzzy processing method of calendar factor is proposed to explore the varying impact of the calendar factor in the holiday period. Finally, our proposed method is combined with LSTM and other three typical machine learning models, and experiments are performed on a load data set from the real power system. The results show that our proposed method can improve the prediction accuracy of the typical machine learning based models. Especially for LSTM, the MAPEs reach to 3.8% and 3.1% in winter and summer, with the improvements of 0.6% and 0.3% respectively. Runhai Jiao, Shuangkun Wang, Hui Lu 0005, Brij B. Gupta |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2020 | Inferring Passengers' Interactive Choices on Public Transits via MA-AL: Multi-Agent Apprenticeship LearningabstractPublic transports, such as subway lines and buses, offer affordable ride-sharing services and reduce the road network traffic. Extracting passengers’ preferences from their public transit choices is important to city planners but technically non-trivial. When traveling by taking public transits, passengers make sequences of transit choices, and their rewards are usually influenced by other passengers’ choices. This process can be modeled as a Markov Game (MG). In this paper, we make the first effort to model travelers’ preferences of making transit choices using MGs. Based on the discovery that passengers usually do not change their policies, we propose novel algorithms to extract reward functions from the observed deterministic equilibrium joint policy of all agents in a general-sum MG to infer travelers’ preferences. First, we assume we have the access to the entire joint policy. We characterize the set of all reward functions for which the given joint policy is a Nash equilibrium policy. In order to remove the degeneracy of the solution, we then attempt to pick reward functions so as to maximize the sum of the deviation between the the observed policy and the sub-optimal policy of each agent. This results in a skillfully solvable linear programming algorithm for the multi-agent inverse reinforcement learning (MA-IRL) problem. Then, we deal with the case where we have access to the equilibrium joint policy through a set of actual trajectories. We propose an iterative algorithm inspired by single-agent apprenticeship learning algorithms and the cyclic coordinate descent approach. We evaluate the proposed algorithms on both a simple Grid Game and a unique real-world dataset (from Shenzhen, China). Results show that when we have access to the full policy, our algorithm can efficiently recover most of the reward structure, especially the interaction of agents. In the case where we only have access to a set of sampled expert trajectories, our algorithm can provide an explanation of the expert trajectories. Measured with respect to the experts’ unknown reward function, the performance of the policy output by our algorithm is close to that of the expert policy. Mingzhou Yang 0001, Xun Zhou 0001, Hui Lu 0005, Zhihong Tian 0001, Jun Luo 0007 |
WWW | 4 |
| 2020 | Deep Reinforcement Learning for Partially Observable Data Poisoning Attack in Crowdsensing SystemsabstractCrowdsensing systems collect various types of data from sensors embedded on mobile devices owned by individuals. These individuals are commonly referred to as workers that complete tasks published by crowdsensing systems. Because of the relative lack of control over worker identities, crowdsensing systems are susceptible to data poisoning attacks which interfering with data analysis results by injecting fake data conflicting with ground truth. Frameworks like TruthFinder can resolve data conflicts by evaluating the trustworthiness of the data providers. These frameworks somehow make crowdsensing systems more robust since they can limit the impact of dirty data by reducing the value of unreliable workers. However, previous work has shown that TruthFinder may also be affected by the data poisoning attack when the malicious workers have access to global information. In this article, we focus on partially observable data poisoning attacks in crowdsensing systems. We show that even if the malicious workers only have access to local information, they can find effective data poisoning attack strategies to interfere with crowdsensing systems with TruthFinder. First, we formally model the problem of partially observable data poisoning attack against crowdsensing systems. Then, we propose a data poisoning attack method based on deep reinforcement learning, which helps malicious workers jeopardize with TruthFinder while hiding themselves. Based on the method, the malicious workers can learn from their attack attempts and evolve the poisoning strategies continuously. Finally, we conduct experiments on real-life data sets to verify the effectiveness of the proposed method. Mohan Li, Yanbin Sun, Hui Lu 0005, Sabita Maharjan, Zhihong Tian 0001 |
IEEE Internet Things J. | 3 |
| 2020 | ConnSpoiler: Disrupting C&C Communication of IoT-Based Botnet Through Fast Detection of Anomalous Domain QueriesabstractThe development of Internet of Things (IoT) dramatically facilitates the integration of computing systems with the physical world. However, as IoT devices are more easy to compromise than desktop computers, cybercriminals have founded IoT-based botnets to launch Distributed Denial of Service (DDoS) attacks with unprecedented traffic volume. To mitigate the damages associated with these attacks, the detection of IoT-based botnet has to preempt the command and control (C&C) communication to prevent the delivery of the attack codes. Motivated by the extensively implementation of domain generation algorithm in botnets, in this article, we propose ConnSpoiler, a lightweight system that detects IoT-based botnets by identifying the stream of algorithmically generated domains (AGDs) in a fast way. ConnSpoiler only needs negligible system resources to take effect and thus can execute well on the resource-restraint IoT devices. By outfitting a powerful statistical algorithm, i.e., threshold random walk, ConnSpoiler has a high probability (about 94%) of detecting infection before the compromised devices connect C&C servers, which can help to prevent the succeeding attacks. Moreover, ConnSpoiler only requires the benign domains to take effect and therefore does not need extra effort to label malicious samples for training phase. We evaluate ConnSpoiler based on real-world DNS traffics collected from two different large ISP networks and show that it accurately identifies devices that are compromised by unknown botnets. Lihua Yin, Chunsheng Zhu, Liming Wang 0001, Zhen Xu 0009, Hui Lu 0005 |
IEEE Trans. Ind. Informatics | 6 |
| 2020 | DHPA: Dynamic Human Preference Analytics Framework: A Case Study on Taxi Drivers' Learning Curve AnalysisabstractMany real-world human behaviors can be modeled and characterized as sequential decision-making processes, such as a taxi driver’s choices of working regions and times. Each driver possesses unique preferences on the sequential choices over time and improves the driver’s working efficiency. Understanding the dynamics of such preferences helps accelerate the learning process of taxi drivers. Prior works on taxi operation management mostly focus on finding optimal driving strategies or routes, lacking in-depth analysis on what the drivers learned during the process and how they affect the performance of the driver. In this work, we make the first attempt to establish Dynamic Human Preference Analytics. We inversely learn the taxi drivers’ preferences from data and characterize the dynamics of such preferences over time. We extract two types of features (i.e., profile features and habit features) to model the decision space of drivers. Then through inverse reinforcement learning, we learn the preferences of drivers with respect to these features. The results illustrate that self-improving drivers tend to keep adjusting their preferences to habit features to increase their earning efficiency while keeping the preferences to profile features invariant. However, experienced drivers have stable preferences over time. The exploring drivers tend to randomly adjust the preferences over time. Menghai Pan, Weixiao Huang, Xun Zhou 0001, Zhenming Liu, Rui Song 0006, Hui Lu 0005, Zhihong Tian 0001, Jun Luo 0007 |
ACM Trans. Intell. Syst. Technol. | 7 |
| 2019 | Effective Recycling Planning for Dockless Sharing BikesabstractBike-sharing systems become more and more popular in the urban transportation system, because of their convenience in recent years. However, due to the high daily usage and lack of effective maintenance, the number of bikes in good condition decreases significantly, and vast piles of broken bikes appear in many big cities. As a result, it is more difficult for regular users to get a working bike, which causes problems both economically and environmentally. Therefore, building an effective broken bike prediction and recycling model becomes a crucial task to promote cycling behavior. In this paper, we propose a predictive model to detect the broken bikes and recommend an optimal recycling program based on the large scale real-world sharing bike data. We incorporate the realistic constraints to formulate our problem and introduce a flexible objective function to tune the trade-off between the broken probability and recycled numbers of the bikes. Finally, we provide extensive experimental results and case studies to demonstrate the effectiveness of our approach. Cong Zhang 0003, Jie Bao 0003, Sijie Ruan, Tianfu He, Hui Lu 0005, Zhihong Tian 0001, Cong Liu 0005, Jianfeng Lin 0004, Xianen Li |
SIGSPATIAL/GIS | 6 |
| 2019 | Dissecting the Learning Curve of Taxi Drivers: A Data-Driven ApproachabstractMany real world human behaviors can be modeled and characterized as sequential decision making processes, such as taxi driver's choices of working regions and times. Each driver possesses unique preferences on the sequential choices over time and improves their working efficiency. Understanding the dynamics of such preferences helps accelerate the learning process of taxi drivers. Prior works on taxi operation management mostly focus on finding optimal driving strategies or routes, lacking in-depth analysis on what the drivers learned during the process and how they affect the performance of the driver. In this work, we make the first attempt to inversely learn the taxi drivers' preferences from data and characterize the dynamics of such preferences over time. We extract two types of features, i.e., profile features and habit features, to model the decision space of drivers. Then through inverse reinforcement learning we learn the preferences of drivers with respect to these features. The results illustrate that self-improving drivers tend to keep adjusting their preferences to habit features to increase their earning efficiency, while keeping the preferences to profile features invariant. On the other hand, experienced drivers have stable preferences over time. Menghai Pan, Xun Zhou 0001, Zhenming Liu, Rui Song 0006, Hui Lu 0005, Jun Luo 0007 |
SDM | 6 |