Zhou Zhou 0007

dblp:67/2535-7 · DBLP profile ↗
← Back
33ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0001-6924-5848ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 16 · 3 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 7 · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AdaTable: Learning to Calibrate for Robust VNF Auto-scaling under Capacity Drift
Weikang Huang, Chenkan Wang, Zhou Zhou 0007, Rong Yang 0008, Qingyun Liu 0001
ICIC (15)5
2026 You can run but you can never hide: A multi-module collaborative detection framework based on network traffic
Chengxiang Si, Zhou Zhou 0007, Zhenyu Cheng 0001, Peishuai Sun
Comput. Networks3
2025 MS-NHHO: A Swarm Intelligence Optimization Algorithm Incorporating Cognitive Science for Malicious Traffic Detection
Zhou Zhou 0007, Chengxiang Si, Qingyun Liu 0001
CogSci2
2025 DygLM: Detecting Lateral Movement Threat via Continuous-time Dynamic Graph Learning
abstract
In an advanced persistent threat (APT) targeting an enterprise, an adversary attempts to move through the network environment starting from their initial point, which is known as Lateral Movement (LM). Detecting LM typically uses authentication logs, which are modeled as dynamic graphs with temporal properties. Most prior research focuses on discrete snapshots, neglecting the evolving nature of the enterprise network. Recently, some researchers have proposed continuous-time graph approaches. However, they often fail to fully incorporate security domain knowledge, and suffer from detection target bias. In this paper, given the limitations of the above approaches, we propose DygLM, a framework based on continuous-time graph representation learning with suspicious LM events as the detection target. It enables self-supervised learning in an attack-free manner, adapting to the evolving network environment. By applying a well-designed encoder to the continuous-time authentication events, we encode the node-associated neighbors, edges, and time intervals to obtain topological as well as temporal representation, which is further enhanced by learning long-term correlations using Transformer. We address the detection target bias by scoring the likelihood of malicious events in the evaluation phase. Through extensive experiments on two large datasets, we demonstrate the effectiveness of DygLM in detecting LM in both transductive and inductive settings.
Zhou Zhou 0007, Qingyun Liu 0001
CSCWD2
2025 MalTAG: Encrypted Malware Traffic Detection Framework via Graph Based Flow Interaction Mining
abstract
As encrypted malware campaigns become more sophisticated, significant challenges arise in effectively detecting malicious communications when relying solely on single-stream or single-feature network traffic analysis methods. Therefore, we propose MalTAG, a graph-based flow interaction mining framework to address existing challenges. MalTAG integrates network traffic into a multi-flow correlated graph and incorpo-rates various feature types rather than relying on a single stream or single type of feature. It effectively captures the contextual correlation properties of malicious activities in both temporal and attribute dimensions. Furthermore, MalTAG achieves graph representation learning through self-supervised contrastive learning without relying on prior label knowledge. This design helps to construct more stable representations of traffic behavior to adapt to the evolving nature of malicious activities. Ultimately, it utilizes various machine learning algorithms to detect malicious traffic comprehensively. Experimental evaluations demonstrate the feasibility and good performance of MalTAG in malware detection and family classification, outperforming existing methods.
Zhou Zhou 0007, Fengyuan Shi 0005, Qingyun Liu 0001
DSN2
2025 APTSniffer: Detecting APT Attack Traffic Using Retrieval-Augmented Large Language Models
abstract
Advanced Persistent Threats (APT) differ from traditional attacks by using more complex and covert strategies for long-term assaults, posing a severe threat to organizational and national security. Due to problems like the shortage of APT traffic data and encrypted traffic obfuscation, existing methods cannot accurately identify APT traffic with just a few traffic samples. To overcome the above limitation, we propose a novel encrypted APT traffic detection model, APTSniffer, which combines large language models (LLM) and retrieval-augmented technology. APTSniffer utilizes the few-shot inference and generalization abilities of large language models by converting raw traffic data into natural language inference examples understandable by the LLM. Experimental results show that, compared to other baseline models, APTSniffer exhibits SOTA performance. It achieves F1 scores above 97% on three APT datasets, making it practically applicable for APT traffic detection tasks.
Chengxiang Si, Zhou Zhou 0007, Chenxu Wang 0006, Peishuai Sun, Qingyun Liu 0001
ICASSP3
2025 PORTIA: A Multi-Granularity APT Detection Model Based on Provenance Graphs
abstract
Advanced Persistent Threats (APTs) have become a major cybersecurity threat due to their stealthy attack methods and long latency periods. Traditional signature-based detection struggles to detect novel attacks, and while unsupervised methods using Graph Neural Networks (GNNs) can model system behavior, they face challenges in handling large-scale provenance graphs and accurately incorporating system operational states for detection.This paper presents PORTIA, a multi-granularity APT detection model based on graph representation learning. PORTIA constructs provenance graphs by integrating temporal information from audit logs and uses a graph mask autoencoder to model normal system behavior, detecting anomalies through embedding shifts. As an unsupervised model, PORTIA can swiftly identify anomalous system states without relying on attack signatures, achieving fine-grained detection by incorporating system operational states. Evaluations on multiple datasets show that PORTIA detects both standard and APT attacks with high precision, outperforming existing detection systems.
Haoqiang Wang, Zhou Zhou 0007, Chengxiang Si, Qingyun Liu 0001
IJCNN5
2025 MOLE: Provenance Graph Generation Framework Based on LLM Prompting
abstract
In the increasingly complex landscape of cyber-attacks, logs have become a critical source of data for detecting system threats. Currently, most log-based detection systems rely on converting audit logs into provenance graphs during the process of attack investigation. However, this construction process is still heavily dependent on manually written code with regular expressions tailored to each specific log type. In this paper, we propose MOLE, a provenance graph generation framework based on prompting with large language models (LLMs). Unlike traditional approaches, MOLE does not rely on prior knowledge and is adaptable to diverse types of log data. The framework automatically generates provenance graph extraction templates through instruction generation and parses logs locally to produce the final provenance graph.MOLE leverages the log patterns and structures learned by LLMs from large-scale data during training. As a result, tasks that previously required several days of manual coding to generate a provenance graph can now be completed in just a few minutes. Furthermore, when processing 50 million log entries, the entire provenance graph generation process consumed only 20k tokens.
Haoqiang Wang, Zhou Zhou 0007, Chengxiang Si, Qingyun Liu 0001
IJCNN4
2025 DarkDC:Analysis of Scanning Behavior Subject Portrait Based on Deep Clustering
Xinyi Ji, Yixiang Zhao, Zhou Zhou 0007, Qingyun Liu 0001
ISCC3
2025 MV-TFNet: Malware Traffic Detection Based on Multi-View Flow Sequence and Time-Frequency Feature Fusion
abstract
The proliferation of malware’s command-andcontrol (C2) communications poses significant challenges to traditional network security defence as adversaries increasingly leverage advanced evasion tactics such as dynamic infrastructure rotation and traffic obfuscation. However, previous works suffer from inadequate temporal modeling, session-completion assumption, and inflexible data augmentation. To address these limitations, this paper proposes MV-TFNet, a novel C2 traffic detection framework that synergizes multi-view flow sequence analysis with temporal-frequency feature fusion. By decomposing host-server interactions into upstream, downstream, and bidirectional views, we employ a dual-path encoder to capture localized temporal dependencies via dilated convolutions and global frequencydomain patterns through spectral analysis. A contrastive learning paradigm, enhanced by label-guided adaptive data augmentation, further optimizes robustness against adversarial noise and distribution shifts. Moreover, MV-TFNet has a protocol-agnostic design, eliminating the dependency on plaintext information or protocol-specific parsing. Extensive evaluations on the MTA and DoHBrw datasets demonstrate MV-TFNet’s superiority over state-of-the-art methods, achieving an F1-score range from 0.94 to 0.99 and a false positive rate below 0.05.
Zhou Zhou 0007, Qingyun Liu 0001
ISCC2
2025 MTDIR: A Malicious Traffic Detection Model Based on the Image Retrieval Perspective
abstract
Traffic typically reflects network behavior, enabling the detection of network attacks through malicious traffic analysis. Existing methods often suffer from feature redundancy during extraction, reducing model efficiency. Additionally, these methods struggle to capture long-range dependencies, impacting feature learning. A single-perspective approach also limits the ability to learn universal patterns from multiple features, restricting the model's applicability. To address these shortcomings, this paper proposes a malicious traffic detection model based on image retrieval (MTDIR), which leverages depthwise separable convolution for multi-level feature retrieval, reducing computational overhead and redundancy. By learning features across multiple dimensions, MTDIR enhances detection performance. Ablation experiments confirm each module's effectiveness. MTDIR outperforms the control group in detection accuracy while maintaining low time overhead and considerable generalization and robustness, making it highly applicable.
Zhou Zhou 0007, Chengxiang Si
ICMR3
2025 Two Heads are Better than One: A Network Attack Detection Model Based on Multimodal and Multimedia Retrieval
abstract
Traffic, as a carrier of network behavior, can be used to detect attacks through malicious traffic detection. However, the existing methods have limited effectiveness in detecting covert attacks and are insufficiently resistant to interference in complex environments. In addition, the redundant information in the feature extraction process will reduce the model's efficiency. Therefore, this paper proposes a network attack detection model based on multimodal and multimedia retrieval (NADMR), which utilizes two modalities, namely, image and time series, to mine the key features of the traffic from a multi-dimensional perspective. Secondly, lightweight spatial attention and split channel attention are designed to extract discriminative features of image modality and analyze feature information of time series modality from multiple domains. Meanwhile, ghost convolution is introduced to improve efficiency. The ablation experiments verify the effectiveness of each technique, and the comparison experiments show that NADMR performs better in detection performance and efficiency. Its good generalization ability and robustness, as well as lower complexity, make it suitable for more scenarios.
Zhou Zhou 0007, Chengxiang Si
ICMR2
2024 P4-FILB: Stateless Load Balancing Mechanism in Firewall and IPv6 Environment with P4
abstract
Load balancers are critical infrastructure in modern distributed systems, and their main function is to evenly distribute client traffic to multiple servers for high availability and scalability. However, current load balancers face challenges in balanced resource utilization. To ensure per-connection consistency, load balancers typically assign subsequent requests from the same client to the same server. While this strategy simplifies session management and reduces the overhead of state synchronization, it also leads to uneven resource utilization.In this paper, we propose P4-FILB, a stateless load balancing scheme that achieves uniform load distribution among servers so that the resources of each server can be optimally balanced when receiving a large number of data streams. The key idea behind the implementation of P4-FILB is that a load balancer between the client and the server keeps track of the size of the packets that are currently being processed by the different servers. It also uses segment routing in IPv6, which uses an ordered list called "segments" to direct packets to servers that are currently utilizing a relatively small amount of resources. And this paper also proposes a two-tier load policy in firewall environment. According to the intelligent routing policy of the firewall, the packets are distributed to different groups of servers, and then the load balancer carries out further request allocation to improve the allocation of network resources. After evaluating the performance of P4-FILB, it is shown that the load balancing scheme performs better than the existing studies in terms of load balancing state between servers. It also reduces the overhead caused by extra packets by means of packets carrying connection information.
Haizhang Zhu, Zhou Zhou 0007, Chengwei Peng, Rong Yang 0008, Qingyun Liu 0001
IPCCC3
2024 LightRL-AD: A Lightweight Online Reinforcement Learning Approach for Autonomous Defense against Network Attacks
abstract
With the rapid growth of the Internet, network structure has become increasingly complex, leading to more diverse and impactful network attacks. Traditional methods of detecting and defending against network attacks struggle with increasingly complex situations due to human decision-making processes. Recent research has started exploring autonomous defense mechanisms for network attacks within software-defined network (SDN) environments. However, these methods typically employ complex reinforcement learning techniques, making them challenging to implement in online deployment environments. In this paper, we propose LightRL-AD, a lightweight online reinforcement learning approach for autonomous defense against network attacks in SDN. LightRL-AD integrates a machine learning-based Intrusion Detection System (IDS), a reinforcement learning-based Intrusion Prevention System (IPS), and a Moving Target Defense (MTD) mechanism. The ML-based IDS classifies network flows into categories such as malicious or benign, while the RL-based IPS utilizes the SARSA algorithm to determine and execute appropriate defensive actions, ensuring robust network security. We employ specific hardware and software to establish a simulated SDN network for our experiments. And we implement LightRL-AD in the network and evaluate its performance. Experimental results demonstrate that LightRL-AD performs better to defend against slow-rate DDoS attacks autonomously.
Fengyuan Shi 0005, Zhou Zhou 0007, Qingyun Liu 0001, Xiuguo Bao
TrustCom2
2024 Property graph representation learning for node classification
abstract
Abstract Graph representation learning (graph embedding) has led to breakthrough results in various machine learning graph-based applications such as node classification, link prediction and recommendation. Many real-world graphs can be characterized as the property graphs, because besides the structure information, there exists rich property information related to each node in the graphs. Many existing graph representation learning methods—e.g. random walk-based methods like and , focus only on the structure of graph for learning the node embedding. Although graph representation learning based on neural networks (e.g. typical methods such as ) uses the property of nodes as the initial features of nodes and then aggregates feature information of the neighbours, their limitation is that the neighbourhood of a node is considered to be uniform—i.e. there is no way to differentiate among neighbours of a node when learning a node embedding. Additionally, their definition of neighbourhood is local, i.e. only nodes connected to the current node are considered as neighbours. Hence, those methods fail to capture implicit/latent relationships among nodes, which are implicit in the given structure. In this study, our aim is to improve the performance of graph representation learning methods on property graphs. We present a new framework called ()—a graph representation learning framework to address above-mentioned limitations. Our proposed framework relies on the notion of latent neighbourhood, as well as systematic sampling of neighbouring nodes to obtain better representation of the nodes. The experimental results on five publicly available graph datasets demonstrate that outperforms state-of-the-art baselines for the task of node classification. We further evaluate the superiority of our proposed formulation by defining a novel quantitative metric to measure the usefulness of the sampled neighbourhood in the graph.
Nayyar Abbas Zaidi, Meijie Du, Zhou Zhou 0007, Gang Li 0009
Knowl. Inf. Syst.4
2023 AHIP: An Adaptive IP Hopping Method for Moving Target Defense to Thwart Network Attacks
abstract
In a static network, attackers can easily launch network attacks on target hosts which have long-term constant IP addresses. In order to defend against attackers effectively, many defense approaches use IP hopping to dynamically transform IP configuration. However, these approaches usually focus on one type of network attacks, scanning attacks or Denial of Service (DoS) attacks, and cannot sense network situations. This paper proposes AHIP, an adaptive IP hopping method for moving target defense (MTD) to defend against different network attacks. We use a trained lightweight one-dimensional convolutional neural network (1D-CNN) detector to judge whether there are no attacks, scanning attacks or DoS attacks in the network, which can adaptively trigger corresponding IP hopping strategy. We use specific hardware and software to create the software defined network (SDN) environment for experiments. The experiments prove that AHIP performs better to thwart network attacks and has lower system overhead.
Fengyuan Shi 0005, Zhou Zhou 0007, Qingyun Liu 0001, Xiuguo Bao
CSCWD2
2023 DarkFT: Automatic Scanning Behavior Analysis with FastText in Darknet Traffic
abstract
Network telescopes (Darknets) collect and record unsolicited Internet-wide traffic destined for a routed but unused address space, which provides a global perspective on Inter-net scanning behavior. However, it’s very difficult to extract meaningful information from Darknet traffic including a large number of unlabeled packets. In recent years, some work has used NLP techniques for self-supervised learning to generate embeddings as a rich representation of the Darknet. However, we found that previous resulting embeddings are not general enough. This paper proposed a new traffic representation model called Darknet Traffic Representations using FastText (DarkFT) which trains contextualized representation from large-scale unlabeled data. The embedding features can be applied to semi-supervised and unsupervised tasks. We conduct experiments on a dataset collected by network telescopes located in Japan and get better accuracy compared with state-of-the-art NLP-based approaches, especially for unknown senders.
Yixiang Zhao, Zhou Zhou 0007, Qingyun Liu 0001
CSCWD2
2023 Shrink: Identification of Encrypted Video Traffic Based on QUIC
abstract
With the increasing prevalence of network videos, video traffic has become a significant portion of overall network traffic. Due to the presence of harmful content such as pornography and violence in network videos, network monitoring is necessary. However, the encryption of videos poses challenges for network monitoring. More and more video service providers are adopting QUIC as the default video transmission protocol to accelerate data transfer speeds. However, the existing methods for identifying encrypted video traffic do not apply to QUIC. Video service providers typically employ Content Delivery Network (CDN) technology to enhance user experience, which can result in missing video chunks for side-channel identification. Additionally, fluctuations in network conditions can lead to the retransmission of video chunks. This paper proposes Shrink, a QUIC-based encrypted video traffic identification method. It effectively extracts video chunks from online QUIC encrypted video traffic and proposes a bucket structure and global-local match to alleviate the issues of video chunks retransmission and loss. Furthermore, a bucket word dictionary is designed to enhance the method’s running speed. Experimental results demonstrate that Shrink performs well in real network environments, exhibiting superior accuracy and speed compared to existing state-of-the-art methods.
Weitao Tang, Meijie Du, Zhao Li 0010, Zhou Zhou 0007, Qingyun Liu 0001
IPCCC5
2023 GoGDDoS: A Multi-Classifier for DDoS Attacks Using Graph Neural Networks
abstract
Distributed Denial of Service (DDoS) attacks are rising, evolving and growing sophistication. Multi-vector which leverages more than one methods is prevalent recently. To cope with multi-vector DDoS attack, it is necessary to classify DDoS attacks for taking robust measures. However, existing ML-based approaches for DDoS traffic multi-classification barely leverage relationships between packets and flows, which are crucial information that can significantly improve multi-classification performance. This paper proposes GoGDDoS, a multi-classifier for DDoS attacks. Concretely, we construct GoG traffic graph to clearly compress relationships between packets and flows. It merges relationship graphs of packets and flows by using graph of graph. Then, we build a two-level Graph Neural Network model to mine potential attack patterns from GoG traffic graph. The experiments with well-known datasets show that GoGDDoS performs better than its counterparts.
Zhou Zhou 0007, Fengyuan Shi 0005, Qingyun Liu 0001
ISCC2
2023 Hunting for Hidden RDP-MITM: Analyzing and Detecting RDP MITM Tools Based on Network Features
abstract
Remote Desktop Protocol (RDP) is commonly used for remote access to windows computers. As more and more people work remotely, the number of users of RDP is increasing, making RDP a growing concern in cybersecurity. The latest way to threaten RDP security is RDP man-in-the-middle (MITM) tools which realize the MITM function in an RDP connection and automate the MITM attack process, significantly reducing the difficulty of network attacks. At the same time, RDP MITM tools can be used for high-interaction RDP honeypots. In order to mitigate this risk, we present the first in-depth study of RDP MITM tools in this paper. By analysis and experiment, we identify network features that can be used to detect RDP MITM tools effectively. Based on packet latency and TLS handshake, we propose a machine learning classifier that can detect RDP MITM tools for securing RDP connections. Finally, we analyze the deployment of RDP MITM tools in the wild and effectively measure the RDP MITM tools using our proposed detection approach.
Zhou Zhou 0007, Fengyuan Shi 0005, Qingyun Liu 0001
ISCC2
2022 GraphDDoS: Effective DDoS Attack Detection Using Graph Neural Networks
abstract
Distributed Denial of Service (DDoS) attacks have occurred frequently in recent years, causing massive damage. It is critical to detect DDoS attacks fast and accurately. Previous Deep Learning (DL) methods for detecting DDoS attacks barely leverage the relationships between packets and between flows in traffic, which are crucial information that can significantly improve detection performance. This paper proposes GraphDDoS, a GNN-based approach for detecting DDoS attacks using endpoint traffic graphs. Concretely, we convert traffic into endpoint traffic graphs, containing information of packets’ relationships (structure of a single flow) and flows’ relationships (burst information and periodic information of multiple flows). Then, converted endpoint traffic graphs are sent to the GNN classifier to learn DDoS attack patterns accurately. The experiments with well-known datasets show that GraphDDoS outperforms the state-of-the-art DL-based approaches. The effectiveness is mainly introduced by the capability of GraphDDoS to learn patterns of attacks structured as graphs.
Zhou Zhou 0007, Meijie Du, Qingyun Liu 0001
CSCWD3
2022 P4-NSAF: defending IPv6 networks against ICMPv6 DoS and DDoS attacks with P4
abstract
Internet Protocol Version 6 (IPv6) is expected for widespread deployment worldwide. Such rapid development of IPv6 may lead to safety problems. The main threats in IPv6 networks are denial of service (DoS) attacks and distributed DoS (DDoS) attacks. In addition to the similar threats in Internet Protocol Version 4 (IPv4), IPv6 has introduced new potential vulnerabilities, which are DoS and DDoS attacks based on Internet Control Message Protocol version 6 (ICMPv6). We divide such new attacks into two categories: pure flooding attacks and source address spoofing attacks. We propose P4-NSAF, a scheme to defend against the above two IPv6 DoS and DDoS attacks in the programmable data plane. P4-NSAF uses Count-Min Sketch to defend against flooding attacks and records information about IPv6 agents into match tables to prevent source address spoofing attacks. We implement a prototype of P4-NSAF with P4 and evaluate it in the programmable data plane. The result suggests that P4-NSAF can effectively protect IPv6 networks from DoS and DDoS attacks based on ICMPv6.
Zhou Zhou 0007, Qingyun Liu 0001, Zhao Li 0010
ICC3
2021 Attributed Heterogeneous Graph Neural Network for Malicious Domain Detection
abstract
Malicious activities on the Internet are one of the most dangerous threats to users and organizations. Because of the flexibility a nd accessibility of domains, cyber criminals often utilize them to launch cyber attacks such as phishing or malware. Most of traditional malicious domain detection methods rely on feature engineering to learn the patterns of malicious domains. However, these methods can be easily evaded by some sophisticated evasion techniques such as Domain-flux or Fast-flux. Some recent studies utilized graph-based models to infer malicious domains and achieved better performance, yet without a fine-grained modeling of DNS scenarios. In this paper, we propose an attributed heterogeneous graph neural network model, GAMD, to detect malicious domains in a semi-supervised learning paradigm. Concretely, we utilize attributed heterogeneous information network to model the DNS scenarios with different types of nodes including domain, host, resolved-IP and different types of relation, such as request and resolution relations. We then design a fine-grained node type-aware feature transformation and edge type-aware aggregation mechanism to fuse the node attributes and structure information simultaneously and complete the inference over DNS graphs. In the experiments, we evaluate the performance of our model on a large-scale realworld passive DNS data and show that the proposed method outperforms the state-of-the-art in most evaluation metrics.
Shuai Zhang 0007, Zhou Zhou 0007, Da Li 0002, Youbing Zhong, Qingyun Liu 0001
CSCWD2
2020 BPA: The Optimal Placement of Interdependent VNFs in Many-Core System
Youbing Zhong, Zhou Zhou 0007, Xuan Liu 0006, Da Li 0002, Meijun Guo, Shuai Zhang 0007, Qingyun Liu 0001, Li Guo 0001
CollaborateCom (2)2
2019 NTS: A Scalable Virtual Testbed Architecture with Dynamic Scheduling and Backpressure
Youbing Zhong, Zhou Zhou 0007, Da Li 0002, Wenliang He, Chao Zheng 0001, Qingyun Liu 0001, Li Guo 0001
CollaborateCom2
2019 Hunting for Invisible SmartCam: Characterizing and Detecting Smart Camera Based on Netflow Analysis
abstract
Nowadays, the rapid growth of cloud computing and IoT enabled services among multiple organizations brings both promising prospects and security & privacy challenges. IP cameras have become a top target for hackers because of their relatively high computing power and throughput. To understand the risks of these threats requires learning about IP cameras-where are they, how many are there? Active scanning is considered to be an effective way, like SHODAN. However, deployment of smart cameras in the network address translation (NAT) environments with dynamic locations is usually desired. To find these Invisible Cameras, CamHunter: (i) introduces three statements of smart cameras when they are online, (ii) concludes the most popular smart cameras in China have very similar communication patterns, (iii) proposes a model to detect smart cameras in a passive way constructed by nineteen feature sets, and (iv) raises alarms for IoT manufacturers. Our real-world experiments demonstrate the effectiveness of CamHunter in finding smart cameras even if they are behind NATs and using encrypted connections like SSL/TLS or private protocols. We argue that CamHunter represents an important view of IoT security and privacy, and it can guide the effort of designing and protecting smart cameras.
Baiyang Li, Yujia Zhu, Qingyun Liu 0001, Zhou Zhou 0007, Li Guo 0001
ICC4
2018 SASD: A Self-Adaptive Stateful Decompression Architecture
abstract
Due to the increasing threats in the current network environment, many researchers have shifted their interests to network content audit, which combines deep packet inspection and natural language processing. However, the performance of network content audit systems is becoming the bottle-neck because of the demand on processing fast growing compressed traffic. While compressed traffic is often split into multiple out-of-order packets for transmission, stateful decompression ensures that the compressed data are processed in a timely manner without waiting for all the compressed traffic to arrive before decompressing. In the meanwhile, hardware innovations lead to new type of devices being invented, which shows promise to fully handle the offloaded traffic for complex calculations at higher throughput than software-based solutions. We consider both software-based and hardware-based solutions for decompressing traffic from network content audit systems and study the workload. We notice that the performance is data-dependant: hardware-based decompression solutions perform better for longer compressed data than software method. On the contrary, software-based decompressing methods are more preferred for the short content in terms of the processing speed. So there is no one-size-fits-all solution. In this paper, we combine the advantages of hardware and software and propose a novel self-adaptive stateful decompression architecture to support fast decompression in accordance with the traffic status and system state. Experiments on real-world traffic show that our proposed architecture can achieve about three times of the data decompression efficiency, compared to the best pure software and hardware algorithm, which can significantly improve the detection efficiency of many network content audit systems.
Zhou Zhou 0007, Qingyun Liu 0001, Yujia Zhu, Da Li 0002, Li Guo 0001
GLOBECOM1
2018 Janus: A User-Level TCP Stack for Processing 40 Million Concurrent TCP Connections
abstract
C10M is an Internet scalability problem regarding how to handle 10 million simultaneous TCP connections on a web server. Although kernel- and user-level approaches have been proposed to increase TCP stack scalability on multicore systems, C10M is still an open problem. In this paper we present Janus, a high-performance user-level TCP stack that focuses on serving massive TCP connections. In addition to adopting well- known techniques, our design (1) separates packet I/O cores from TCP processing cores to achieve high scalability and flexibility on a multicore system and (2) lets each application run as a per- connection coroutine together with a packet processing loop, which greatly improves cache affinity and saves memory. We demonstrate that Janus can accept 1.86 million new connections per second while maintaining 40 million concurrent connections and significantly outperforms Linux and state-of-the-art user-space network stacks in both throughput and connection concurrency.
Chao Zheng 0001, Qiuwen Lu, Zhou Zhou 0007, Qinyun Liu
ICC5
2015 Limited Dictionary Builder: An approach to select representative tokens for malicious URLs detection
abstract
Cybercriminals use Malicious Uniform Resource Locators (URLs) as the entry to implement a variety of web attacks, such as phishing, spamming, and malware distribution, which may lead to huge finance and data loss. Thus, malicious URLs should be detected as accurately and quickly as possible. Heuristic-based detection approaches are one of the most popular methods to achieve the above goals. The detection results come from the usage of many heuristic features in this approach. However, tremendous new pages and meaningless tokens lead to the explosion of feature sets, and exhaust memory space finally. In this paper, we try to address the problem by selecting some representative members from the initial feature set, which should have the best predictive ability among the same number of selected features. For each feature, we give an evaluation method of O(1) complexity to measure its predictive ability. Then we make the selection based on all the measured values with linear complexity. Experimental results show that our approach can achieve almost the same false negative rate using only 8.3% features for malicious URLs detection, comparing with prior approaches. Moreover, our approach may work efficiently in the big data era, as it can handle 20 thousand URLs per second in our experiments on average.
Hongzhou Sha, Zhou Zhou 0007, Qingyun Liu 0001, Tingwen Liu, Chao Zheng 0001
ICC2
2014 GuidedTracker: Track the victims with access logs to finding malicious web pages
abstract
Malicious web pages have become a malignant tumour for the Internet, which spread malicious code, steal people's private information, and deliver spamming advertisements. And how to distinguish them from the huge number of normal web pages effectively remains a huge challenge in the era of big data. To detect malicious pages, one needs to first collect candidate web pages that are live on the web; then filter massive legitimate pages using fast filters and finally examine the remaining pages using precisely but slow analyzer. However, there are new challenges recently for these conventional techniques, including large scale, imbalance data and the usage of cloaking techniques. To cope with these challenges, the malicious URL detection system should perform more efficiently. In this paper, we propose a system, named GuidedTracker, to search for suspicious malicious pages. GuidedTracker starts from the seed set which includes known malicious pages. Then, it automatically figures out those victims based on the seed set and the visit relation database. Finally, the access records of these victims are used to identify other malicious pages. In this way, GuidedTracker increase the percentages of malicious URLs in the input URL stream submitted to the precisely analyzer. To our best knowledge, GuidedTracker is the first to introduce visit relations to tackle the malicious URL detection problem. The introduction of visit relations limits the scope of URL inspection and enables this approach to have the ability of self-learning. Experimental results show that the overall "toxicity" can be improved by 6.97%-50.38% compared with full inspection of access logs.
Hongzhou Sha, Qingyun Liu 0001, Zhou Zhou 0007, Chao Zheng 0001
GLOBECOM3
2013 File-aware P2P traffic classification: An aid to network management
Zhou Zhou 0007
Peer-to-Peer Netw. Appl.2
2011 File-Aware P2P Traffic Classification and Management
abstract
As P2P dominates Internet traffic in recent years, ISPs are striving to balance between providing the basic networking service for P2P users and properly managing network bandwidth usage. However, current P2P traffic management strategies are unable to satisfy both requirements. A file-aware P2P traffic classification method is presented in this paper. It can identify a file and the associated flows. Based on the file-level concurrent flow information, ISPs can adopt more efficient and flexible strategies to manage P2P traffic. We offered two alternatives: limiting the aggregated bandwidth consumption or the number of concurrent flows that a peer can use to download a particular file.
Zhou Zhou 0007
ANCS1
2010 A High-Performance URL Lookup Engine for URL Filtering Systems
abstract
URL filtering systems provide a simple and effective way to prevent people from browsing undesirable or malicious Websites. These systems require a well-designed URL lookup method as the core operation. A high-performance URL lookup engine is proposed in this paper for URL filtering systems. It combines a URL compression algorithm with a multiple string matching based (Wu-Manber-like) matching algorithm. Using this method, the proposed URL lookup engine can achieve high URL lookup performance and efficient memory utilization for storing the ever-increasing URL blacklist with the ability of prefix matching. Experiments with actual URL blacklists and requested URL sets show that our engine can save about 80% memory usage for storing URL blacklists, and reduce 58%-162% URL lookup time compared with the state-of-the-art URL lookup methods.
Zhou Zhou 0007, Yunde Jia
ICC1