Chao Zheng 0001

dblp:54/1147-1 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
8since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 9 · 1 first-author · 5 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021
YearPublicationVenuePosition
2024 Seeing the Attack Paths: Improved Flow Correlation Scheme in Stepping-Stone Intrusion
abstract
Stepping-stones are widely used by attackers to conceal their identities and gain unauthorized access to restricted targets. Numerous strategies have been suggested to identify stepping-stones and counteract evasive behaviors, with flow correlation standing out as the key technique used in many deanonymization methods. Existing attempts to tackle the flow correlation issue rely on long-flow observations and feature-based methods. Yet, these approaches meet substantial limitations in their suitability across diverse scenarios, network noises influence and accuracy, particularly in the stepping-stone environment. In this paper, we introduce an improved flow correlation model, FlowLinker, which aims to resolve these issues. Specifically, we combined multidimensional statistical features and cumulative flow sequences to construct a robust traffic feature. Then, we leveraged the triplet network to produce an optimized representation, amplifying the difference between unrelated representations. Consequently, it significantly reduces the false correlation rate with flows that seem similar but are unrelated. The experiments on real-world datasets from different network environments we collected show that FlowLinker outperforms other state-of-the-art methods with shorter lengths of flow observations.
Chao Zheng 0001, Zhao Li 0010, Jinqiao Shi
CSCWD2
2024 Scrutinizing Code Signing: A Study of in-Depth Threat Modeling and Defense Mechanism
abstract
Abuse of code signing has garnered attention from security researchers, as evidenced by threat modeling efforts targeting the public key infrastructure trust infrastructure of code signing and empirical studies examining issues surrounding the revocation of code signing certificates. However, current research on code signing remains inadequate in bridging the gap between attack strategies and defensive measures. This shortfall is primarily due to the predominant focus on quantitative measurements in academic studies, often at the expense of a thorough analysis of the underlying code signing mechanisms. Moreover, the misalignment of some threat models and measurement outcomes with real-world attack scenarios further hampers efforts to enhance defenses against code signing abuse. To the best of our knowledge, this article represents the first comprehensive and in-depth analysis of code signing security from both offensive and defensive perspectives. Commencing with a profound understanding of code signing and its verification mechanisms, we constructed an integrated threat model encompassing eight typical attack patterns and distilled a set of critical security properties that directly influence the security of code signing. Proceeding from this foundation, we systematically reviewed and analyzed various defensive strategies associated with these security properties, meticulously discussing their strengths, limitations, and the specific attack types they effectively defend against or mitigate. Lastly, this article conducts a risk statistical analysis based on actual security incidents, with the related results directly impacting the prioritization of defensive mechanism deployments. This ensures, at a practical level, the effectiveness and relevance of the defense strategies implemented.
Tiantian Ji, Binxing Fang, Xiang Cui, Fan Gu, Chao Zheng 0001
IEEE Internet Things J.7
2024 A Semantics-Based Approach on Binary Function Similarity Detection
abstract
As a fundamental component of Internet of Things (IoT) devices, firmware plays an essential role. Nowadays, the development of IoT firmware relies extensively on third-party components and substantially enhances development efficiency. However, these components are not inherently secure, and their vulnerabilities can adversely affect the security of IoT firmware. Existing research adopts binary code similarity analysis to detect known vulnerabilities in firmware. However, it encounters significant challenges, primarily in extracting function features from the limited semantic information within binary code. Another challenge is the need for real-world datasets to assess the model’s performance in practical scenarios, such as firmware supply chain analysis. We present a detection model named PDG2VEC based on Program Dependence Graphs (PDGs) to tackle these challenges. PDG2VEC extracts function features at the variable level on PDG and assesses function similarity by evaluating whether two functions can represent each other. We conducted evaluations using three datasets, including one we created to simulate a firmware supply chain scenario. The experimental results demonstrate that PDG2VEC exhibits resilience to cross-architecture challenges and captures more precise semantics than other approaches. Furthermore, PDG2VEC outperforms state-of-the-art tools in the supply chain analysis scenario, with a 16% higher AUC value average against baseline approaches.
Binxing Fang, Zehui Xiong, Yuwei Liu 0001, Chao Zheng 0001, Qinnan Zhang
IEEE Internet Things J.6
2023 Detecting Fake-Normal Pornographic and Gambling Websites through one Multi-Attention HGNN
abstract
The rapid development of pornographic and gambling websites, fueled by the widespread abuse of information technology, has become a growing concern. They pose a serious threat to the physical and mental health of children and can also endanger personal property. Therefore, it is necessary to detect them. However, pornographic and gambling websites become more and more tricky, which shows fake-normal to evade censorship and challenges traditional content-based detection methods. Therefore, it is essential to rely on information about relationships between websites.We propose HMAN, one Multi-Attention Heterogeneous Graph Neural Network (HGNN) model to detect pornographic and gambling websites by integrating content features and structural information, even if they present fake-normal. By one multi-attention mechanism consisting of explicit weight, self-attention and attention mechanism, content features can be selectively utilized with the assistance of structural information. The experimental results show that our method achieves the best 95.1% Macro-Avg-F1 and outperforms all baselines. We also illustrate that all extracted metapaths do contribute to the detection, where the hyperlink, title/meta terms and IP address are relatively important.
Xiaoqing Ma, Chao Zheng 0001, Zhao Li 0010, Jiangyi Yin, Qingyun Liu 0001, Xunxun Chen
CSCWD2
2022 Accelerate State Sharing of Network Function with RDMA
abstract
State sharing can help network functions (NFs) provide support for clustered deployment and minimize the jitter caused by elastic scaling, one of the central futures of NFV. But the network overhead of remote access and the dynamically changing workload restrict the performance of the state sharing framework. The state-of-the-art state sharing frameworks mainly use the affinity strategy to migrate the target state from another node to itself. This strategy is unsuitable for asymmetric routing scenario because multiple nodes will access the same state. This paper presents RedKV, a fast and flexible Remote Direct Memory Access (RDMA) Enhanced Distributed Key-Value store that supports low-latency state sharing both in symmetric and asymmetric routing scenarios. RedKV has two unconventional designs: First, it scatters states over cluster nodes without affinity strategy and uses a unified addressing (UA) space to hide the complexity of state access. Second, it accelerates the operation with RDMA and optimizes the process according to the characteristics of RDMA and NF to mitigate the cost of remote access. Our evaluation shows that RedKV can help network functions share states with an added latency overhead of 2.26µs, outperforming state-of-art solutions by 4.6 times. The pause time caused by the scaling event has also been reduced by 81.5%.
Chenming Chang, Chao Zheng 0001, Liang Zhan, Qingyun Liu 0001
GLOBECOM3
2022 A Lightweight Graph-based Method to Detect Pornographic and Gambling Websites with Imperfect Datasets
abstract
With the widespread abuse of information technology, pornographic and gambling websites develop rapidly. They affect the physical and mental health of children and endanger personal property. Therefore, it is necessary to detect them. However, the existing detection methods ignored that imperfect datasets are common in the scenario of pornographic and gambling websites which are hence adverse to the detection. Those imperfections specifically include sparse samples, mismatch and imbalanced datasets. In addition, over-reliance on visual features incurred high overhead.To overcome these shortcomings, we innovatively propose a lightweight graph-based method to detect pornographic and gambling websites through semi-supervised learning of textual content. The semi-supervised learning is to solve sparse samples and mismatch datasets, while the graph-based approach can combine the semi-supervised part with community discovery to deal with imbalanced datasets. Specifically, we perform the detection process with the utilization of modified TF-IDF and Louvain during the iteration and updating by the EM algorithm. The experimental results show that our method achieves the best 92.01% Macro-Avg-F1 with the shortest CPU time and outperforms all baselines. We also illustrate that the designed components in our model do contribute to the detection.
Xiaoqing Ma, Chao Zheng 0001, Zhao Li 0010, Jiangyi Yin, Qingyun Liu 0001, Xunxun Chen
TrustCom2
2021 A P4-Based Packet Scheduling Approach for Clustered Deep Packet Inspection Appliances
abstract
Packet scheduling approach enables clustered deep packet inspection (DPI) appliances to achieve stateful packet processing through efficient cooperation. In this paper, we propose a novel methodology of packet scheduling, called P4CLUS, for clustering DPI appliances efficiently. P4CLUS eliminates unnecessary packet transmission between peer nodes via the idea of packet path prediction. It performs sketch-based measurement on traffic handling nodes to record their flow distributions and then compresses the results to a compact packet forwarding table. Our approach guarantees per-flow consistency with lower bandwidth overhead and is effective in the practice of eliminating asymmetric traffic in stateful DPI. We construct the P4CLUS prototype with P4 on Protocol-Independent Switch Architecture (PISA). Our evaluation shows that P4CLUS can reduce the bandwidth overhead by up to 73.75% than the hash-based method, and only requires 2MB SRAM and 80KB TCAM resources at most.
Qingyun Liu 0001, Chao Zheng 0001
ICCCN4
2021 CDNFinder: Detecting CDN-hosted Nodes by Graph-Based Semi-Supervised Classification
abstract
As a crucial internet infrastructure, Content Delivery Network (CDN) is widely deployed. Detecting CDN-hosted nodes from network traffic is important for Quality of Service (QoS), malware detection and firewall rule-sets. Current researches use hand-crafted rules, classification or clustering methods. However, those methods relying on plaintext are limited by the invisibility of plaintext due to encryption, as well as the limitations of DNS Resource Records, such as unreliability. Besides, those methods don't dig the structural information of domains and IPs. To overcome those shortcomings, we present CDNFinder, a novel method to detect CDN-hosted nodes by graph-based semi-supervised classification. Based on the active datasets collected in 10 vantage points, we construct the graph and extract innovative attributes. By modifying Graph Neural Network (GNN), CDNFinder outperforms classical machine learning methods, especially in recall rate (around 98%). Meanwhile, CDNFinder shortens the runtime of classical GNN algorithm by about 31% with no loss in metrics.
Xiaoqing Ma, Chao Zheng 0001, Zhao Li 0010, Qingyun Liu 0001, Xunxun Chen
ISCC2
2020 MAAN: A Multiple Attribute Association Network for Mobile Encrypted Traffic Classification
Fengzhao Shi, Chao Zheng 0001, Qingyun Liu 0001
SecureComm (1)2
2019 NTS: A Scalable Virtual Testbed Architecture with Dynamic Scheduling and Backpressure
Youbing Zhong, Zhou Zhou 0007, Da Li 0002, Wenliang He, Chao Zheng 0001, Qingyun Liu 0001, Li Guo 0001
CollaborateCom5
2018 Janus: A User-Level TCP Stack for Processing 40 Million Concurrent TCP Connections
abstract
C10M is an Internet scalability problem regarding how to handle 10 million simultaneous TCP connections on a web server. Although kernel- and user-level approaches have been proposed to increase TCP stack scalability on multicore systems, C10M is still an open problem. In this paper we present Janus, a high-performance user-level TCP stack that focuses on serving massive TCP connections. In addition to adopting well- known techniques, our design (1) separates packet I/O cores from TCP processing cores to achieve high scalability and flexibility on a multicore system and (2) lets each application run as a per- connection coroutine together with a packet processing loop, which greatly improves cache affinity and saves memory. We demonstrate that Janus can accept 1.86 million new connections per second while maintaining 40 million concurrent connections and significantly outperforms Linux and state-of-the-art user-space network stacks in both throughput and connection concurrency.
Chao Zheng 0001, Qiuwen Lu, Zhou Zhou 0007, Qinyun Liu
ICC1
2018 Hashing Incomplete and Unordered Network Streams
Chao Zheng 0001, Qingyun Liu 0001, Binxing Fang
IFIP Int. Conf. Digital Forensics1
2016 CookieMiner: Towards real-time reconstruction of web-downloading chains from network traces
abstract
Network traces are one of the most exhaustive data sources for the forensic investigation of computer security incidents. Recent advances in capturing the network traces techniques have facilitated the forensic processing, including the reconstruction. Unfortunately, off-line Web-downloading chain reconstruction could not meet the demands for real time processing. Furthermore, the packets in prior studies are mainly captured on the client or server side, which is difficult to monitor the network traffic in the real-time. Consequently, how to online reconstruct from the packets captured by the gateway is getting challenging. In this paper, based on the packets from the gateway, we propose a novel system, CookieMiner. CookieMiner first identifies the web-downloading resources, and then use the cookies to reconstruct the web-downloading chains reversely. All cookies would be firstly split into a series of tokens by semicolon and a threshold value will be set by the SET_K algorithm according to the number of tokens. Next, the HTTP packets whose tokens' number is greater than the threshold will be sorted by timestamp and then inserted into the corresponding web-downloading chain. Finally, the most frequent URLs are extracted as entry points from all the chains with the same web-downloading resources. In addition, by a user study involving 6 pair-wise downloading applications, CookieMiner can reconstruct the web-downloading chains and find their entry points with high precision and low false positive rate.
Peng Zhang 0001, Chao Zheng 0001, Qingyun Liu 0001
ICC3
2015 Limited Dictionary Builder: An approach to select representative tokens for malicious URLs detection
abstract
Cybercriminals use Malicious Uniform Resource Locators (URLs) as the entry to implement a variety of web attacks, such as phishing, spamming, and malware distribution, which may lead to huge finance and data loss. Thus, malicious URLs should be detected as accurately and quickly as possible. Heuristic-based detection approaches are one of the most popular methods to achieve the above goals. The detection results come from the usage of many heuristic features in this approach. However, tremendous new pages and meaningless tokens lead to the explosion of feature sets, and exhaust memory space finally. In this paper, we try to address the problem by selecting some representative members from the initial feature set, which should have the best predictive ability among the same number of selected features. For each feature, we give an evaluation method of O(1) complexity to measure its predictive ability. Then we make the selection based on all the measured values with linear complexity. Experimental results show that our approach can achieve almost the same false negative rate using only 8.3% features for malicious URLs detection, comparing with prior approaches. Moreover, our approach may work efficiently in the big data era, as it can handle 20 thousand URLs per second in our experiments on average.
Hongzhou Sha, Zhou Zhou 0007, Qingyun Liu 0001, Tingwen Liu, Chao Zheng 0001
ICC5
2014 GuidedTracker: Track the victims with access logs to finding malicious web pages
abstract
Malicious web pages have become a malignant tumour for the Internet, which spread malicious code, steal people's private information, and deliver spamming advertisements. And how to distinguish them from the huge number of normal web pages effectively remains a huge challenge in the era of big data. To detect malicious pages, one needs to first collect candidate web pages that are live on the web; then filter massive legitimate pages using fast filters and finally examine the remaining pages using precisely but slow analyzer. However, there are new challenges recently for these conventional techniques, including large scale, imbalance data and the usage of cloaking techniques. To cope with these challenges, the malicious URL detection system should perform more efficiently. In this paper, we propose a system, named GuidedTracker, to search for suspicious malicious pages. GuidedTracker starts from the seed set which includes known malicious pages. Then, it automatically figures out those victims based on the seed set and the visit relation database. Finally, the access records of these victims are used to identify other malicious pages. In this way, GuidedTracker increase the percentages of malicious URLs in the input URL stream submitted to the precisely analyzer. To our best knowledge, GuidedTracker is the first to introduce visit relations to tackle the malicious URL detection problem. The introduction of visit relations limits the scope of URL inspection and enables this approach to have the ability of self-learning. Experimental results show that the overall "toxicity" can be improved by 6.97%-50.38% compared with full inspection of access logs.
Hongzhou Sha, Qingyun Liu 0001, Zhou Zhou 0007, Chao Zheng 0001
GLOBECOM4