Fangliang Xu

dblp:132/1722 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-4099-560XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-authorComputer networks · 3Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SVRM: Composing Various Network Service Fuzzing Corpus with One Single Model
abstract
Discovering vulnerabilities in network service is of great significance. Currently, coverage-guided fuzzing (CGF) is widely regarded as the most effective method. However, the efficiency of CGF depends on the quality of initial corpus. The initial corpus is a set of valid input examples used to initiate the fuzzing process. Constructing high-quality initial corpus typically requires manual efforts to understand the implementation details and corresponding protocol specifications, making it difficult to generalize across different protocol implementations.To generate high-quality corpus tailored to service under test (SUT), this paper proposes a protocol-independent smart generation method. The paper introduces a novel service communication model and utilizes active learning algorithms to automatically construct the model. By analyzing the minimum spanning tree of the model, we achieve automatic generation of high-quality corpus that adapts to the SUT.We conduct experiment by generating adaptive corpus for 6 targets of 6 different protocols in ProFuzzBench. Compared to the corpus provided by ProFuzzBench, the corpus generated by our system improve the state coverage of modern protocol fuzzers by 37.1% and discover known real protocol vulnerabilities at a speed 2.47x faster.
Wenfeng Lin, Zhiyuan Jiang, Fangliang Xu, Yunfei Su, Lingchu Mao, Chaojing Tang
ICASSP3
2025 RPFUZZ: Efficient network service fuzzing via pruning redundant mutation
abstract
Coverage-guided fuzzing (CGF) has proven its outstanding performance on vulnerability detection. However, existing approaches exhibit limitations when handling network service. Restricted by network I/O duration and chronology, long packet sequences crafted by fuzzers incur a substantial execution cost. Test cases with such non-coverage-improving mutations (i.e. redundant mutation) can significantly reduce fuzzing throughput and compromise vulnerability discovery. To address this issue, we propose RPFUZZ, a novel network fuzzing framework designed to systematically reduce redundant mutations: (1) We propose redundant mutation pruning for network service fuzzing. By early terminating redundant mutations’ execution, RPFUZZ can achieve higher throughput. (2) To detect redundant mutation, we propose redundant mutation oracle. This oracle dynamically judges whether a test case is redundant according to current code coverage and value of service-related variables (SRVs). (3)To identify SRVs, we propose an integrated approach combining dynamic call stack analysis with static value-flow graph (VFG) analysis. To evaluate the performance of RPFUZZ, we implement a prototype on top of NYX-NET. We conduct thorough experiments on ProFuzzBench, a benchmark that consists of 12 real-world network services. The results indicate that RPFUZZ achieves over 185% improvement in throughput and 1.02% rise in code coverage compared with NYX-NET. Besides, RPFUZZ has successfully uncovered 1753 unique crashes across 6 network services, including an unreported vulnerability (assigned to CVE-2024-57392) in ProFTPD, which has been well tested. • We propose redundant mutation pruning technique for network service fuzzing. By pruning mutated suffix packet sequence which is non-coverage-improving, network service fuzzer can achieve higher throughout. This is achieved by redundant mutation oracle, which leverage code coverage and identified service-related variables’ (SRVs) value to decide whether continuing current execution is advisable. • To precisely identify SRVs in network services, we propose an identification method combining with call stack analysis and value-flow graph analysis. This method is based on SRV’s programming features, which can be applied in various network service. • Based on technique above, We implement RPFUZZ. RPFUZZ achieved more than 185.92% (average 56.39%) throughput enhancement over NYX-NET, while improving maximum 4.27% code coverage (average +1.02%). It successfully identified 1753 unique crashes across 6 targets without ASAN and a buffer overflow vulnerability in ProFTPD (assigned CVE-2024-57392).
Wenfeng Lin, Fangliang Xu, Zhiyuan Jiang, Chaojing Tang
Comput. Secur.2
2024 S2Vul: Vulnerability Analysis Based on Self-supervised Information Integration
abstract
The analysis of static vulnerabilities, which consists of detection, classification, and localization, is a perpetually significant concern in software security. The advancement of neural networks has led to a greater emphasis on vulnerability detection research. However, most research faced obstacles in attempting to perform satisfactorily on real-world datasets. Furthermore, an additional obstacle is the substantial reliance of the studies on labels, which requires considerable effort for labeling, restricts the model’s scalability, and potentially results in adverse effects due to inaccurate labels.This study introduces an approach for efficiently training the model to adjust to various vulnerability analysis tasks. By synthesizing a wide range of information extracted from the source code, the model allows the network to collect substantial data to identify vulnerabilities. In addition, unsupervised representation learning is employed to mitigate the influence of labels in the method. The model could be applied to a wide range of vulnerability-related tasks by training task-specific classifiers at a minimal cost. The model was evaluated using two public vulnerability datasets. The test results validate the efficacy of our model, which achieves an accuracy of more than 70% in localization, an accuracy of more than 98% in detection, and a greater accuracy above 70% in classification when evaluating the real-world dataset, surpassing the performance of peer approaches.
Fangliang Xu
ISSRE3
2022 An Android Malware Detection and Classification Approach Based on Contrastive Lerning
Fangliang Xu, Mantun Chen
Comput. Secur.4
2020 Reducing network cost of data repair in erasure-coded cross-datacenter storage
Han Bao 0006, Yijie Wang 0001, Fangliang Xu
Future Gener. Comput. Syst.3
2020 An Adaptive Erasure Code for JointCloud Storage of Internet of Things Big Data
abstract
JointCloud is a cross-cloud cooperation architecture for integrated Internet service customization. The customized cross-cloud storage service based on this architecture is called JointCloud storage. Storing the Internet of Things (IoT) big data in erasure-coded JointCloud storage systems ensures that data can be accessed when several cloud services interrupt. However, because existing erasure codes cannot adapt the generator matrix and data placement scheme to different network environments and encoding parameters, they usually incur a large network resource consumption (NRC) for repairing data in JointCloud storage systems. As a result, the availability of IoT applications running on JointCloud storage systems is impaired. In this article, to minimize the NRC of repairing data, we propose an adaptive erasure code for JointCloud storage of IoT big data called ACIoT. Specifically, we first propose the concept of average weighted locality (AWL) of a stripe of erasure-coded data, which is proportional to the average NRC of repairing this stripe in JointCloud storage systems. Then, we propose an active parallel trial-and-error algorithm to calculate the optimal generator matrix and data placement scheme to achieve the lowest AWL, under different network environments and encoding parameters. By encoding and placing each stripe of data with the optimal generator matrix and data placement scheme, ACIoT can achieve the minimum NRC. The experiments show that, compared with several state-of-the-art erasure codes, ACIoT reduces the NRC by 26.4%-44.7%.
Han Bao 0006, Yijie Wang 0001, Fangliang Xu
IEEE Internet Things J.3
2019 LAR: Locality-Aware Reconstruction for erasure-coded distributed storage systems
abstract
Summary Many modern distributed storage systems adopt erasure coding to protect data from frequent server failures for cost reason. Reconstructing data in failed servers efficiently is vital to these erasure‐coded storage systems. To this end, tree‐structured reconstruction mechanisms where blocks are transmitted and combined through a reconstruction tree have been proposed. However, existing tree‐structured reconstruction mechanisms build reconstruction trees from the perspective of available network bandwidths between servers, which are fluctuating and difficult to measure. Besides, these reconstruction mechanisms cannot reduce data transmission. In this study, we overcome these limitations by proposing LAR, a locality‐aware tree‐structured reconstruction mechanism. LAR builds reconstruction trees from the perspective of data locality, which is stable and easy to obtain. More importantly, by building reconstruction trees that combine blocks closer to each other first, LAR can reduce the data transmitted through the network core and hence speed up reconstruction. We prove that a minimum spanning tree is an optimal reconstruction tree that minimizes core bandwidth usage. We also design and implement a general reconstruction framework that supports all tree‐structured reconstruction mechanisms and nearly all erasure codes. Large‐scale simulations on commonly deployed network topologies show that LAR consumes 20%–61% less core bandwidth than previous reconstruction mechanisms. Thorough experiments on a testbed consisting of 40 physical servers show that LAR improves proactive recovery throughput by 23% at least and improves degraded read rate by up to 68%.
Fangliang Xu, Yijie Wang 0001, Xiaoqiang Pei, Xingkong Ma
Concurr. Comput. Pract. Exp.1
2018 Incremental encoding for erasure-coded cross-datacenters cloud storage
Fangliang Xu, Yijie Wang 0001, Xingkong Ma
Future Gener. Comput. Syst.1
2018 TA-Update: An Adaptive Update Scheme with Tree-Structured Transmission in Erasure-Coded Storage Systems
abstract
Erasure coding has received considerable attentions due to the better tradeoff between the space efficiency and reliability. The frequent update of the stored data in the distributed storage systems has posed a new challenge for erasure codes: how to update the erasure-coded data in a general, efficient and adaptive way. However, existing update schemes of erasure codes are inadequate to meet these requirements, since their code-related update manners lead to a low generality, their star-structured data transmission manners lead to a low update efficiency, and their redo manners when encountering the node failure lead to a low adaptivity. In this paper, we propose an adaptive update scheme with the tree-structured transmission, called TA-Update, which consists of a code-independent update framework and three algorithms: the rack-aware tree construction algorithm, the top-down data processing algorithm and the rollback-based failure processing algorithm. For generality, we propose a code-independent update framework with the tree structure to support the MDS code with any coding parameter. For efficiency, a rack-aware tree construction algorithm is proposed to achieve the high available bandwidth, which organizes the data node and parity nodes as an update tree. Moreover, a top-down data processing algorithm is proposed to achieve the high transmission and computation efficiency, which pipelines the data transmission along the update tree and distributes the encoding computations among all the participating nodes. For adaptivity, we propose a rollback-based failure processing algorithm to achieve high adaptivity, which handles the node failure during update with the existing update tree in a rollback manner. To evaluate the performance of TA-Update, we conduct experiments on HDFS-RAID under various parameter settings on both 30 physical and 200 virtual machines. Extensive experiments confirm that TA-Update could support the various erasure codes with any parameter, improve the update efficiency by 30 percent and the adaptivity by 47 percent on average compared with the state-of-the-art approaches under various parameter settings.
Yijie Wang 0001, Xiaoqiang Pei, Xingkong Ma, Fangliang Xu
IEEE Trans. Parallel Distributed Syst.4
2017 A cloud-assisted publish/subscribe service for time-critical dissemination of bulk content
abstract
Summary Characterized by the increasing arrival rate of live content, emergency applications pose a great challenge: how to disseminate data with diverse sizes to interested users in a real‐time manner. Most file sharing applications focus on the dissemination of bulk content with less consideration of users' interests. On the other hand, existing publish/subscribes are designed for notifying interested users with small‐sized content. To bridge this gap, we propose CAPS, a cloud‐assisted publish/subscribe service for time‐critical bulk content dissemination. In CAPS, a hybrid 2‐layer architecture is proposed to knit servers in the cloud and clients in the internet. Through dividing each event into attribute‐value pairs and the data content, CAPS provides both event matching service and data distribution in a parallel manner. To improve the upload bandwidth of data distribution, we propose a helper‐based content distribution protocol, where the servers not only guide the clients with similar interests to exchange their received data blocks but also contribute their own upload capacities to clients. Moreover, a volume‐aware helper renting scheme is proposed to adaptively adjust the scale of servers according to the churn of data volume, leading to a high‐performance price ratio. So as to evaluate the performance of CAPS, about 1000 virtual machines are deployed in our Cloud‐Stack testbed. Extensive experiments confirm that CAPS can linearly reduce the download completion time with the growing number of servers, adaptively adjust the upload capacity in tens of seconds according to the change of the workloads, and ensure reliable data dissemination even if a large number of nodes frequently churn or instantaneously fail. Compared with the state‐of‐the‐art approaches, CAPS demonstrates better performance under various parameter settings.
Xingkong Ma, Yijie Wang 0001, Xiaoqiang Pei, Fangliang Xu
Concurr. Comput. Pract. Exp.4
2017 A decentralized redundancy generation scheme for codes with locality in distributed storage systems
abstract
Summary The increasing data volume in a large number of applications presents a dire need for supporting the reliable data management in distributed storage systems. Existing classical erasure codes, such as the Reed‐Solomon codes and locally reconstruction codes, are widely adopted by many distributed storage systems. However, existing researches mainly focus on proposing new optimized codes, ignoring the optimization of the encoding process with the classical codes, where inefficient encoding process greatly degrades the encoding performance of the distributed storage systems. Thus, how to complete the encoding process in an efficient way has become the challenge for adopting the classical codes. In this paper, we propose a decentralized redundancy generation scheme on the basis of the codes with locality, called D2CP, where a 2‐step framework is proposed to support both the data patterns (replication to encodinganddirect encoding) and codes with locality with any parameter set. For improving the insertion throughput, D2CP adopts a data placement technique with consistent hashing to guide the selection of nodes. For reducing the network traffic cost, D2CP adopts a data sending scheduling technique to schedule the transmission of the source nodes and a cooperative parity generation technique to generate the parity data cooperatively. To evaluate the performance of D2CP, we conduct experiments on our RAID distributed storage system under various parameter settings with both 30 physical and 200 virtual servers. Extensive experiments confirm that D2CP can improve the encoding throughput by 20% and 32% and reduce the network traffic cost by 16% and 33% compared with the typical approaches on average for the 2 data patterns respectively.
Xiaoqiang Pei, Yijie Wang 0001, Xingkong Ma, Fangliang Xu
Concurr. Comput. Pract. Exp.4
2017 Efficient in-place update with grouped and pipelined data transmission in erasure-coded storage systems
Xiaoqiang Pei, Yijie Wang 0001, Xingkong Ma, Fangliang Xu
Future Gener. Comput. Syst.4
2016 T-Update: A tree-structured update scheme with top-down transmission in erasure-coded systems
abstract
Erasure coding has received considerable attention due to the better tradeoff between the space efficiency and reliability. However, it consumes large network traffic and long time to complete the update, involving updates of both data nodes and parity nodes. Existing solutions to this problem mainly focus on proposing new class of codes with lower update complexity to reduce the network traffic, ignoring the optimization of data transmission structure. In fact, the data transmission structure has great impact on the update. In this paper, we propose T-Update, a tree-structured update scheme with top-down transmission that minimizes the update time for erasure-coded data with no additional network traffic. Specially, we propose a rack-aware tree construction technique to construct an update tree to organize the data connections, with the data node as the root and the parity nodes as the children. To maximize the update efficiency, we propose a top-down data transmission technique to guide the data transmission and distribute the data computation for updating the parity nodes. To evaluate the performance of T-Update, we conduct experiments on HDFS-RAID under various parameter settings on both 30 physical and 200 virtual servers. Extensive experiments confirm that T-Update reduces the update time by 27% and 32% on average compared with two typical update schemes respectively.
Xiaoqiang Pei, Yijie Wang 0001, Xingkong Ma, Fangliang Xu
INFOCOM4
2016 Repairing multiple failures adaptively with erasure codes in distributed storage systems
abstract
Summary Repairs of multiple failures in distributed storage systems have posed the challenges for erasure coding: how to minimize the repair time with the least extra repair network traffic cost. However, existing repair schemes designed for single failure suffer from the high network traffic cost due to the serial repairs for multiple failures. Repair schemes designed for multiple failures suffer from long repair time due to the centralized repair structure. In this paper, we propose a decentralized adaptive repair scheme, called DARS, to minimize the repair time with the least extra network traffic cost. Specially, we propose a three‐layer repair model to support the repairs for both the single and multiple failures. For low repair time, a bandwidth‐aware node selection technique is proposed to guide the selection of nodes, and a line‐structured data transmission technique is proposed to organize the data transmission between the providers and the newcomer. For the least extra network traffic cost, a core‐based data distribution technique is proposed to organize the data transmission between the coordinator and other newcomers, and an intersection provider adjustment technique is proposed to adaptively adjust the number of intersection providers. Moreover, we adopt the ‘lazy repair’ within a stripe to further reduce the repair network traffic cost. We implement and evaluate DARS on our raid distributed storage system under various parameter settings with 30 physical machines and 200 virtual machines. Experimental results confirm that DARS reduces the repair time by 29% and 55% on average compared with tree‐structured repair and CORE, respectively. Copyright © 2015 John Wiley & Sons, Ltd.
Xiaoqiang Pei, Yijie Wang 0001, Xingkong Ma, Fangliang Xu
Concurr. Comput. Pract. Exp.4
2015 Scalable and elastic total order in content-based publish/subscribe systems
Xingkong Ma, Yijie Wang 0001, Xiaoqiang Pei, Fangliang Xu
Comput. Networks4
2014 MCRTREE: A Mutually Cooperative Recovery Scheme for Multiple Losses in Distributed Storage Systems Based on Tree Structure
abstract
To guarantee the reliability of distributed storage systems, erasure coding, as a redundant scheme, has received increasingly attention because it can greatly improve the space efficiency compared with the replica schemes. However, it takes a long time and consumes a lot of network bandwidth for erasure coding to repair the lost data on failed nodes. The state-of-art studies focus on the repairing optimization for the single-node-failure context. Real-world experiments have clearly shown that multi-node failures indeed happen in cloud storage systems. Borrowing single-node repairing techniques to the multi-node setting faces challenges on the efficiency. We propose a mutually cooperative recovery scheme MCRTREE based on the tree structure for multiple node failures. MCRTREE improves the bandwidth utilization and reduces the repair time by the construction of regeneration trees between each new node (denoted as newcomers) and alive nodes (denoted as providers). Further, MCRTREE reduces the size of the data volumes to be transmitted for the repair process. Numerical experiments show that MCRTREE consumes less storage cost and the maintenance bandwidth compared with other redundancy recovery schemes. Trace-driven simulation results reveal that the MCRTREE reduces the regeneration time by 30% - 50%, improves the successful regeneration probability by 10% - 20% and the data availability by 10% - 20% compared with the typical repair schemes.
Xiaoqiang Pei, Yijie Wang 0001, Xingkong Ma, Yongquan Fu, Fangliang Xu
NAS5