Shuhui Fan

dblp:275/5256 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
5since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Fuzzing JavaScript Engines with a Graph-based IR
abstract
Mutation-based fuzzing effectively discovers defects in JS engines. High-quality mutations are key for the performance of mutation-based fuzzers. The choice of the underlying representation (e.g., a sequence of tokens, an abstract syntax tree, or an intermediate representation) defines the possible mutation space and subsequently influences the design of mutation operators. Current program representations in JS engine fuzzers center around abstract syntax trees and customized bytecode-level intermediate languages. However, existing efforts struggle to generate semantically valid and meaningful mutations, limiting the discovery of defects in JS engines.
Zhiyuan Jiang, Shuhui Fan, Shenglin Xu, Peidai Xie, Shaojing Fu, Mathias Payer
CCS4
2024 Edge-feature Modeling-based Topological Graph Neural Networks for Phishing Scams Detection on Ethereum
abstract
Detecting phishing scams has become an important task in blockchain-based cryptocurrency applications. While many network representation learning-based approaches have been proposed for this task, they suffer from various issues including (1) the requirement of handcrafted features, which may not capture complex relationships and patterns in graph data, and/or (2) considering only node features while ignoring the more significant edge features, and/or (3) incapability of preserving complete network topology, which affects the generalization ability. In this paper, we propose a novel Edge-feature modeling-based Topological Graph Neural Network (ETGNN) to detect phishing scams on Ethereum, which avoids all aforementioned issues of existing approaches. Specifically, ETGNN involves two key components, one responsible for learning weighted features of nodes and edges in the Ethereum transaction graph, and the other responsible for incorporating global topological information of the graph using persistent homology. Finally, phishing scams are detected based on these two learned features. The experimental results demonstrate that ETGNN outperforms the state-of-the-art method with an improvement rate of 14.38% on F1-score.
Shuhui Fan, Shaojing Fu, Yuchuan Luo, Ming Xu 0002
IWQoS1
2024 Fuzzing JavaScript engines with a syntax-aware neural program model
Zhiyuan Jiang, Shuhui Fan, Shaojing Fu, Peidai Xie
Comput. Secur.4
2022 Smart Contract Scams Detection with Topological Data Analysis on Account Interaction
abstract
The skyrocketing market value of cryptocurrencies has prompted more investors to pour funds into cryptocurrencies to seek asset hedging. However, the anonymity of blockchain makes cryptocurrency naturally a tool of choice for criminals to commit smart contract scams. Consequently, smart contract scam detection is particularly critical for investors to avoid economic loss. Previous methods mainly leverage specific code logic of smart contracts and/or design rules based on abnormal transaction behaviors for scam detection. Although these methods gain success at detecting particular scams, they perform worse when applied to scams with highly similar codes. Besides, well-designed decision rules rely on expert knowledge and tedious data collection steps, which causes poor flexibility. To combat these challenges, we consider the problem of smart contract scam detection via mining topological features of account interaction information that dynamically evolves. We adopt interactive features extracted from dynamic interaction information of accounts and propose a framework named TTG-SCSD to utilize the features and Topological Data Analysis for smart contract scams detection. The TTG-SCSD constructs discrete dynamic interaction graphs for each contract and designs interactive features that characterize account behaviors. The features are modeled combined with a topology quantification mechanism to capture contract intentions in transactions. Experimental results on real-world transaction datasets from Ethereum show that TTG-SCSD obtains better generalizability and improves the performance of the bare versions of the comparison methods.
Shuhui Fan, Shaojing Fu, Yuchuan Luo, Xuyun Zhang, Ming Xu 0002
CIKM1
2021 Al-SPSD: Anti-leakage smart Ponzi schemes detection in blockchain
Shuhui Fan, Shaojing Fu, Xiaochun Cheng
Inf. Process. Manag.1
2020 Tree2tree Structural Language Modeling for Compiler Fuzzing
Shuhui Fan, Hongzuo Xu, Peidai Xie
ICA3PP (1)2
2020 Expose Your Mask: Smart Ponzi Schemes Detection on Blockchain
abstract
The anonymity of blockchain has caused Ponzi schemes to be transferred to smart contract platforms by scammers. These Ponzi schemes wearing the mask of smart contracts caused huge losses to people, which makes the detection of smart Ponzi schemes attract people's attention. Recent methods mainly focus on machine learning technology to enable automatic detection for smart Poniz schemes. However, there are some problems with their methods. Firstly, the gradient boosting algorithm in machine learning they used have the problem of prediction shift due to target leakage when processing category features and calculating gradient estimates. Secondly, they ignored the imbalance and repetitiveness of Ponzi schemes on smart contract platforms. These problems can directly lead to model overfitting and affect the generalization ability of trained models. This paper proposes a novel Ponzi schemes detection method on smart contract platform for blockchain. Our method addresses the above issues with the following strategies. Firstly, we leverage ordered target statistic (TS) to process the category features of smart contract. Secondly, we solve the imbalance of dataset through a data augmentation method. Thirdly, with the idea of ordered boosting algorithm, we train a PonziTect model to fight prediction shift caused by target leakage. Based on the above ideas, the experimental results fully manifest the effectiveness and reliability of our model in detecting smart Ponzi schemes on the blockchain. Specifically, our model achieves 98% F-score on the real-world dataset, which significantly outperforms the existing methods. Using our method, we estimate that there are about 532 Ponzi schemes on Ethereum.
Shuhui Fan, Shaojing Fu, Chengzhang Zhu
IJCNN1
2020 DSmith: Compiler Fuzzing through Generative Deep Learning Model with Attention
abstract
Compiler fuzzing is a technique to test the functionalities of compiler. It requires well-formed test cases (i.e., programs) that have correct lexicons and syntax to pass the parsing stage of a compiler. Recently, advanced compiler fuzzing methods generate effective test cases by deep neural networks, which learn the language model of regular programs to guarantee test case quality. However, most of these methods fail to capture long-distance dependencies of syntax (e.g., paired curly braces) in a program. As a result, they may generate test cases with syntax errors, which cannot pass the parsing stage to test the compiler functionality. In this paper, we propose a framework, namely DSmith, to capture long-distance dependencies of syntax for a robust test case generation. Specifically, DSmith memorizes the hidden state of each token in a program and leverages the interactions of these hidden states to embed the long-distance dependencies between tokens. It then adopts an encoder-decoder architecture with the embedding of these long-distance dependencies to build a language model of regular programs. Finally, DSmith uses the built language model to generate test cases according to four novel generation strategies, which significantly increase the diversity of test cases. Extensive experiments show that DSmith increases the parsing pass rate of the generated programs by an average of 19% and significantly improves the code coverage of the compiler, compared with state-of-the-art methods. Benefiting from the high pass rate and broad code coverage, DSmith has found eleven brand new bugs in currently supported GCC compiler versions.
Shuhui Fan, Peidai Xie, Aizhi Liu
IJCNN3