Shihao Xia

dblp:314/7084 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0006-7334-7701ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SymGPT: Auditing Smart Contracts via Combining Symbolic Execution with Large Language Models
abstract
This paper introduces SymGPT , a tool that combines LLMs with symbolic execution to automatically verify smart contracts’ compliance with ERC rules. We begin by empirically analyzing 132 ERC rules from three major ERC standards, examining their content, security implications, and natural language descriptions. Based on this study, SymGPT instructs an LLM to translate ERC rules into a domain-specific language, synthesizes constraints from the translated rules to model potential rule violations, and performs symbolic execution for violation detection. Our evaluation shows that SymGPT identifies 5,783 ERC rule violations in 4,000 real-world contracts, including 1,375 violations with clear attack paths for financial theft. Furthermore, SymGPT outperforms six automated techniques and a security-expert auditing service, underscoring its superiority over current smart contract analysis methods.
Shihao Xia, Tingting Yu 0001, Yiying Zhang 0005, Nobuko Yoshida, Linhai Song
Proc. ACM Program. Lang.1
2025 How to Save My Gas Fees: Understanding and Detecting Real-World Gas Issues in Solidity Programs
abstract
The execution of smart contracts on Ethereum, a public blockchain system, incurs a fee called gas fee for its computation and data storage. When programmers develop smart contracts (e.g., in the Solidity programming language), they could unknowingly write code snippets that unnecessarily cause more gas fees. These issues, or what we call gas wastes, can lead to significant monetary losses for users. This paper takes the initiative in helping Ethereum users reduce their gas fees in two key steps. First, we conduct an empirical study on gas wastes in open-source Solidity programs and Ethereum transaction traces. Second, to validate our study findings, we develop a static tool called PeCatch to effectively detect gas wastes in Solidity programs, and manually examine the Solidity compiler’s code to pinpoint implementation errors causing gas wastes. Overall, we make 11 insights and four suggestions, which can foster future tool development and programmer awareness, and fixing our detected bugs can save $0.76 million in gas fees daily.
Shihao Xia, Boqin Qin, Nobuko Yoshida, Tingting Yu 0001, Yiying Zhang 0005, Linhai Song
IEEE Trans. Software Eng.2
2024 DeepPepPI: A deep cross-dependent framework with information sharing mechanism for predicting plant peptide-protein interactions
Zhaowei Wang 0005, Jun Meng, Qiguo Dai, Shihao Xia, Ruirui Yang, Yushi Luan
Expert Syst. Appl.5
2023 TGAAL: Combining Transformer-based GAN and active learning to identify the coding potential of sORFs in plant lncRNAs
abstract
Some small open reading frames (sORFs) in plant long non-coding RNAs (lncRNAs) are capable of encoding small peptides, which play key roles in the growth and development of organisms. Therefore, it is particularly important to identify the coding potential of sORFs in plant lncRNAs. However, existing methods often ignore the differences in length distribution between coding sORFs (csORFs) and non-coding sORFs (non-csORFs), which may lead to incorrect identification of csORFs. To address this issue, we propose a novel method to identify the coding potential of sORFs in plant lncRNAs, named Transformer Generative Adversarial Active Learning (TGAAL), which combines Transformer-based Generative Adversarial Network (TGAN) and active learning based on KL-topk sampling strategy. TGAN can generate sORF sequences in a specific length interval, which have the same class as the input sORFs. Meanwhile, using active learning based on KL-topk sampling strategy, samples with high confidence can be selected for data augmentation. 5-fold cross-validation shows that KL-topk sampling strategy significantly improves the prediction performance compared with commonly adopted sampling strategies. The experimental results show that TGAAL significantly outperforms existing methods in identifying the coding potential of sORFs in Arabidopsis thaliana, reaching 0.7761, 0.7906 and 0.7529 unweighted average recall in three sORF length intervals, respectively.
Jun Meng, Shihao Xia, Zhaowei Wang 0005, Yushi Luan
BIBM3
2023 A multi-granularity information-enhanced pre-training method for predicting the coding potential of sORFs in plant lncRNAs
abstract
Small open reading frames (sORFs) are nucleotide sequences that may be translated into small peptides. Recently, increasing studies have demonstrated that peptides encoded by sORFs in plant long noncoding RNAs (lncRNAs) play a vital role in growth regulation and disease treatment. To accelerate the discovery of lncRNA-encoded peptides, it is essential to predict translatable sORFs in lncRNAs (lncRNA-sORFs) by computational methods. As only a few translatable plant lncRNA-sORFs have been discovered to date, there is a lack of effective methods for characterizing the coding potential of lncRNA-sORFs in data-scarce scenarios. Therefore, a novel method for plant lncRNA-sORFs coding potential prediction using the pre-trained bidirectional encoder representations from transformer (LSCPP-BERT) is proposed. Firstly, the BERT model is trained to extract multi-granularity context information from large-scale unlabeled lncRNA-sORFs through two pre-training tasks. Then, the pre-trained model can be fine-tuned with two additional linear layers for classification. The LSCPP-BERT is featured by a self-supervised pre-training scheme and multi-granularity context information, aiming to enhance the representational power of the network. In addition, an extra pre-training task called contextual relation of lncRNA-sORFs prediction (CRSP) is presented to extract sentence-level information. Experiment results show that the accuracy of LSCPP-BERT is increased by 8.14% compared with state-of-the-art methods. We hope that the proposed method can serve as a reliable tool for the prediction of coding lncRNA-sORFs, thereby further contributing to drug development and agronomical applications.
Shihao Xia, Jun Meng, Zhaowei Wang 0005, Zhaojing Qin, Yushi Luan
BIBM1
2022 Who goes first? detecting go concurrency bugs via message reordering
abstract
Go is a young programming language invented to build safe and efficient concurrent programs. It provides goroutines as lightweight threads and channels for inter-goroutine communication. Programmers are encouraged to explicitly pass messages through channels to connect goroutines, with the purpose of reducing the chance of making programming mistakes and introducing concurrency bugs. Go is one of the most beloved programming languages and has already been used to build many critical infrastructure software systems in the data-center environment. However, a recent study shows that channel-related concurrency bugs are still common in Go programs, severely hurting the reliability of the programs.
Shihao Xia, Yu Liang 0002, Linhai Song, Hong Hu 0004
ASPLOS2