Chaochen Shi

dblp:249/5046 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0002-5543-1655ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2025 MM-SCS: Leveraging Multimodal Features to Enhance Smart Contract Code Search
abstract
Semantic code search technology allows searching for existing code snippets through natural language, which can greatly improve programming efficiency. Smart contracts, programs that run on the blockchain, have a code reuse rate of more than 79%, which means developers have a great demand for semantic code search tools. However, the existing code search models still have a semantic gap between code and query and perform poorly on specialized queries of smart contracts. In this paper, we propose a Multi-Modal Smart contract Code Search (MM-SCS) model. Specifically, we construct a Contract Elements Dependency Graph (CEDG) for MM-SCS as an additional modality to capture the data flow and control flow information of the code. To make the model more focused on the key contextual information, we use a multi-head attention network to generate embeddings for code features. In addition, we use a fine-tuned pretrained model to ensure the model's effectiveness when the training data is small. We compared MM-SCS with four state-of-the-art models on a dataset with 470K (code, docstring) pairs collected from Github and Etherscan. Experimental results show that MM-SCS achieves an MRR (Mean Reciprocal Rank) of 0.572, outperforming four state-of-the-art models UNIF, DeepCS, CARLCS-CNN, and TAB-CS by 34.2%, 59.3%, 36.8%, and 14.1%, respectively. Additionally, the search speed of MM-SCS is second only to UNIF, reaching 0.34s/query.
Chaochen Shi, Yong Xiang 0001, Jiangshan Yu, Longxiang Gao
IEEE Trans. Software Eng.1
2024 Long-Term Over One-Off: Heterogeneity-Oriented Dynamic Verification Assignment for Edge Data Integrity
abstract
EdgeIntelligence (EI), a burgeoning research area, motivates App vendors to cache data replicas on geographically distributed edge servers to deliver better services. On the downside, this benefit also incurs more data integrity audit overhead on App vendors, which calls for more efficientEdgeDataIntegrity (EDI) verification approaches. However, existing EDI solutions totally rely on an implicitresource homogeneity assumption-edge servers have identical resource availability throughout EDI inspection execution in each round-but it rarely holds in reality. The edge servers with insufficient computation and/or communication capacity greatly limit overall EDI verification efficiency from a round perspective. Thus, in this work, we release the identified impractical assumption and accordingly study the EDIDynamicVerificationAssignment (DVA) problem for the first time. The problem aims to maximize the number of data replicas being verified in the long term under the constraints of verification delay in resource-limited environments. In this way, App vendors merely need to check the integrity of selected data replicas in each round for efficiency improvement. Specifically, we first formalize the DVA problem as a delay-constrained long-term stochastic optimization problem and further prove its$\mathcal {NP}$-hardness. To resolve the problem efficiently, we decompose it to an easy-to-handle form and then develop a polynomial-timePriority-based approach named DVA-P with a theoretical analysis of its time complexity and performance bound. Finally, experimental evaluations validate that DVA-P can be seamlessly incorporated into existing EDI solutions to enhance overall verification efficiency while guaranteeing verification performance.
Yao Zhao 0006, Youyang Qu, Yong Xiang 0001, Chaochen Shi, Feifei Chen 0001, Longxiang Gao
IEEE Trans. Mob. Comput.4
2023 Blockchain search engine: Its current research status and future prospect in Internet of Things network
Jine Tang, Xinming Lu, Yong Xiang 0001, Chaochen Shi, Junhua Gu
Future Gener. Comput. Syst.4
2023 Machine translation-based fine-grained comments generation for solidity smart contracts
Chaochen Shi, Yong Xiang 0001, Jiangshan Yu, Keshav Sood, Longxiang Gao
Inf. Softw. Technol.1
2023 CoSS: Leveraging Statement Semantics for Code Summarization
abstract
Automated code summarization tools allow generating descriptions for code snippets in natural language, which benefits software development and maintenance. Recent studies demonstrate that the quality of generated summaries can be improved by using additional code representations beyond token sequences. The majority of contemporary approaches mainly focus on extracting code syntactic and structural information from abstract syntax trees (ASTs). However, from the view of macro-structures, it is challenging to identify and capture semantically meaningful features due to fine-grained syntactic nodes involved in ASTs. To fill this gap, we investigate how to learn more code semantics and control flow features from the perspective of code statements. Accordingly, we propose a novel model entitled CoSS for code summarization. CoSS adopts a Transformer-based encoder and a graph attention network-based encoder to capture token-level and statement-level semantics from code token sequence and control flow graph, respectively. Then, after receiving two-level embeddings from encoders, a joint decoder with a multi-head attention mechanism predicts output sequences verbatim. Performance evaluations on Java, Python, and Solidity datasets validate that CoSS outperforms nine state-of-the-art (SOTA) neural code summarization models in effectiveness and is competitive in execution efficiency. Further, the ablation study reveals the contribution of each model component.
Chaochen Shi, Borui Cai, Yao Zhao 0006, Longxiang Gao, Keshav Sood, Yong Xiang 0001
IEEE Trans. Software Eng.1
2022 Towards Accurate Knowledge Transfer between Transformer-based Models for Code Summarization
abstract
Automatic code summarization generates high-level natural language descriptions of code snippets, which can benefit software maintenance and code comprehension.Recently, Transformer-based models achieved state-of-the-art performance on code summarization tasks.However, there are data gaps in neural model training for some programming languages.To fill this gap, we propose a novel transfer learning approach to accurately transfer knowledge between Transformer-based models.We train a discriminator to identify which heads of the multi-head attention module should be transferred.On this basis, we define a transfer strategy of parameter matrices.We evaluated the proposed transfer learning approach on four state-of-the-art Transformer-based code summarization models.Experimental results show that models with transferred knowledge outperform original models up to 10.70% in BLEU, 5.36% in ROUGE-L, and 4.34% in METEOR.
Chaochen Shi, Yong Xiang 0001, Jiangshan Yu, Longxiang Gao
SEKE1
2022 A Bytecode-based Approach for Smart Contract Classification
abstract
With the development of blockchain technologies, the number of smart contracts deployed on blockchain platforms is growing exponentially, which makes it difficult for users to find desired services by manual screening. The automatic classification of smart contracts can provide blockchain users with keyword-based contract searching and helps to manage smart contracts effectively. Current research on smart contract classification focuses on Natural Language Processing (NLP) solutions which are based on contract source code. However, more than 94% of smart contracts are not open-source, so the application scenarios of NLP methods are very limited. Meanwhile, NLP models are vulnerable to adversarial attacks. This paper proposes a classification model based on features from contract bytecode instead of source code to solve these problems. We also use feature selection and ensemble learning to optimize the model. Our experimental studies on over 11K real-world Ethereum smart contracts show that our model can classify smart contracts without source code and has better performance than baseline models. Our model also has good resistance to adversarial attacks compared with NLP-based models. In addition, our analysis reveals that account features used in many smart contract classification models have little effect on classification and can be excluded.
Chaochen Shi, Yong Xiang 0001, Jiangshan Yu, Longxiang Gao, Keshav Sood, Robin Doss
SANER1
2019 A Hidden Markov Model-Based Method for Virtual Machine Anomaly Detection
Chaochen Shi, Jiangshan Yu
ProvSec1