Guohua Shen

dblp:93/786 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Security and privacy · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 LevDetectCode: Zero-Shot Detection of AI-Generated Code Using Levenshtein Distance
abstract
AI-generated code is spreading rapidly, making it harder to trace the true origins of code. Text-based tools such as DetectGPT perform well on texts but struggle with code. We propose a new approach: under masked reconstruction, code produced by LLMs exhibits a much wider spread in Levenshtein distance than human-written code. Our key innovation is to treat this Levenshtein distance dispersion as an indicator for token-probability differences, leading us to develop LevDetectCode. Our zero-shot detector relies solely on string-level Levenshtein distance, requiring neither GPUs nor access to model internals. Experiments on Python snippets from MBPP-train and HumanEval, as well as Java data from CodeContest_Java_test, using GPT-3.5, GPT-4, and Deepseek-V3, show a 5–10 point AUROC improvement over other zero-shot baselines, with an additional case study for the C[Formula: see text] dataset, which also demonstrates strong effectiveness. Furthermore, it runs nearly 3000 times faster and significantly reduces memory usage. These results demonstrate that Levenshtein distance can outperform other zero-shot methods that rely on model-based computations, offering a practical and portable solution for detecting AI-generated code.
Jiazhou Fu, Guohua Shen, Yaoshen Yu
Int. J. Softw. Eng. Knowl. Eng.2
2025 Automatic IoT permission assignment with transformer models under spatiotemporal constraints
Guohua Shen, Jian Xie 0004, Jiazhou Fu
J. Inf. Secur. Appl.2
2024 ASTSDL: predicting the functionality of incomplete programming code via an AST-sequence-based deep learning model
Yaoshen Yu, Guohua Shen, Weiwei Li 0001, Yichao Shao
Sci. China Inf. Sci.3
2024 LoHDP: Adaptive local differential privacy for high-dimensional data publishing
abstract
Summary The increasing availability of high‐dimensional data collected from numerous users has led to the need for multi‐dimensional data publishing methods that protect individual privacy. In this paper, we investigate the use of local differential privacy for such purposes. Existing solutions calculate pairwise attribute marginals to construct probabilistic graphical models for generating attribute clusters. These models are then used to derive low‐dimensional marginals of these clusters, allowing for an approximation of the distribution of the original dataset and the generation of synthetic datasets. Existing solutions have limitations in computing the marginals of pairwise attributes and multi‐dimensional distribution on attribute clusters, as well as constructing relational dependency graphs that contain large clusters. To address these problems, we propose LoHDP, a high‐dimensional data publishing method composed of adaptive marginal computing and an effective attribute clustering method. The adaptive local marginal calculates anyk‐dimensional marginals required in the algorithm. In particular, methods such as sampling‐based randomized response are used instead of privacy budget splits to perturb user data. The attribute clustering method measures the correlation between pairwise attributes using an effective method, reduces the search space during the construction of the dependency graph using high‐pass filtering technology, and realizes dimensionality reduction by combining sufficient triangulation operation. We demonstrate through extensive experiments on real datasets that our LoHDP method outperforms existing methods in terms of synthetic dataset quality.
Guohua Shen, Mengnan Cai, Feifei Guo, Linlin Wei
Concurr. Comput. Pract. Exp.1
2024 FuEPRe: a fusing embedding method with attention for post recommendation
Xinbo Zhang, Guohua Shen, Yaoshen Yu
Serv. Oriented Comput. Appl.2
2023 Test Input Selection for Deep Neural Network Enhancement Based on Multiple-Objective Optimization
abstract
Deep Neural Networks (DNNs) have been applied in many domains, such as autonomous driving and image recognition. However, due to the lower-than-expected performance of the DNN models in the real application, researchers are committed to sampling a test subset from test data with the limited labeling effort to retrain the DNN models for enhancement. Existing test input selection methods aim at selecting the test inputs according to the probability that is classified incorrectly by the DNN model. However, the test inputs selected by using existing methods might have similar features, making the DNN model unable to learn more diverse features when retraining. To address this limitation, this paper proposes Multiple-Objective Optimization-Based Test Input Selection (MOTS) to select more effective test subset to retrain the DNN model for enhancement. Different from existing works, this work not only considers the uncertainty of the test input but also takes the diversity of the test subset into account. Then MOTS uses a multiple-objective optimization algorithm NSGA-II to solve this problem, which ensures the test subset has more diverse features and is more helpful for retraining DNN models. This paper conducts the experiment on two popular DNN models and three widely-used datasets. The experiment results indicate that MOTS achieves 114%, 72%, 55%, 41% average accuracy improvement under four different sampling ratios 1%, 3%, 5%, 10% compared with five baseline methods. Therefore, MOTS is very effective in improving the quality of the DNN models compared with the state-of-the-art methods and the diversity of the test subset contributes greatly to the effectiveness of retraining DNN models.
Yao Hao, Hongjing Guo, Guohua Shen
SANER4
2023 An accident prediction architecture based on spatio-clock stochastic and hybrid model for autonomous driving safety
abstract
Summary Collaborative and autonomous driving vehicles combine hardware and software complex processes, also are heavily dependent on and influenced by the world of physical and cyber interactions. They have enabled many new features and advanced functionalities, such as stochastic and hybrid natures, mobile spatial topologies, and time‐critical dependability. However, the existing modeling and verification techniques have not established faith in proving correctness and safety. Spatial and time collision avoidance remains crucial obstacles on the path to becoming ubiquitous and dependable. In order to ensure safety, we first design an accident prediction architecture in system design‐time and run‐time stages. We apply it on collaborative and autonomous overtaking systems involving spatial‐ and time‐critical accident predictions. Then, we develop a novel and dedicated spatio‐clock stochastic specification language (SCSSL) to describe safety invariants and guards in domain‐specific autonomous driving systems. Next, we create the spatio‐clock stochastic and hybrid automata models based on SCSSL in order to model inherently stochastic and hybrid behaviors. To illustrate the effectiveness of spatio‐clock consistency stochastic specification and verification, we adopt statistical model checking natively to provide reliable predictions for the incoming collision instants and positions. Finally, we present an illustrative overtaking case study to verify spatio‐clock stochastic and hybrid related properties and ensure correct modeling, and demonstrate the significance of our proposed approach.
Jinyong Wang, Tiexin Wang, Guohua Shen, Jian Xie 0004
Concurr. Comput. Pract. Exp.5
2023 Maneuver Conditioned Vehicle Trajectory Prediction Using Self-Attention
abstract
Forecasting the motion of surrounding vehicles is necessary for a self-driving vehicle to plan a safe and efficient trajectory for the future. Like experienced human drivers, the self-driving vehicle needs to perceive the interaction of surrounding vehicles and decide the best trajectory from many choices. However, previous methods either lack modeling of interactions or ignore the multi-modal nature of this problem. In this paper, we focus on two important cues of trajectory prediction: interaction and maneuver, and propose Maneuver conditioned Attentional Network named MAN. MAN learns the interactions of all vehicles in a scenario in parallel by self-attention social pooling and the attentional decoder generates the future trajectory conditioned on the predicted maneuver among 3 classes: Lane Changing Left (LCL), Lane Changing Right (LCR) and Lane Keeping (LK). Experiments demonstrate the improvement of our model in prediction on the publicly available NGSIM and HighD datasets. We also present quantitative analysis to study the relationship between maneuver prediction accuracy and trajectory error.
Junan Huang, Guohua Shen, Jinyong Wang, Xiaohua Yin
Int. J. Comput. Intell. Appl.3
2023 Shrinking the Semantic Gap: Spatial Pooling of Local Moment Invariants for Copy-Move Forgery Detection
abstract
Copy-move forgery is a manipulation of copying and pasting specific patches from and to an image, with potentially illegal or unethical uses. Recent advances in the forensic methods for copy-move forgery have shown increasing success in detection accuracy and robustness. However, for images with high self-similarity or strong signal corruption, the existing algorithms often exhibit inefficient processes and unreliable results. This is mainly due to the inherent semantic gap between low-level visual representation and high-level semantic concept. In this paper, we present a very first study of trying to mitigate the semantic gap problem in copy-move forgery detection, with spatial pooling of local moment invariants for midlevel image representation. Our detection method expands the traditional works on two aspects: 1) we introduce the bag-of-visual-words model into this field for the first time, may meaning a new perspective of forensic study; 2) we propose a word-to-phrase feature description and matching pipeline, covering the spatial structure and visual saliency information of digital images. Extensive experimental results show the superior performance of our framework over state-of-the-art algorithms in overcoming the related problems caused by the semantic gap.
Chao Wang 0028, Yaoshen Yu, Guohua Shen, Yushu Zhang 0001
IEEE Trans. Inf. Forensics Secur.5
2022 SIA-Net: Scalable Interaction-Aware Network for Vehicle Trajectory Prediction Based on Self-Attention
abstract
In order to navigate through different scenarios safely and efficiently, self-driving vehicles must predict future trajectories of other vehicles, which is a challenging task due to the implicit vehicle interactions in the driving scenario. Because there is no predefined number of surrounding vehicles, the model must be scalable to cope with scenarios of different vehicle numbers with high accuracy and low computation cost for jointly predicting future trajectories of all vehicles. However, previous methods mainly focus on predicting a single trajectory of the target vehicle, which makes them subject to accuracy and computation speed. In this paper, we propose SIA-Net that predicts future trajectories of all vehicles in the scenario independent of the vehicle number. SIA-Net learns the implicit interactions of all vehicles by self-attention social pooling and generates each trajectory through one forward propagation by attentional decoder. Experiments demonstrate the improvement of our model in prediction accuracy on the publicly available NGSIM and INTERACTION datasets while keeping the computation cost low. We also present qualitative analysis to study the mechanism of our model.
Junan Huang, Guohua Shen, Gaoyang Hua
ICTAI3
2022 Statistical Model Checking for Stochastic and Hybrid Autonomous Driving Based on Spatio-Clock Constraints
abstract
Autonomous driving vehicles are a kind of typical cyber-physical systems integrating complex interactions between hardware and software components such as collaborative computation, distributed communication, and spatio-clock synchronous control with surrounding traffic environment. They can percept the environment, communicate with surroundings, and react fast enough to control independently. The purpose of autonomous driving emergence is to improve driving safety, reduce environmental pollution, and ease the traffic congestion. However, new features with surrounding open and dynamic environment make systems design and verification becoming more and more complex than ever, such as stochastic communication delay, hardware spontaneous failure distribution, and natively hybrid behaviors described by ordinary differential equations. Spatial and time collision avoidance remains crucial obstacles on the path to becoming ubiquitous and dependable. In this paper, we adopt statistical model checking (SMC) to enlighten possible hazards affected by stochastic and hybrid features in the design phase of autonomous driving systems. In order to provide safety and accountability, we first propose a dedicated multi-lane spatio-clock stochastic specification language (MLSCL) to describe safety invariants and guards in domain-specific autonomous driving systems. Then, we present the semantic mapping rules between MLSCL and UPPAAL SMC models, and design the spatio-clock stochastic and hybrid automata based on MLSCL in order to model inherently stochastic and hybrid behaviors. Finally, we present an illustrative lane-change case study to verify spatio-clock stochastic and hybrid-related properties adopting SMC, and demonstrate the effectiveness of our proposed approach.
Jinyong Wang, Yi Zhu 0008, Guohua Shen
Int. J. Softw. Eng. Knowl. Eng.4
2022 ASTENS-BWA: Searching partial syntactic similar regions between source code fragments via AST-based encoded sequence alignment
Yaoshen Yu, Guohua Shen, Weiwei Li 0001, Yichao Shao
Sci. Comput. Program.3
2022 SafeOSL: Ensuring memory safety of C via ownership-based intermediate language
abstract
Abstract The unsafe features of C make it a big challenge to ensure memory safety of C programs, and often lead to memory errors that can result in vulnerabilities. Various formal verification techniques for ensuring memory safety of C have been proposed. However, most of them either have a high overhead, such as state explosion problem in model checking, or have false positives, such as abstract interpretation. In this article, by innovatively borrowing ownership system from Rust, we propose a novel and sound static memory safety analysis approach, named SafeOSL. Its basic idea is an ownership‐based intermediate language, called ownership system language (OSL), which captures the features of the ownership system in Rust. Ownership system specifies the relations among variables and memory locations, and maintains invariants that can ensure memory safety. The semantics of OSL is formalized in K‐framework, which is a rewriting‐logic based tool. C programs to be checked are first transformed into OSL programs and then detected by OSL semantics. Experimental results have demonstrated that SafeOSL is effective in detecting memory errors of C. Moreover, the translations and experiments indicate that the intermediate language OSL could be reused by other programming languages to detect memory errors.
Xiaohua Yin, Shuanglong Kan, Guohua Shen, Zhe Chen 0011, Yang Liu 0003, Fei Wang 0032
Softw. Pract. Exp.4
2021 A security policy model transformation and verification approach for software defined networking
Yunfei Meng, Guohua Shen, Changbo Ke
Comput. Secur.3
2021 Improved Entity Linking for Simple Question Answering Over Knowledge Graph
abstract
Question Answering systems over Knowledge Graphs (KG) answer natural language questions using facts contained in a knowledge graph, and Simple Question Answering over Knowledge Graphs (KG-SimpleQA) means that the question can be answered by a single fact. Entity linking, which is a core component of KG-SimpleQA, detects the entities mentioned in questions, and links them to the actual entity in KG. However, traditional methods ignore some information of entities, especially entity types, which leads to the emergence of entity ambiguity problem. Besides, entity linking suffers from out-of-vocabulary (OOV) problem due to the limitation of pre-trained word embeddings. To address these problems, we encode questions in a novel way and encode the features contained in the entities in a multilevel way. To evaluate the enhancement of the whole KG-SimpleQA brought by our improved entity linking, we utilize a relatively simple approach for relation prediction. Besides, to reduce the impact of losing the feature during the encoding procedure, we utilize a ranking algorithm to re-rank (entity, relation) pairs. According to the experimental results, our method for entity linking achieves an accuracy of 81.8% that beats the state-of-the-art methods, and our improved entity linking brings a boost of 5.6% for the whole KG-SimpleQA.
Guohua Shen, Haijuan Wang
Int. J. Softw. Eng. Knowl. Eng.2
2021 Supporting Requirements to Code Traceability Creation by Code Comments
abstract
Requirements-to-code tracing is an important and costly task that creates trace links from requirements to source code. These trace links help engineers reduce the time and complexity of software maintenance. Code comments play an important role in software maintenance tasks. However, few studies have focused intensively on the impact of code comments on requirements-to-code trace links creation. Different types of comments have different purposes, so how different types of code comments provide different improvements for requirements-to-code trace links creation? We focus on learning whether code comments and different types of comments can improve the quality of trace links creation. This paper presents a study to evaluate the contribution of code comments and different types of code comments to the creation of trace links. More specifically, this paper first experimentally evaluates the impact of code comments on requirements-to-code trace links creation, and then divides code comments into six categories to evaluate its impact on trace links creation. The results show that the precision increases by an average of 15% (based on the same recall) after adding code comments (even for different trace links creation techniques), and the type of Purpose comments contributes more to the tracing task than the other five. This empirical study provides evidence that code comments are effective in tracing links creation, and different types of code comments contribute differently. Purpose comments can be used to improve the accuracy of requirements-to-code trace links creation.
Guohua Shen, Haijuan Wang, Yaoshen Yu
Int. J. Softw. Eng. Knowl. Eng.1
2021 Analyzing close relations between target artifacts for improving IR-based requirement traceability recovery
abstract
Requirement traceability is an important and costly task that creates trace links from requirements to different software artifacts. These trace links can help engineers reduce the time and complexity of software maintenance. The information retrieval (IR) technique has been widely used in requirement traceability. It uses the textual similarity between software artifacts to create links. However, if two artifacts do not share or share only a small number of words, the performance of the IR can be very poor. Some methods have been developed to enhance the IR by considering relations between target artifacts, but they have been limited to code rather than to other types of target artifacts. To overcome this limitation, we propose an automatic method that combines the IR method with the close relations between target artifacts. Specifically, we leverage close relations between target artifacts rather than just text matching from requirements to target artifacts. Moreover, the method is not limited to the type of target artifacts when considering the relations between target artifacts. We conduct experiments on five public datasets and take account of trace links between requirements and different types of software artifacts. Results show that under the same recall, the precisions on the five datasets improve by 40%, 8%, 20%, 4%, and 6%, respectively, compared with the baseline method. The precision on the five datasets improves by an average of 15.6%, showing that our method outperforms the baseline method when working under the same conditions.
Haijuan Wang, Guohua Shen, Yaoshen Yu
Frontiers Inf. Technol. Electron. Eng.2
2020 Automatic traceability link recovery via active learning
abstract
Traceability link recovery (TLR) is an important and costly software task that requires humans establish relationships between source and target artifact sets within the same project. Previous research has proposed to establish traceability links by machine learning approaches. However, current machine learning approaches cannot be well applied to projects without traceability information (links), because training an effective predictive model requires humans label too many traceability links. To save manpower, we propose a new TLR approach based on active learning (AL), which is called the AL-based approach. We evaluate the AL-based approach on seven commonly used traceability datasets and compare it with an information retrieval based approach and a state-of-the-art machine learning approach. The results indicate that the AL-based approach outperforms the other two approaches in terms of F-score.
Tianbao Du, Guohua Shen, Yaoshen Yu, Dexiang Wu
Frontiers Inf. Technol. Electron. Eng.2
2020 SDN-Based Security Enforcement Framework for Data Sharing Systems of Smart Healthcare
abstract
As novel healthcare paradiagm, smart healthcare can provide more efficient and high quality medical services for patients. However, smart healthcare needs patients to share their physiological information for online diagnoses, if the data sharing system of smart healthcare lacks effective security mechanisms, these sensitive information might be abused by illegal or malicious users. Moreover, smart healthcare needs to confront some brand-new challenges, such as resource-constrained IoT things, identity theft attacks and insider attacks. To tackle these problems, we propose a SDN-based security enforcement framework for data sharing systems of smart healthcare. In our framework, each patient has a dedicated virtual machine in data sharing system, each virtual machine provides a group data services which can be released to those authorized service consumers or IoT things. In additon, virtual machine is protected by the SDN-based gateway which provides a firewall mechanism and guarantees only authorized things can access patient's virtual machine. Since each thing has a unique MAC address, thus our framework can effectively authenticate resource-constrained IoT things and tackle the problems caused by identity theft. To validate the effectiveness and feasibility of our framework, we implement an experimental system using POX controller and Mininet emulator. The experimental results illustrate our framework is effective under different test scenarios. As increasing the scale of information flow model, the framework can still work well and its performance can be still acceptable.
Yunfei Meng, Guohua Shen, Changbo Ke
IEEE Trans. Netw. Serv. Manag.3
2019 Deep image reconstruction from human brain activity
abstract
The mental contents of perception and imagery are thought to be encoded in hierarchical representations in the brain, but previous attempts to visualize perceptual contents have failed to capitalize on multiple levels of the hierarchy, leaving it challenging to reconstruct internal imagery. Recent work showed that visual cortical activity measured by functional magnetic resonance imaging (fMRI) can be decoded (translated) into the hierarchical features of a pre-trained deep neural network (DNN) for the same input image, providing a way to make use of the information from hierarchical visual features. Here, we present a novel image reconstruction method, in which the pixel values of an image are optimized to make its DNN features similar to those decoded from human brain activity at multiple layers. We found that our method was able to reliably produce reconstructions that resembled the viewed natural images. A natural image prior introduced by a deep generator neural network effectively rendered semantically meaningful details to the reconstructions. Human judgment of the reconstructions supported the effectiveness of combining multiple DNN layers to enhance the visual quality of generated images. While our model was solely trained with natural images, it successfully generalized to artificial shapes, indicating that our model was not simply matching to exemplars. The same analysis applied to mental imagery demonstrated rudimentary reconstructions of the subjective content. Our results suggest that our method can effectively combine hierarchical neural representations to reconstruct perceptual and subjective images, providing a new window into the internal contents of the brain.
Guohua Shen, Tomoyasu Horikawa, Kei Majima, Yukiyasu Kamitani
PLoS Comput. Biol.1
2017 Quantitative risk analysis of safety-critical embedded systems
Yinling Liu, Guohua Shen, Zhibin Yang 0005
Softw. Qual. J.2
2012 Feature modeling and Verification based on Description Logics
Guohua Shen, Changbao Tian, Qiang Ge
SEKE1
2009 A Semantic Model for Matchmaking of Web Services Based on Description Logics
abstract
Matchmaking plays an important role in Web services interactions. The matchmaking based on keywords easily leads to low precision, Meanwhile, the current semantic service discovery methods perform service I/O based profile matching, there exists no matchmaker that performs an integrated service matching by additional reasoning on logically defined preconditions, effects, Qos and so on. In this paper, the semantic web services are described based on Description Logics, and the services description model is designed, which describes the various aspects (such as IOPEs, Qos and so on) of the web services. So the services matchmaking is transformed into the match of concepts. The service match algorithm is proposed and the description logics reasoner RacerPro is adopted for Web services discovery. We show how the semantic matching between providers and a requester is performed by a case study.
Guohua Shen
Fundam. Informaticae1