Gaigai Tang

dblp:263/9696 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-1798-5718ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Discrepancy-aware contrastive learning with mixture of experts for cross-modal image-text semantic alignment
Gaigai Tang, Kaiyuan Qi, Guangfeng Su, Huiyun Zhang
Neurocomputing3
2025 Automatic generation of industrial internet attack graphs with graph neural networks and Bayesian models
Gaigai Tang, Kaiyuan Qi, Guangfeng Su, Huiyun Zhang
Comput. Networks2
2024 MRC-VulLoc: Software source code vulnerability localization based on multi-choice reading comprehension
Gaigai Tang, Lin Yang 0031, Long Zhang 0004, Hongyu Kuang
Comput. Secur.1
2023 Leveraging User-Defined Identifiers for Counterfactual Data Generation in Source Code Vulnerability Detection
abstract
Software vulnerability detection is a critical aspect of ensuring the security and reliability of software systems. However, traditional vulnerability detection approaches often have limitations due to the scarcity and need for more diversity in labeled data. This research introduces a novel approach to overcome these challenges by utilizing user-defined identifiers in the source code to generate counterfactual training data. User-defined identffiers, such as variable and function names, contain essential information about the intentions and logic of the program. By perturbing these identifiers while maintaining the syntactic and semantic structure of the code, we create a diverse set of counterfactual examples that simulate potential vulnerabilities. When combined with existing labeled data, these counterfactual examples enrich the training process for vulnerability detection models. To evaluate the effectiveness of our approach, we conduct experiments on various datasets, achieving state-of-the-art performance on the VulDeePecker and Draper datasets. Our approach also outperforms models that utilize the same pre-trained language model in terms of accuracy.
Hongyu Kuang, Long Zhang 0004, Gaigai Tang, Lin Yang 0031
SCAM4
2021 Interpretation of Learning-Based Automatic Source Code Vulnerability Detection Model Using LIME
Gaigai Tang, Long Zhang 0004, Lianxiao Meng, Weipeng Cao, Meikang Qiu, Shuangyin Ren, Lin Yang 0031
KSEM1
2021 An Automatic Source Code Vulnerability Detection Approach Based on KELM
abstract
Traditional vulnerability detection mostly ran on rules or source code similarity with manually defined vulnerability features. In fact, these vulnerability rules or features are difficult to be defined accurately, which usually cost much expert labor and perform weakly in practical applications. To mitigate this issue, researchers introduced neural networks to automatically extract features to improve the intelligence of vulnerability detection. Bidirectional Long Short-term Memory (Bi-LSTM) network has proved a success for software vulnerability detection. However, due to complex context information processing and iterative training mechanism, training cost is heavy for Bi-LSTM. To effectively improve the training efficiency, we proposed to use Extreme Learning Machine (ELM). The training process of ELM is noniterative, so the network training can converge quickly. As ELM usually shows weak precision performance because of its simple network structure, we introduce the kernel method. In the preprocessing of this framework, we introduce doc2vec for vector representation and multilevel symbolization for program symbolization. Experimental results show that doc2vec vector representation brings faster training and better generalizing performance than word2vec. ELM converges much quickly than Bi-LSTM, and the kernel method can effectively improve the precision of ELM while ensuring training efficiency.
Gaigai Tang, Lin Yang 0031, Shuangyin Ren, Lianxiao Meng
Secur. Commun. Networks1
2021 An Approach of Linear Regression-Based UAV GPS Spoofing Detection
abstract
A prominent security threat to unmanned aerial vehicle (UAV) is to capture it by GPS spoofing, in which the attacker manipulates the GPS signal of the UAV to capture it. This paper introduces an anti‐spoofing model to mitigate the impact of GPS spoofing attack on UAV mission security. In this model, linear regression (LR) is used to predict and model the optimal route of UAV to its destination. On this basis, a countermeasure mechanism is proposed to reduce the impact of GPS spoofing attack. Confrontation is based on the progressive detection mechanism of the model. In order to better ensure the flight security of UAV, the model provides more than one detection scheme for spoofing signal to improve the sensitivity of UAV to deception signal detection. For better proving the proposed LR anti‐spoofing model, a dynamic Stackelberg game is formulated to simulate the interaction between GPS spoofer and UAV. In particular, for GPS spoofer, it is worth mentioning that for the scenario that the UAV is cheated by GPS spoofing signal in the mission environment of the designated route is simulated in the experiment. In particular, UAV with the LR anti‐spoofing model, as the leader in this game, dynamically adjusts its response strategy according to the deception’s attack strategy when upon detection of GPS spoofer’s attack. The simulation results show that the method can effectively enhance the ability of UAV to resist GPS spoofing without increasing the hardware cost of the UAV and is easy to implement. Furthermore, we also try to use long short‐term memory (LSTM) network in the trajectory prediction module of the model. The experimental results show that the LR anti‐spoofing model proposed is far better than that of LSTM in terms of prediction accuracy.
Lianxiao Meng, Lin Yang 0031, Shuangyin Ren, Gaigai Tang, Long Zhang 0004, Wu Yang 0001
Wirel. Commun. Mob. Comput.4
2020 An Optimization of Deep Sensor Fusion Based on Generalized Intersection over Union
Lianxiao Meng, Lin Yang 0031, Gaigai Tang, Shuangyin Ren, Wu Yang 0001
ICA3PP (2)3
2020 A Comparative Study of Neural Network Techniques for Automatic Software Vulnerability Detection
abstract
Software vulnerabilities are usually caused by design flaws or implementation errors, which could be exploited to cause damage to the security of the system. At present, the most commonly used method for detecting software vulnerabilities is static analysis. Most of the related technologies work based on rules or code similarity (source code level) and rely on manually defined vulnerability features. However, these rules and vulnerability features are difficult to be defined and designed accurately, which makes static analysis face many challenges in practical applications. To alleviate this problem, some researchers have proposed to use neural networks that have the ability of automatic feature extraction to improve the intelligence of detection. However, there are many types of neural networks, and different data preprocessing methods will have a significant impact on model performance. It is a great challenge for engineers and researchers to choose a proper neural network and data preprocessing method for a given problem. To solve this problem, we have conducted extensive experiments to test the performance of the two most typical neural networks (i.e., Bi-LSTM and RVFL) with the two most classical data preprocessing methods (i.e., the vector representation and the program symbolization methods) on software vulnerability detection problems and obtained a series of interesting research conclusions, which can provide valuable guidelines for researchers and engineers. Specifically, we found that 1) the training speed of RVFL is always faster than Bi-LSTM, but the prediction accuracy of Bi-LSTM model is higher than RVFL; 2) using doc2vec for vector representation can make the model have faster training speed and generalization ability than using word2vec; and 3) multi-level symbolization is helpful to improve the precision of neural network models.
Gaigai Tang, Lianxiao Meng, Shuangyin Ren, Qiang Wang 0020, Lin Yang 0031, Weipeng Cao
TASE1