VLDB 2026 Research / reviewers in the wild / expert
Lin Yang 0031
dblp:20/2970-31
· DBLP profile ↗
21ranked-venue papers
0as first author
18since 2021 · last 2026
0000-0002-6956-8177ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 6 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Security and privacy · 6 · 6 since 2021Computer networks · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fusion Is Not A Simple Ensemble! Towards The Evolving Views in Insider Threat DetectionabstractInsider threat detection (ITD) is notoriously difficult: malicious actions are rare, context-dependent, and deliberately hidden within massive volumes of legitimate user behavior. Existing ITD methods rely on single- or fused-view models, which lack extensibility and therefore fail to leverage the supervisory signals from newly introduced complementary views. While ensembling is a natural next step, its direct application to ITD confronts three core obstacles: scalability bottlenecks from independently trained sub - models, semantic misalignment across heterogeneous feature spaces, and view imbalance, where strong views overshadow weaker yet informative ones. In this work, we propose Insight-LLM, the first extensible multi-view fusion framework tailored for ITD. Insight-LLM encodes each view with frozen pre-trained backbones and aligns heterogeneous representations into a unified semantic space via a lightweight ViewAdapter, enabling coherent cross-view reasoning without incurring additional training overhead. A context-adaptive fusion module dynamically re-weights views to emphasize subtle yet semantically consistent threat signals, and the fused representation is integrated with task prompts for lightweight LLM fine-tuning. Experiments on CERT datasets show that Insight-LLM improves F1 by up to 4.8% and reduces false positives by 61%, while decreasing training time per newly added view by up to 83.2% compared with the simple Ensemble method. Chengyu Song, Lin Yang 0031, Jianming Zheng, Jingjing Zhang 0005, Hongyu Kuang, Jinzhi Liao, Mengchun Zhao |
WWW | 2 |
| 2026 | Federated Learning With Drift Correction and Convergence Acceleration
Jihao Yang, Lin Yang 0031, Wen Jiang 0002, Laisen Nie |
IEEE Internet Things J. | 2 |
| 2025 | Parse-LLM: A Prior-Free LLM Parser for Unknown System LogsabstractLog parsing extracts structured information from unstructured logs and serves as a fundamental pre-processing step for various log-based analytics and monitoring tasks. Recent advances have leveraged Large Language Models (LLMs) to handle log format complexities and enhance parsing performance. However, these methods heavily rely on labeled data, which is often scarce in rapidly evolving industrial systems, limiting their applicability in real-world scenarios. Moreover, the sheer volume of logs results in slow parsing and high computational costs, further hindering the deployment of LLM-based log parsing systems. To address these issues, we propose Parse-LLM, an unsupervised end-to-end log parsing framework based on LLMs Specifically, we first developed a Log Decomposer Agent that leverages Chain-of-Thought (CoT) reasoning and callable tools, enabling the LLM to autonomously separate log headers from content. Next, we introduce the Hybrid Log Partition module, which segments logs by balancing commonalities and differences. Finally, we developed a novel Variation-aware Log Parsing module that allows the LLM to harness additional supervisory signals through comparative analysis of similar logs. Comprehensive experiments conducted on large-scale public datasets show that Parse-LLM outperforms state-of-the-art log parsers in an unsupervised setting, offering an effective and scalable solution for the practical application of unsupervised log parsing. Chengyu Song, Lin Yang 0031, Jianming Zheng, Jinzhi Liao, Linru Ma |
CIKM | 2 |
| 2025 | Dynamic Feedback-Based Cost-Sensitive Learning for Imbalanced Intrusion DetectionabstractIn network intrusion detection systems, the persistent challenge of class imbalance critically undermines detection efficacy for minority attacks. To address this, we propose a Dynamic Feedback Based Cost Sensitive Learning (DFCSL) method. This method employs a two-stage detection framework to achieve hierarchical classification between normal traffic and anomalous attacks. Additionally, we introduce a dynamic weight update mechanism based on recall feedback, which adaptively enhances the model’s focus on hard-to-recognize categories, mitigating the limitations of traditional static weight settings. Moreover, a feature selection strategy based on Shapley values is designed to identify high-discriminative features, reduce input dimensionality, and lower the computational cost of the model. Experiments conducted on three datasets—CICIDS2017, CSE-CIC-IDS2018, and CIDDS-001 that the proposed method not only improves the detection performance of minority class attacks but also maintains strong recognition capabilities for majority class attacks. Xiaohang Ma, Lin Yang 0031 |
SMC | 2 |
| 2025 | Intrusion Detection for Internet of Things: An Anchor Graph Clustering ApproachabstractIntrusion detection systems are a crucial technique for securing the Internet of Things (IoT) from malicious attacks. Additionally, due to the continuous emergence of new vulnerabilities and unknown attack types, only a small number of attack samples in the IoT environments can be captured for analysis. In this work, we introduce an anchor graph clustering (AGC) method for intrusion detection to address the challenge of limited labeled samples in the IoT environments. AGC initially transforms the raw data into the embedding space to obtain more representative anchors. Then, AGC unifies anchor graph construction, anchor graph learning, and graph clustering into a unified framework, solving the resulting optimization problem through an iterative solution algorithm. Finally, AGC leverages the powerful analytical capabilities of graph learning to achieve fine-grained classification of low-quality labels. Experimental results on both real and synthetic datasets confirm that AGC can identify intrusions with high precision, while also being time-efficient in detection. Long Zhang 0004, Lin Yang 0031, Linru Ma, Zhoumin Lu, Wen Jiang 0002 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | VulCausal: Robust Vulnerability Detection Using Neural Network Models from a Causal Perspective
Hongyu Kuang, Jingjing Zhang 0005, Long Zhang 0004, Lin Yang 0031 |
KSEM (3) | 6 |
| 2024 | Insider Threat Defense Strategies: Survey and Knowledge Integration
Chengyu Song, Jingjing Zhang 0005, Linru Ma, Xinxin Hu, Jianming Zheng, Lin Yang 0031 |
KSEM (5) | 6 |
| 2024 | MRC-VulLoc: Software source code vulnerability localization based on multi-choice reading comprehension
Gaigai Tang, Lin Yang 0031, Long Zhang 0004, Hongyu Kuang |
Comput. Secur. | 2 |
| 2024 | Intrusion Detection for Unmanned Aerial Vehicles Security: A Tiny Machine Learning ModelabstractUnmanned Aerial Vehicles (UAVs) are vulnerable to network attacks. Designing an effective intrusion detection system (IDS) for UAVs is crucial. However, UAVs have limited computing resources and need to deal with massive amounts of network data, which further increases the difficulty of detection. Moreover, most existing IDSs have large parameters. In this study, we develop a tiny machine learning-based IDS to solve the above issue. We first establish an improved fuzzy rough set (FRS) model based on adaptive neighborhoods. Then, using the proposed FRS model, we employ a feature selection (FS) method to select optimal features and reduce overall computational cost of the IDS. Furthermore, we proposed a tiny intrusion detection model that attains high-precision detection via shallow deep learning. Additionally, the proposed method can address intrusion detection problems in scenarios with partial data missing. According to the evaluations, the proposed method can effectively address intrusion detection in UAVs. Lin Yang 0031, Long Zhang 0004, Laisen Nie |
IEEE Internet Things J. | 2 |
| 2023 | Leveraging User-Defined Identifiers for Counterfactual Data Generation in Source Code Vulnerability DetectionabstractSoftware vulnerability detection is a critical aspect of ensuring the security and reliability of software systems. However, traditional vulnerability detection approaches often have limitations due to the scarcity and need for more diversity in labeled data. This research introduces a novel approach to overcome these challenges by utilizing user-defined identifiers in the source code to generate counterfactual training data. User-defined identffiers, such as variable and function names, contain essential information about the intentions and logic of the program. By perturbing these identifiers while maintaining the syntactic and semantic structure of the code, we create a diverse set of counterfactual examples that simulate potential vulnerabilities. When combined with existing labeled data, these counterfactual examples enrich the training process for vulnerability detection models. To evaluate the effectiveness of our approach, we conduct experiments on various datasets, achieving state-of-the-art performance on the VulDeePecker and Draper datasets. Our approach also outperforms models that utilize the same pre-trained language model in terms of accuracy. Hongyu Kuang, Long Zhang 0004, Gaigai Tang, Lin Yang 0031 |
SCAM | 5 |
| 2023 | Robust Anomaly-Based Insider Threat Detection Using Graph Neural NetworkabstractMisuse or malicious access to critical assets of information systems by insiders usually causes significant loss to organizations. The issue of insider threat detection for information systems has received many researchers’ attention in both security and data mining fields, and a lot of related research results were presented. However, there are still many challenges in capturing the behavior difference between malicious insiders and normal users accurately, such as lack of labeled insider threats, the subtle and adaptive nature of insider threats, complexity, heterogeneity, sparsity of the underlying data, etc. To detect insider threats with large and complex audit data, a Multi-Edge Weight Relational Graph Neural Network method (MEWRGNN) for robust anomaly detection is proposed in this paper. Unlike most existing approaches, the MEWRGNN adopts several graph neural networks to capture the contextual relationship of user behaviors over a period of time, which is a critical factor for achieving accurate anomaly identification. The MEWRGNN achieves a certain degree of interpretability through ranking the contribution of different edge-representation features. Evaluation experimental results demonstrate that the MEWRGNN can learn a model from limited sample data sets, and achieve quick and accurate insider threat detection performance. In addition, other feature ranking results allow providing security analysts with understandable insights for investigating the detected insider threats. Junchao Xiao, Lin Yang 0031, Fuli Zhong, Xiaolei Wang 0003 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2022 | MADDC: Multi-Scale Anomaly Detection, Diagnosis and Correction for Discrete Event LogsabstractAnomaly detection for discrete event logs can provide critical information for building secure and reliable systems in various application domains, such as large scale data centers, autonomous driving, and intrusion detection. However, the task is very challenging due to the lack of a clear understanding and definition of anomaly in the specific problem space, and the log data is often highly complex with temporal correlation. Existing deep learning based methods mostly suffer from such issues as overfitting, uncertainty or low interpretability; consequently, the detection results may be inaccurate, with little information to help security analysts diagnose the reported anomalies with high confidence. To tackle this challenge, in this research, we propose a general framework named MADDC, which aims to (1) accurately perform Multi-scale Anomaly Detection, Diagnosis and Correction for discrete event logs, and (2) help analysts further mitigate anomalies based on diagnosis results. Specifically, we first design a new anomaly critic for LSTM variational autoencoder based model to alleviate overfitting and reduce false negatives during anomaly detection. As one of our main contributions, we then introduce process mining technique to build process-centric workflow models in an unsupervised manner, which forms the ‘normal’ context of an event sequence and help perform accurate and consistent anomaly diagnosis through global sequence alignment. Experiments on publicly available datasets show that MADDC not only outperformed several representative methods in terms of detection accuracy, but also could improve the visibility to abnormal deviations from normal execution, hence helping security analysts understand anomalies and make further corrections. Xiaolei Wang 0003, Lin Yang 0031, Linru Ma, Junchao Xiao, Jiyuan Liu 0003, Yuexiang Yang |
ACSAC | 2 |
| 2022 | Automated Reliability Analysis of Redundancy Architectures Using Statistical Model Checking
Hongbin He, Hongyu Kuang, Lin Yang 0031, Qiang Wang 0020, Weipeng Cao |
KSEM (3) | 3 |
| 2021 | Interpretation of Learning-Based Automatic Source Code Vulnerability Detection Model Using LIME
Gaigai Tang, Long Zhang 0004, Lianxiao Meng, Weipeng Cao, Meikang Qiu, Shuangyin Ren, Lin Yang 0031 |
KSEM | 8 |
| 2021 | Image-Based Insider Threat Detection via Geometric TransformationabstractInsider threat detection has been a challenging task over decades; existing approaches generally employ the traditional generative unsupervised learning methods to produce normal user behavior model and detect significant deviations as anomalies. However, such approaches are insufficient in precision and computational complexity. In this paper, we propose a novel insider threat detection method, Image-based Insider Threat Detector via Geometric Transformation (IGT), which converts the unsupervised anomaly detection into supervised image classification task, and therefore the performance can be boosted via computer vision techniques. To illustrate, our IGT uses a novel image-based feature representation of user behavior by transforming audit logs into grayscale images. By applying multiple geometric transformations on these behavior grayscale images, IGT constructs a self-labelled dataset and then trains a behavior classifier to detect anomaly in a self-supervised manner. The motivation behind our proposed method is that images converted from normal behavior data may contain unique latent features which remain unchanged after geometric transformation, while malicious ones cannot. Experimental results on CERT dataset show that IGT outperforms the classical autoencoder-based unsupervised insider threat detection approaches, and improves the instance and user based Area under the Receiver Operating Characteristic Curve (AUROC) by 4% and 2%, respectively. Lin Yang 0031, Xiaolei Wang 0003, Linru Ma, Junchao Xiao |
Secur. Commun. Networks | 2 |
| 2021 | An Automatic Source Code Vulnerability Detection Approach Based on KELMabstractTraditional vulnerability detection mostly ran on rules or source code similarity with manually defined vulnerability features. In fact, these vulnerability rules or features are difficult to be defined accurately, which usually cost much expert labor and perform weakly in practical applications. To mitigate this issue, researchers introduced neural networks to automatically extract features to improve the intelligence of vulnerability detection. Bidirectional Long Short-term Memory (Bi-LSTM) network has proved a success for software vulnerability detection. However, due to complex context information processing and iterative training mechanism, training cost is heavy for Bi-LSTM. To effectively improve the training efficiency, we proposed to use Extreme Learning Machine (ELM). The training process of ELM is noniterative, so the network training can converge quickly. As ELM usually shows weak precision performance because of its simple network structure, we introduce the kernel method. In the preprocessing of this framework, we introduce doc2vec for vector representation and multilevel symbolization for program symbolization. Experimental results show that doc2vec vector representation brings faster training and better generalizing performance than word2vec. ELM converges much quickly than Bi-LSTM, and the kernel method can effectively improve the precision of ELM while ensuring training efficiency. Gaigai Tang, Lin Yang 0031, Shuangyin Ren, Lianxiao Meng |
Secur. Commun. Networks | 2 |
| 2021 | A Systematic Approach to Formal Analysis of QUIC Handshake Protocol Using Symbolic Model CheckingabstractAs a newly proposed secure transport protocol, QUIC aims to improve the transport performance of HTTPS traffic and enable rapid deployment and evolution of transport mechanisms. QUIC is currently in the IETF standardization process and will potentially carry a significant portion of Internet traffic in the emerging future. An important safety goal of QUIC protocol is to provide effective data service for users. To aim this safety requirement, we propose a formal analysis method to analyze the safety of QUIC handshake protocol by using model checker SPIN and cryptographic protocol verifier ProVerif. Our analysis shows the counterexamples to safety properties, which reveal a design flaw in the current protocol specification. To this end, we also propose and verify a possible fix that is able to mitigate these flaws. Jingjing Zhang 0005, Xianming Gao, Lin Yang 0031, Qiang Wang 0020 |
Secur. Commun. Networks | 3 |
| 2021 | An Approach of Linear Regression-Based UAV GPS Spoofing DetectionabstractA prominent security threat to unmanned aerial vehicle (UAV) is to capture it by GPS spoofing, in which the attacker manipulates the GPS signal of the UAV to capture it. This paper introduces an anti‐spoofing model to mitigate the impact of GPS spoofing attack on UAV mission security. In this model, linear regression (LR) is used to predict and model the optimal route of UAV to its destination. On this basis, a countermeasure mechanism is proposed to reduce the impact of GPS spoofing attack. Confrontation is based on the progressive detection mechanism of the model. In order to better ensure the flight security of UAV, the model provides more than one detection scheme for spoofing signal to improve the sensitivity of UAV to deception signal detection. For better proving the proposed LR anti‐spoofing model, a dynamic Stackelberg game is formulated to simulate the interaction between GPS spoofer and UAV. In particular, for GPS spoofer, it is worth mentioning that for the scenario that the UAV is cheated by GPS spoofing signal in the mission environment of the designated route is simulated in the experiment. In particular, UAV with the LR anti‐spoofing model, as the leader in this game, dynamically adjusts its response strategy according to the deception’s attack strategy when upon detection of GPS spoofer’s attack. The simulation results show that the method can effectively enhance the ability of UAV to resist GPS spoofing without increasing the hardware cost of the UAV and is easy to implement. Furthermore, we also try to use long short‐term memory (LSTM) network in the trajectory prediction module of the model. The experimental results show that the LR anti‐spoofing model proposed is far better than that of LSTM in terms of prediction accuracy. Lianxiao Meng, Lin Yang 0031, Shuangyin Ren, Gaigai Tang, Long Zhang 0004, Wu Yang 0001 |
Wirel. Commun. Mob. Comput. | 2 |
| 2020 | An Optimization of Deep Sensor Fusion Based on Generalized Intersection over Union
Lianxiao Meng, Lin Yang 0031, Gaigai Tang, Shuangyin Ren, Wu Yang 0001 |
ICA3PP (2) | 2 |
| 2020 | A Comparative Study of Neural Network Techniques for Automatic Software Vulnerability DetectionabstractSoftware vulnerabilities are usually caused by design flaws or implementation errors, which could be exploited to cause damage to the security of the system. At present, the most commonly used method for detecting software vulnerabilities is static analysis. Most of the related technologies work based on rules or code similarity (source code level) and rely on manually defined vulnerability features. However, these rules and vulnerability features are difficult to be defined and designed accurately, which makes static analysis face many challenges in practical applications. To alleviate this problem, some researchers have proposed to use neural networks that have the ability of automatic feature extraction to improve the intelligence of detection. However, there are many types of neural networks, and different data preprocessing methods will have a significant impact on model performance. It is a great challenge for engineers and researchers to choose a proper neural network and data preprocessing method for a given problem. To solve this problem, we have conducted extensive experiments to test the performance of the two most typical neural networks (i.e., Bi-LSTM and RVFL) with the two most classical data preprocessing methods (i.e., the vector representation and the program symbolization methods) on software vulnerability detection problems and obtained a series of interesting research conclusions, which can provide valuable guidelines for researchers and engineers. Specifically, we found that 1) the training speed of RVFL is always faster than Bi-LSTM, but the prediction accuracy of Bi-LSTM model is higher than RVFL; 2) using doc2vec for vector representation can make the model have faster training speed and generalization ability than using word2vec; and 3) multi-level symbolization is helpful to improve the precision of neural network models. Gaigai Tang, Lianxiao Meng, Shuangyin Ren, Qiang Wang 0020, Lin Yang 0031, Weipeng Cao |
TASE | 6 |
| 2018 | An improvement to generalized regret based decision making method considering unreasonable alternativesabstractRegret decision theory is a classic theory for decision problem. Recently, Yager proposed a generalized regret based decision-making method, which calculates the effective regret associated with an alternative by aggregating this alternative's all regrets across all the possible states of nature. The generalized regret based decision-making method that can be applied in many fields is understandable and effective. However, as Yager pointed, an issue limits the application of this method, that is, the generalized regret based decision-making method is lack of indifference to irrelevant alternatives. In this paper, we analyze the cause of this issue, that is, unreasonable alternatives may change other alternatives' regrets by changing the maximal payoff under the occurrence of a state of nature. Furthermore, a new method based on original model is proposed to reduce the impact of unreasonable alternatives according to a parameter called impact factor defined to measure an alternative's quality. Finally, several numerical examples are illustrated to show this new method's effectiveness. Xinyang Deng, Lin Yang 0031, Wen Jiang 0002 |
Int. J. Intell. Syst. | 3 |