Linshu Ouyang

dblp:192/1117 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
4since 2021 · last 2022
0000-0001-8948-9436ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-authorSecurity and privacy · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2022 BinMLM: Binary Authorship Verification with Flow-aware Mixture-of-Shared Language Model
abstract
Binary authorship analysis is a significant problem in many software engineering applications. In this paper, we formulate a binary authorship verification task to accurately reflect the real-world working process of software forensic experts. It aims to determine whether an anonymous binary is developed by a specific programmer with a small set of support samples, and the actual developer may not belong to the known candidate set but from the wild. We propose an effective binary authorship verification framework, BinMLM. BinMLM trains the RNN language model on consecutive opcode traces extracted from the control-flow-graph (CFG) to characterize the candidate developers' programming styles. We build a mixture-of-shared architecture with multiple shared encoders and author-specific gate layers, which can learn the developers' combination preferences of universal programming patterns and alleviate the problem of low training resources. Through an optimization pipeline of external pre-training, joint training, and fine-tuning, our framework can eliminate additional noise and accurately distill developers' unique styles. Extensive experiments show that BinMLM achieves promising results on Google Code Jam (GCJ) and Codeforces datasets with different numbers of programmers and supporting samples. It significantly outperforms the baselines built on the state-of-the-art feature set (4.73% to 19.46% improvement) and remains robust in multi-author collaboration scenarios. Furthermore, Bin-MLM can perform organization-level verification on a real-world APT malware dataset, which can provide valuable auxiliary information for exploring the group behind the APT attack.
Qige Song, Yongzheng Zhang 0002, Linshu Ouyang
SANER3
2021 Incremental Learning for Mobile Encrypted Traffic Classification
abstract
With the rising popularity of mobile networks and applications, network traffic classification has gradually become essential to mobile network management and cyberspace security. Existing state-of-the-art methods have achieved high accuracy in the closed-world mobile encrypted traffic classification, where the classifier only needs to process the classes seen in the training. When we update the dataset with new mobile applications, these methods must retrain a new classifier from scratch to learn the knowledge of all applications because directly fine-tuning the existing classifier would lead to the catastrophic forgetting problem. Thus, it is challenging to incrementally add new applications to the classification system while preserving the learned knowledge of the existing classifier. To tackle this issue, we propose an incremental learning framework based on the one vs rest (OvR) strategy and neural network classifiers. Moreover, we adopt a sample selection algorithm to balance the conflict between the growing training effort caused by new applications and the high classification accuracy. The experimental results demonstrate that our proposed framework achieves incremental learning with high classification accuracy like the closed-world method, and the selection algorithm significantly reduces training efforts to meet the dataset scale control and classification accuracy requirement in the lifetime incremental learning.
Tianning Zang, Yongzheng Zhang 0002, Yuan Zhou 0008, Linshu Ouyang
ICC5
2021 Phishing Web Page Detection with Semi-Supervised Deep Anomaly Detection
Linshu Ouyang, Yongzheng Zhang 0002
SecureComm (2)1
2021 Phishing Web Page Detection with HTML-Level Graph Neural Network
abstract
Phishing web page is one of the most serious threats to the users of the Internet. Traditional phishing web page detection methods rely on manually designed features. Recently, deep learning-based methods using HTML as input have achieved significant detection performance improvement. They usually treat HTML codes as sequences of characters and utilize Convolutional Neural Network (CNN) or Recurrent Neural Network (RNN) for classification. However, CNN and RNN typically can only extract local features in the HTML code sequences while failing to model the long-range semantics that is crucial for phishing detection. In this paper, we propose a novel Graph Neural Network (GNN) based phishing web page detection method that can effectively utilize the inherent structural information of HTML to capture the long-range semantics. We first naturally represent an HTML as a graph according to its Document Object Model (DOM) and utilize RNN to extract the local features of node attributes. Then we adopt GNN to model the long-range relations between nodes based on these local features and the graph structure. Our proposed model combines the advantage of RNN and GNN to better understand the intention of HTML codes. Extensive experiments on a real-world dataset demonstrate that the accuracy of our method outperforms other state-of-the-art methods by a large margin.
Linshu Ouyang, Yongzheng Zhang 0002
TrustCom1
2020 Gated POS-Level Language Model for Authorship Verification
abstract
Authorship verification is an important problem that has many applications. The state-of-the-art deep authorship verification methods typically leverage character-level language models to encode author-specific writing styles. However, they often fail to capture syntactic level patterns, leading to sub-optimal accuracy in cross-topic scenarios. Also, due to imperfect cross-author parameter sharing, it's difficult for them to distinguish author-specific writing style from common patterns, leading to data-inefficient learning. This paper introduces a novel POS-level (Part of Speech) gated RNN based language model to effectively learn the author-specific syntactic styles. The author-agnostic syntactic information obtained from the POS tagger pre-trained on large external datasets greatly reduces the number of effective parameters of our model, enabling the model to learn accurate author-specific syntactic styles with limited training data. We also utilize a gated architecture to learn the common syntactic writing styles with a small set of shared parameters and let the author-specific parameters focus on each author's special syntactic styles. Extensive experimental results show that our method achieves significantly better accuracy than state-of-the-art competing methods, especially in cross-topic scenarios (over 5\% in terms of AUC-ROC).
Linshu Ouyang, Yongzheng Zhang 0002, Yipeng Wang 0001
IJCAI1
2020 Unified Graph Embedding-Based Anomalous Edge Detection
abstract
Detecting anomalous edges in graph-structured data plays an important role in many fields such as finance, social network, and network security. Recently, graph embedding based anomaly detection methods show promising results. These methods typically encode graph structure information into vector representation and apply general anomaly detection methods. However, since the parameters in these two parts are learned separately with different objectives, the learned representation may contain some information irrelevant to the task. It would be ideal if we can combine representation learning and anomaly detection into one objective function to force the model to focus on learning task relevant patterns. In this paper, we propose a novel end-to-end neural network architecture that can accurately estimate the probability distribution of edges in the graph based on its local structure. An edge has a high chance to be considered an anomaly if the probability of its existence is low. Extensive experiments on several public datasets at different scales show that the accuracy and scalability of our method outperform other methods by a large margin.
Linshu Ouyang, Yongzheng Zhang 0002, Yipeng Wang 0001
IJCNN1