Junnan Yin

dblp:330/5555 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0003-1456-1212ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 BTRFormer: Hierarchical Learning of Encrypted Traffic Using a Masked Autoencoder with Block-Based Traffic Representation
abstract
Encrypted traffic classification (ETC) is essential for ensuring network security and efficient management. Despite advances in deep learning, ETC remains challenging as existing models struggle to learn robust, discriminative representations from content-encrypted, highly imbalanced traffic.To address these challenges, we propose BTRFormer, a novel ETC approach that capitalizes on the inherent properties of encryption algorithms to enhance classification accuracy. At the core of BTRFormer lies a block-based, multi-layer traffic representation that adopts a 4×4 block as the fundamental unit, inspired by the encryption algorithm’s use of 16-byte blocks for encryption operations. This representation preserves the intrinsic structure of encrypted payloads, facilitating the model’s ability to learn deep semantic features. Subsequently, a transformer-based model is employed to learn from the multi-layer representation, capturing intra-block, inter-block, and inter-packet dependencies through block-wise attention mechanisms. Finally, BTRFormer leverages a pre-training phase on large-scale unlabeled data, followed by fine-tuning with a minimal amount of labeled samples to improve generalization and adaptability. Experimental results show that BTRFormer significantly outperforms SOTA methods on six real-world datasets, highlighting its effectiveness in encrypted traffic classification and secure network management.
Junnan Yin, Lei Cui 0003, Zhiyu Hao, Peng Liu 0044, Xiao-chun Yun
ICNP1
2024 Revisiting Open DNS Resolver Vulnerabilities to Reflection-Based DDoS Threats
abstract
DNS, as a vital component of the Internet, is frequently exploited for malicious activities. Millions of open DNS resolvers are exposed with public access, posing significant risks. Amplification vulnerability in UDP-based DNS protocol has been abused by miscreants to launch reflection amplification Distributed Denial of Service (DDoS) attacks. In reflection amplification attacks, forged DNS request packets are continuously sent to open resolvers, triggering amplified attack traffic against the targeted victim. To defend against such attacks, resolvers can take measures to reject anomalous requests and limit the size of responses. Measures such as source address verification and response rate limiting prove effective in mitigating the risk of resolver exploitation. However, implementing these measures requires software and hardware updates or configuration changes, potentially incurring additional costs. Currently, it remains unclear how many open resolvers are adequately protected and how many still pose the potential for exploitation. In this paper, we conducted a thorough measurement on open resolvers about the actual potential of abuse. Our measurement results indicated that 14.9% of open resolvers are susceptible to exploitation for reflection-based DDoS attacks and thousands of resolvers are still exposed to reflection amplification attacks with no mitigation measure.
Kehong Liu, Junnan Yin, Letian Du, Tianning Zang
CSCWD3
2024 APIBeh: Learning Behavior Inclination of APIs for Malware Classification
abstract
Malware classification involves categorizing mal-ware samples based on their characteristics. While deep learning techniques applied to malware execution traces, mainly API calls, have shown potential in this field, they still perform poorly. This is primarily because they treat all APIs equally and train classifiers directly on native APIs, which inadequately capture the under-lying family-related semantics. In this paper, we first investigate the behaviors of multiple malware families and observe that different families exhibit divergent behaviors, with each family consistently favoring certain behaviors over time. Motivated by this, we propose APIBeh, a new embedding method designed to enhance malware classification. APIBeh first utilizes Benignity Degree Algorithm to identify and exclude insignificant, likely benign APIs from sequences. Then, it introduces the concept of Behavior Inclination, which quantifies the association between an API and malicious behaviors, facilitating high-level behavior encoding for each API. This Behavior Inclination embedding is then concatenated with raw embedding to represent an API, and fed into a DL model for classifier training. Experimental results show that APIBeh outperforms existing embedding methods in classification performance, e.g., 3.18% boost in weighted f1-score over a recent study using word2vec. In addition, it offers robustness to concept drift and adversarial attacks.
Lei Cui 0003, Yiran Zhu, Junnan Yin, Zhiyu Hao, Wei Wang 0428, Peng Liu 0044, Xiao-chun Yun
ISSRE3
2024 API2Vec++: Boosting API Sequence Representation for Malware Detection and Classification
abstract
Analyzing malware based on API call sequences is an effective approach, as these sequences reflect the dynamic execution behavior of malware. Recent advancements in deep learning have facilitated the application of these techniques to mine valuable information from API call sequences. However, these methods typically operate on raw sequences and may not effectively capture crucial information, especially in the case of multi-process malware, due to theAPI call interleaving problem. Furthermore, they often fail to capture contextual behaviors within or across processes, which is particularly important for identifying and classifying malicious activities. Motivated by this, we present API2Vec++, a graph-based API embedding method for malware detection and classification. First, we construct a graph model to represent the raw sequence. Specifically, we design the Temporal Process Graph (TPG) to model inter-process behaviors and the Temporal API Property Graph (TAPG) to model intra-process behaviors. Compared to our previous graph model, the TAPG model exposes operations with associated behaviors within the process through node properties and thus enhances detection and classification abilities. Using these graphs, we develop a heuristic random walk algorithm to generate numerous paths that can capture fine-grained malicious familial behavior. By pre-training these paths using the BERT model, we generate embeddings of paths and APIs, which can then be used for malware detection and classification. Experiments on a real-world malware dataset demonstrate that API2Vec++ outperforms state-of-the-art embedding methods and detection/classification methods in both accuracy and robustness, particularly for multi-process malware.
Lei Cui 0003, Junnan Yin, Jiancong Cui, Yuede Ji, Peng Liu 0044, Zhiyu Hao, Xiao-chun Yun
IEEE Trans. Software Eng.2
2023 SynCPFL: Synthetic Distribution Aware Clustered Framework for Personalized Federated Learning
abstract
Federated Learning (FL) is a promising machine learning paradigm for collaborative training on cross-soils in a privacy-protected manner. However, the existence of non-IID data causes problems such as performance degradation and thus becomes one of the key challenges in FL recently. To address this problem, we propose a clustered personalized federated learning method named as SynCPFL. SynCPFL groups clients sharing with the similar data distribution together, thereby facilitating collaboration and producing a better-personalized model for each client. In contrast to existing clustered federated learning methods, SynCPFL does not require multiple rounds of interaction between clients and server, so that the communication overhead is reduced a lot, thereby saving resources of clients. We evaluate SynCPFL on benchmark datasets, the experimental results demonstrate that SynCPFL outperforms existing methods.
Junnan Yin, Yuyan Sun, Lei Cui 0003, Zhengyang Ai, Hongsong Zhu
CSCWD1