Luming Yang

dblp:38/2888 · DBLP profile ↗
← Back
20ranked-venue papers
12as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 7 · 4 first-author · 6 since 2021Computer networks · 5 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Multivariate Gaussian Representation Learning for Medical Action Evaluation
abstract
Fine-grained action evaluation in medical vision faces unique challenges due to the unavailability of comprehensive datasets, stringent precision requirements, and insufficient spatiotemporal dynamic modeling of very rapid actions. To support development and evaluation, we introduce CPREval-6k, a multi-view, multi-label medical action benchmark containing 6,372 expert-annotated videos with 22 clinical labels. Using this dataset, we present GaussMedAct, a multivariate Gaussian encoding framework, to advance medical motion analysis through adaptive spatiotemporal representation learning. Multivariate Gaussian Representation projects the joint motions to a temporally scaled multi-dimensional space, and decomposes actions into adaptive 3D Gaussians that serve as tokens. These tokens preserve motion semantics through anisotropic covariance modeling while maintaining robustness to spatiotemporal noise. Hybrid Spatial Encoding, employing a Cartesian and Vector dual-stream strategy, effectively utilizes skeletal information in the form of joint and bone features. The proposed method achieves 92.1% Top-1 accuracy with real-time inference on the benchmark, outperforming the baseline by +5.9% accuracy with only 10% FLOPs. Cross-dataset experiments confirm the superiority of our method in robustness.
Luming Yang, Haoxian Liu, Siqing Li
AAAI1
2025 Bidirectional Time-Frequency Pyramid Network for Enhanced Robust EEG Classification
abstract
Existing EEG recognition models suffer from poor cross-paradigm generalization due to dataset-specific constraints and individual variability. To overcome these limitations, we propose Bite (Bidirectional Time-Freq Pyramid Network), an end-to-end unified architecture featuring robust multistream synergy, pyramid time-frequency attention (PTFA), and bidirectional adaptive convolutions. The framework uniquely integrates: 1) Aligned time-frequency streams maintaining temporal synchronization with STFT for bidirectional modeling, 2) PTFA-based multiscale feature enhancement amplifying critical neural patterns, 3) BiTCN with learnable fusion capturing forward/backward neural dynamics. Demonstrating enhanced robustness, Bite achieves state-of-the-art performance across four divergent paradigms (BCICIV-2A/2B, HGD, SD-SSVEP), excelling in both withinsubject accuracy and cross-subject generalization. As a unified architecture, it combines robust performance across both MI and SSVEP tasks with exceptional computational efficiency. Our work validates that paradigm-aligned spectral-temporal processing is essential for reliable BCI systems. Just as its name suggests, Bite “takes a bite out of EEG.” We publicly release the model as BiteEEG, and the source code is available at https://github.com/cindy-hong/BiteEEG.
Jiahui Hong, Siqing Li, Muqing Jian, Luming Yang
BIBM4
2025 MM4flow: A Pre-trained Multi-modal Model for Versatile Network Traffic Analysis
abstract
Network traffic analysis is a critical research area, playing an essential role in enhancing network security and ensuring high-quality network services. Existing methods, which primarily rely on a single modality, face two significant limitations. First, while existing approaches may achieve strong performance in specific tasks, they often lack sufficient adaptability for diverse tasks. Second, existing pre-trained models are only trained with GB-scale traffic, with which increases the risk of over-fitting and limiting the models' overall performance. To address these challenges, we propose MM4flow, a pre-trained multi-modal model designed for versatile network traffic analysis. We divide network flows into two modalities: raw byte streams and transmission patterns, which encapsulate the content and behavior information, respectively. MM4flow is composed of two key stages: uni-modal pre-training and multi-modal fine-tuning. We develop an efficient data collection scheme enabling TB-scale traffic pre-training. Leveraging a real-world traffic that exceeds 70 TB, MM4flow conducts uni-modal pre-training on each modality with a modified BERT architecture tailored for network flows. For specific downstream tasks, we introduce a modal fusion module based on cross-attention mechanisms. The fusion module facilitates effective integration of multi-modal information, enabling MM4flow to fully utilize both content and behavior cues during fine-tuning with minimal labeled dataset. We evaluate MM4flow on six public datasets covering six various tasks. Extensive experiments demonstrate that MM4flow achieves superior accuracy than baselines. Especially, compared to existing pre-trained models, MM4flow achieves an 84% improvement in accuracy for website identification under encrypted tunnels. Moreover, the pre-trained MM4flow significantly reduces the reliance on high-quality labeled training data for downstream tasks.
Luming Yang, Lin Liu 0018, Junjie Huang 0001, Zhuotao Liu, Shiyu Liang, Shaojing Fu
CCS1
2025 SemiAF: Semi-Supervised App Fingerprinting on Unknown Traffic via Graph Neural Network
abstract
Application Fingerprinting (AF) enables the identification of applications via traffic analysis, aiding network administrators in comprehending user behavior. However, a large volume of unknown traffic in real network environments poses significant challenges for AF methods in both differentiating unknown traffic and characterizing unknown apps. To address these challenges, we introduce a Semi-Supervised App Fingerprinting (SemiAF) deep learning framework. Specifically, to tackle the unknown traffic differentiating challenge, a semi-supervised contrastive learning method is employed to differentiate and cluster unknown applications. To characterize the features of unknown apps, we present a novel In-Flow Burst Interaction (IFBI) graph where each node represents a fine-grained action. By decomposing the unknown app traffic into combinations of these fine-grained actions, a deeper understanding of the network patterns can be achieved. Furthermore, we introduce an explainable neural network framework, revealing the network traffic interactions and inherent relationships. Extensive experimental results in three scenarios demonstrate that SemiAF outperforms state-of-the-art methods in unknown application recognition.
Xiaodong Lei, Jiangyong Shi, Luming Yang
DSN6
2025 Disrupting explicit encoding paradigms: property-interactive transformers decode T-cell receptor specificity beyond dataset biases
abstract
The human immune response relies on the unique ability of T-cell receptors (TCRs) to specifically bind to peptides, a process essential for immune surveillance and response. Although deep learning methods for prediction of TCR-peptide binding have proliferated, many encoder-based approaches learn dataset biases, greatly overestimating the model results, and ignoring the biochemical mechanisms and spatial properties affecting binding. Through our analysis, we found that interaction pairs generated by cross-mapping the amino acid properties between TCR and peptide implicitly simulate spatial structure, enabling machine learning models to capture information more effectively. Based on this insight, we developed T-cell receptor cross (TCRoss), a transformer-based model for large-scale learning. In addition, we observed that incorporating environmental information into the dataset not only mitigates learning biases but also improves performance. Experiments show that TCRoss consistently outperforms existing models in both observed contexts and de novo peptide scenarios. Wet-lab validation using T-cell activation assays confirmed the model's predictions for nonbinding peptides and provided critical experimental evidence for model assessment. Biophysical validation confirms that high-attention residue pairs correspond to crystallographically observed binding interfaces.
Luming Yang, Haoxian Liu, Alec Calanche, Sohret M. Gokcek, Nicholas Sansoterra, Munir Akkaya, Billur Akkaya, Alper Yilmaz 0001
Briefings Bioinform.1
2025 The Analysis of Encrypted Video Stream Based on Low-Dimensional Embedding Method
abstract
In recent years, encrypted video streaming takes up an increasing proportion of mobile network traffic, with encrypted video streams playing a significant role in illegal video detection. However, there are challenges in performing content analysis of encrypted video streams, including label limitations and complex calculations. In this paper, we proposed a low-dimensional embedding method based on Byte Rate Sequences (BRS), named EVS2vec (Encrypted Video Stream to Vector), to solve these problems effectively. It can represent the content of encrypted video streams with low-dimensional vectors by mapping the indefinite-length sequence into a low-dimensional Euclidean space. EVS2vec can thereby be applied for not only supervised analysis but also unsupervised analysis. Furthermore, using BRS can also save the time overhead on fine-grained network traffic parsing. In order to ensure the content-related distinguishability of the embedding result, inspired by contrastive learning, we designed a network structure based on Recurrent Neural Network (RNN) with self-attention mechanism in EVS2vec and trained it using triplet network. The experiments on a public dataset show that EVS2vec saves storage overhead while containing enough video content information. EVS2vec can achieve a high accuracy of similarity threshold, reaching 96.89%. An 8-dimensional fingerprint for each video is constructed. Moreover, classification and clustering analysis can also be performed with acceptable results.
Luming Yang, Shaojing Fu, Lin Liu 0018, Yuchuan Luo
IEEE Trans. Inf. Forensics Secur.1
2025 Robustness Matters: Pre-Training Can Enhance the Performance of Encrypted Traffic Analysis
abstract
Models with large-scale parameters and pre-training have been leveraged for encrypted traffic analysis. However, existing researches primarily focused on accuracy, often overlooking the role of large-scale pre-trained parameters in enhancing robustness. While machine learning (ML) and deep learning (DL) models trained from scratch can achieve high accuracy, they exhibit limited robustness. When subjected to network noise in real-world, their identification results can fluctuate significantly, which is unacceptable. Unfortunately, current robustness evaluation methods neglect samples diversity and employ unreasonable noise settings. This field still lacks a reasonable quantitative description of models robustness. In this paper, we propose the PA-curve to display the distribution of sample’s correct-decision stability, which can simultaneously reflect the model’s accuracy and robustness. By calculating the area under the PA-curve, called PA-area, we enable the quantitative assessment of robustness for encrypted traffic analysis. Furthermore, we design a pre-trained model based on packet length sequence, and pre-trained it on TB-scale traffic. By fine-tuning on limited labeled training data, it can achieve downstream analysis tasks. We conduct experiments on five encrypted traffic datasets with different tasks. Besides accuracy, we analyzed the robustness of the pre-trained model and existing methods under common network disturbances, including packet loss, retransmission, and disorder. Experimental results demonstrated that, compared to ML-based and DL-based models trained from scratch, the pre-trained model can not only achieve high accuracy, but also exhibit greater resilience to network noise. The source code is available at https://anonymous.4open.science/r/BERT-ps-4630.
Luming Yang, Lin Liu 0018, Junjie Huang 0001, Jiangyong Shi, Shaojing Fu, Jinshu Su
IEEE Trans. Inf. Forensics Secur.1
2025 unFlowS: An Unsupervised Construction Scheme of Flow Spectrum for Network Traffic Detection
abstract
In recent years, the construction of behavior-based analysis models is hindered by issues such as insufficient data, difficulty in labeling, and the complexity of behavior types. In reality, specific cyber threats often require manual analysis of raw network traffic, which is a complex and inefficient process. Flow spectrum can simplify the complex analysis process of raw network flow by mapping it from a high-dimensional space to a one-dimensional spectral space. However, the existing flow spectrum cannot adapt to the open-world scenarios and behavior-based detection for unknown cyber threats. To address these challenges, we propose a new flow spectrum construction scheme, named unFlowS, to effectively represent network flows and assist analysts to understand the behaviors of network traffic. unFlowS-Net, an unsupervised flow-based detection model we designed as the core of our scheme, can transform network flows into spectral lines. It makes unFlowS possible to detect unknown cyber threats. We further build spectral vectors for spectral lines generated by network flow sets, enabling the visualization of network behaviors within a period of time and automatic behavior-based detection. Experimental results demonstrated that unFlowS-Net can achieve better performance than state-of-the-art methods on unsupervised flow-based detection. Based on spectral vectors, not only can it intuitively display the network behavior characteristic of the target host, but also automatically detect suspicious network behaviors.
Luming Yang, Lin Liu 0018, Junjie Huang 0001, Jiangyong Shi, Shaojing Fu, Shize Guo
IEEE Trans. Inf. Forensics Secur.1
2024 ExpMD: an Explainable Framework for Traffic Identification Based on Multi-Domain Features
abstract
Network traffic identification has a significant impact on the QoS (Quality of Service) of network. However, there are challenges in performing network traffic identification, including underutilization of information and weak of interpretability. To address these issues, in this paper, we propose an Explainable framework based on Multi-Domain features for network traffic identification, named ExpMD. The network information is divided into four feature domains: metadata domain, byte-distribution (BD) domain, textual domain, and temporal domain. On these feature domains, we construct four identification models independently, thereby achieving the extraction and utilization of multi-domain information of network flows. Subsequently, efficient identification of network traffic can be achieved through the ensemble of models from four domains. In addition, we employ post-hoc explainable approaches to attribute features across multiple feature domains, thus mitigating the black-box issue of network traffic identification. We run a set of experiments in the identification tasks of application network traffic. The extensive evaluation demonstrates that our framework not only achieves an accuracy of 97. 5%, but also generates high-fidelity, sparse, complete, and stable explanation results for network flows. Furthermore, each sub-model performed at a state-of-the-art effect within their feature domains, respectively.
Luming Yang, Lin Liu 0018, Shaojing Fu
IWQoS1
2023 EVS2vec: A Low-dimensional Embedding Method for Encrypted Video Stream Analysis
abstract
The rise in video streaming has led to an increase in network traffic, with encrypted video streams playing a significant role in illegal video detection. However, there are challenges in performing content analysis of encrypted video streams, including label limitations and complex calculations. In this paper, we proposed a low-dimensional embedding method based on Byte Rate Sequences (BRS), named EVS2vec (Encrypted Video Stream to Vector), to solve these problems effectively. It can represent the content of encrypted video streams with low-dimensional vectors by mapping the indefinite-length sequence into a low-dimensional Euclidean space. EVS2vec can thereby be applied for not only supervised analysis but also unsupervised analysis. Furthermore, using BRS can also save the time overhead on fine-grained network traffic parsing. In order to ensure the content-related distinguishability of the embedding result, inspired by contrastive learning, we designed a network structure based on Recurrent Neural Network (RNN) with self-attention mechanism in EVS2vec and trained it using siamese network. The experiments on a public dataset show that EVS2vec saves storage overhead while containing enough video content information. EVS2vec can achieve a high accuracy of similarity threshold, reaching 96.71%. An 8-dimensional fingerprint for each video is constructed. Moreover, classification and clustering analysis can also be performed with acceptable results.
Luming Yang, Shaojing Fu, Lin Liu 0018, Yuchuan Luo
SECON1
2023 EzBoost: Fast And Secure Vertical Federated Tree Boosting Framework via EzPC
abstract
Federated learning (FL) has emerged as a prominent methodology for collaboratively training machine learning models among multiple participants while alleviating data privacy leakage through data localization. However, recent studies have shown that the transferred intermediate parameters still contain sensitive information that needs to be further protected. More-over, real-world institutions often possess diverse data attributes, necessitating the adoption of Vertical Federated Learning (VFL) for cooperative learning tasks. Existing researches in VFL have proposed some frameworks with privacy-preservation functionality, yet they suffer from high participant overhead or low model accuracy, etc. To address these challenges, in this paper, we propose EzBoost, a fast and secure vertical federated tree boosting framework built upon XGBoost. Specifically, we leverages the efficient Secure Multi-party Computation (MPC) framework, EzPC, to facilitate the design and implementation of EzBoost. By carefully designing our framework with two non-collusive servers for secure two-party computation, EzBoost significantly accelerates the runtime of model training and querying at most 20×, and reduces the participant overheads at most 300×. In addition, we identify a potential privacy leakage problem in recent researches and propose a more robust solution for addressing it. Through comprehensive security analysis and comparative experiments with existing approaches, we demonstrate that EzBoost achieves stronger privacy-preservation, higher accuracy and higher efficiency simultaneously.
Xinwen Gao, Shaojing Fu, Lin Liu 0018, Yuchuan Luo, Luming Yang
TrustCom5
2023 DEV-ETA: An Interpretable Detection Framework for Encrypted Malicious Traffic
abstract
Abstract Traffic encrypted technology enables Internet users to protect their data secrecy, but it also brings a challenge to malicious package detection. To tackle this issue, researchers have investigated into encrypted traffic analysis (ETA) in recent years. Existing works, however, only focus on the accuracy of malicious flow identification. Using ETA as a technical black box, they pay little attention to the internal details and explanation of models. In this paper, we, for the first time, introduce interpretable machine learning into ETA. We aim to provide a reasonable explanation for detection results, so as to enable one to understand and further trust network security analysts. We develop a complete analysis framework, named DEV-ETA (detection, explanation and verification of ETA). DEV-ETA applies post hoc interpretation methods to explain the detection results and verify the explanation using the joint distribution of support features on the dataset. We run thorough experiments to explain the detection result using three popular explanation approaches, namely SHAP, LIME and MSS, and we verify the explanation via the feature distribution plot. The experimental results show that our design can interpret the detection result of ETA model instead of just simply treating the model as a black box.
Luming Yang, Shaojing Fu, Kaitai Liang
Comput. J.1
2022 FlowSpectrum: a concrete characterization scheme of network traffic behavior for anomaly detection
Luming Yang, Shaojing Fu, Xuyun Zhang, Shize Guo, Chi Yang
World Wide Web1
2021 A Clustering Method of Encrypted Video Traffic Based on Levenshtein Distance
abstract
In order to detect the playback of illegal videos, it is necessary for supervisors to monitor the network by analyzing traffic from devices. However, many popular video sites, such as YouTube, have applied encryption to protect users’ privacy, which makes it difficult to analyze network traffic at the same time. Many researches suggest that DASH (Dynamic Adaptive Streaming over HTTP) will leak the information of video segmentation, which is related to the video content. Consequently, it is possible to analyze the content of encrypted video traffic without decryption. At present, most of the encrypted video traffic analysis adopts supervised learning methods, and there is little research on its unsupervised methods. Analysts are usually faced with unlabeled data, in reality, so the existing approaches will not work. The encrypted video traffic analysis methods based on unsupervised learning are required. In this paper, we proposed a clustering method based on Levenshtein distance for title analysis of encrypted video traffic. We also run a thorough set of experiments that verify the robustness and practicability of the method. As far as I am concerned, it is the first work to apply cluster analysis for encrypted video traffic analysis.
Luming Yang, Shaojing Fu, Yuchuan Luo
MSN1
2020 Markov Probability Fingerprints: A Method for Identifying Encrypted Video Traffic
abstract
Detecting illegal video plays an important role in preventing and countering crime in daily life. It is effective for supervisors to monitor the network by analyzing traffic from devices. In this way, illegal video can be detected when it is played on the network. Most Internet traffic is encrypted, which brings difficulties to traffic analysis. However, many researches suggest that even if the video traffic is encrypted, the segmentation prescribed by Dynamic Adaptive Streaming over HTTP (DASH) causes content-dependent fragments, which can be used to identify the encrypted video traffic without decryption. This paper presents Markov probability fingerprint for video, and then designs an algorithm for encrypted video streaming title identification. We demonstrate that an external attacker can identify the video title by analyzing the fragment sequence of encrypted video traffic. Based on the m-order Markov chain, we use the transition tensor of the fragment sequence generated by the video traffic as the video fingerprint, and prove its effectiveness. Then we explore approaches that can further improve the performance of methods in terms of discrimination accuracy. We make promising observations that the higher-order Markov chain, larger training set, and more detailed binning of fragments contribute to encrypted video traffic discrimination. We run a thorough set of experiments that illustrate that our method can achieve an outstanding accuracy rate up to 97.5%, which is superior to previous work.
Luming Yang, Shaojing Fu, Yuchuan Luo, Jiangyong Shi
MSN1
2019 Revisiting signal processing with spectrogram analysis on EEG, ECG and speech signals
Gaopeng Zhang, Luming Yang, V. S. Balaji, Elamaran Vellaiappan, Arunkumar N.
Future Gener. Comput. Syst.3
2011 DENNC: A Wireless Malicious Detection Approach Based on Network Coding
abstract
In wireless networks, communications among nodes are vulnerable to attacks launched by malicious nodes. Presently, Existing malicious node detection approaches either need special hardware or depend on node listening, node encryption or node identity authentication, resulting high costs of networks. In this paper, we present a novel network coding-based malicious detection approach called DENNC for wireless networks. The key idea is to use the characteristic of information exchange to validate the information packets. The neighboring nodes of the sending node may judge the malicious behaviors by checking the correctness of the data packets and related hash value. Our approach requires no superfluity hardware and does not use complicated secret key encryption mechanisms. Analysis reveals that the proposed approach can detect the malicious node in highly probability.
Hong Song 0004, Weiping Wang 0003, Luming Yang
TrustCom4
2011 Identifying the nature of stomach diseases by ultrasonography based on genetic neural network
Xianlai Chen, Luming Yang, Shu-chu Wang, Jianxin Wang 0001
Expert Syst. Appl.2
2008 A New Anonymity Measure Based on Partial Entropy
abstract
With the development of Internet applications, a number of anonymous communication systems have been realized to protect the identity of communication participants. Therefore, it is essential to give a theoretically based and practically usable objective numerical measure for the provided level of anonymity. In this paper some typical anonymity measures are analyzed and limitations of these measures was highlighted. Then a new anonymity measure based on partial entropy is proposed, in which the anonymity is measured by using the entropy of the probability distribution of some distinct subjects in anonymity set. The results of analysis and calculation show that the new measure is preferable for anonymity evaluation.
Guihua Duan, Weiping Wang 0003, Jianxin Wang 0001, Luming Yang
ICC4
2007 Study on consistent query answering in inconsistent databases
Luming Yang
Frontiers Comput. Sci. China2