Longtao He

dblp:33/3525 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0003-2225-6832ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 1 since 2021Security and privacy · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Odysseus: A Context-Level Pre-training Framework for Out-of-Distribution Encrypted Traffic Classification
Wenqi Dong, Longtao He, Gaopeng Gou, Zhen Li 0011, Junzheng Shi, Jianshuo Liu, Gang Xiong 0001
IWQoS3
2024 CETP: A novel semi-supervised framework based on contrastive pre-training for imbalanced encrypted traffic classification
Longtao He, Gaopeng Gou, Jing Yu 0007, Juncheng Guo, Gang Xiong 0001
Comput. Secur.2
2022 ReGFFM : A New Graph Model Designed with Spatial Information for Potential Social Connection Mining
abstract
Mining of potential social connection is an important task in the study of social network and has drawn attentions worldwide. Traditionally, graph model is used to infer potential social connections. However, when the graph model constituted by known social connections is sparse or even separated, it is difficult to mine potential social connections accurately. With the rapid development of LBSN (location based social network), a large amount of user spatial information could be obtained, which provides us a chance to settle the issue in another view. In this paper, a new graph model namely Reconstruction Graph model with Fusion Feature, ReGFFM for short, is designed for mining potential social connections with the help of users’ spatial information, which could reduce the negative effect caused by the sparsity of social connection graph. Experimental results have demonstrated the effectiveness of our model finally.
Kai Zhang 0079, Chenghai He, Longtao He
CSCWD7
2022 Cdga: A GAN-based Controllable Domain Generation Algorithm
abstract
Recently Command and Control (C&C) servers have attracted considerable attention in botnets and domain generation algorithms (DGAs) further enhance the stealth of C&C servers. However, Algorithmically Generated Domains (AGDs) generated by DGAs can be easily detected by previous DGA detection approaches. More specifically, the previous DGAs are hard to satisfy domain name rules, low repetition rate, and anti-detection in practical scenarios simultaneously. Designing an outstanding DGA has become a crucial issue from the botnet owner’s perspective. To mitigate these problems, we propose Cdga, a Controllable DGA via Generative Adversarial Networks (GAN), which is a popular backbone model for text generation in the natural language processing (NLP) community.Controllable text generation approaches are adopted by Cdga to ensure no repetition in the generated domain names and compliance with the domain rules. In addition to cheating DGA detectors, GANs are exploited to equip Cdga with a powerful anti-detection ability. Furthermore, our proposed method uses the technique of NLP to force the AGDs to meet language rules, where the generated domain names are difficult for recognition by human. By utilizing the time-dependent seed, Cdga can dynamically generate domain names, ensuring that the malware can connect to the C&C server conditioned on a specific time stamp. Experimental results demonstrate that the domain names generated by our method are realistic enough to be resistant to the state-of-the-art DGA detectors.
You Zhai, Jian Yang 0030, Longtao He, Liqun Yang, Zhoujun Li 0001
TrustCom4
2020 A Multi-source Self-adaptive Transfer Learning Model for Mining Social Links
Kai Zhang 0079, Longtao He, Chenglong Li 0001, Xiaoyu Zhang 0002
KSEM (2)2
2019 FS-Net: A Flow Sequence Network For Encrypted Traffic Classification
abstract
With more attention paid to user privacy and communication security, the volume of encrypted traffic rises sharply, which brings a huge challenge to traditional rule-based traffic classification methods. Combining machine learning algorithms and manual-design features has become the mainstream methods to solve this problem. However, these features depend on professional experience heavily, which needs lots of human effort. And these methods divide the encrypted traffic classification problem into piece-wise sub-problems, which could not guarantee the optimal solution. In this paper, we apply the recurrent neural network to the encrypted traffic classification problem and propose the Flow Sequence Network (FS-Net). The FS-Net is an end-to-end classification model that learns representative features from the raw flows, and then classifies them in a unified framework. Moreover, we adopt a multi-layer encoder-decoder structure which can mine the potential sequential characteristics of flows deeply, and import the reconstruction mechanism which can enhance the effectiveness of features. Our comprehensive experiments on the real-world dataset covering 18 applications indicate that FS-Net achieves an excellent performance (99.14% TPR, 0.05% FPR and 0.9906 FTF) and outperforms the state-of-the-art methods.
Chang Liu 0049, Longtao He, Gang Xiong 0001, Zigang Cao, Zhen Li 0011
INFOCOM2
2019 Large-scale Detection of Privacy Leaks for BAT Browsers Extensions in China
abstract
Although browser extensions bring users a better experience, it creates a hidden danger of privacy leakage. A common privacy leakage detection method is realized through detecting private data transmission. However, only the unintended transmission is considered to be a privacy leak. Therefore, the real challenge is to determine whether or not the transmission is user intended. In order to address this problem, we check the rationality of private data transmission by establishing a privacy model based on classification for extensions to confirm the scope of private data that can be uploaded and domains that can be sent to. Furthermore, we present BEDS (Browser Extension Detection System), a Chromium based extension dynamic detection system. BEDS first builds a privacy model for each extension and then records the extension's network logs and browser API logs when accessing specified pages. Finally, BEDS determines whether there exists a privacy leak according to the strict privacy leakage judgment rules. We test our implementation in large scale on extensions in browsers developed by China's three major Internet companies and complete 15 months of continuous tracking. After examining a total of 14,487 extensions, 1,897 privacy leaks are identified, all results have been inspected by manual and the accuracy of BEDS is over 97%. A number of domains that illegally collect private user data are discovered and tracked. Our results show that about 47,000 Chinese IPs upload private information to suspicious servers every day.
Longtao He, Zhoujun Li 0001, Liqun Yang, Yu Wang 0206
TASE2
2018 MaMPF: Encrypted Traffic Classification Based on Multi-Attribute Markov Probability Fingerprints
abstract
With the explosion of network applications, network anomaly detection and security management face a big challenge, of which the first and a fundamental step is traffic classification. However, for the sake of user privacy, encrypted communication protocols, e.g. the SSL/TLS protocol, are extensively used, which results in the ineffectiveness of traditional rule-based classification methods. Existing methods cannot have a satisfactory accuracy of encrypted traffic classification because of insufficient distinguishable characteristics. In this paper, we propose the Multi-attribute Markov Probability Fingerprints (MaMPF), for encrypted traffic classification. The key idea behind MaMPF is to consider multi-attributes, which includes a critical feature, namely “length block sequence” that captures the time-series packet lengths effectively using power-law distributions and relative occurrence probabilities of all considered applications. Based on the message type and length block sequences, Markov models are trained and the probabilities of all the applications are concatenated as the fingerprints for classification. MaMPF achieves 96.4% TPR and 0.2% FPR performance on a real-world dataset from campus network (including 950,000+ encrypted traffic flows and covering 18 applications), and outperforms the state-of-the-art methods.
Chang Liu 0049, Zigang Cao, Gang Xiong 0001, Gaopeng Gou, Siu-Ming Yiu, Longtao He
IWQoS6
2016 Information fusion-based method for distributed domain name system cache poisoning attack detection and identification
abstract
In this study, the authors consider the detection and identification problems of distributed domain name system (DNS) cache poisoning attack. In the considered distributed attack, multiple cache servers are invaded simultaneously and the attack intensity for each cache server is slight. It is difficult to detect and identify the distributed attack by the existing local information‐based detection methods, as the abnormal features for each cache server are indistinctive under distributed attack. To handle this problem, they propose an information fusion‐based detection and identification methods. They find that the entropies of the query Internet protocol (IP) addresses for all cache servers are approximately stationary and statistically independent under normal cases. When distributed attack happens, they show the fact that the correlation of the entropies among all cache servers could increase dramatically. On the basis of this feature, they make use of principal component analysis to design the detection and identification methods. Specifically, attack is true when the maximum eigenvalue of the normalised entropies matrix exceeds a threshold, and the attacked servers are identified by the main loading vector. At last, they take a large‐scale DNS in China and a simulation as two examples to show the effectiveness of their methods.
Xianglei Dang, Longtao He
IET Inf. Secur.4
2005 The wide window string matching algorithm
Longtao He, Binxing Fang, Jie Sui
Theor. Comput. Sci.1
2004 Linear Nondeterministic Dawg String Matching Algorithm
Longtao He, Binxing Fang
SPIRE1