Huajun Cui

dblp:189/4016 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-5579-195XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 1 first-author · 4 since 2021Security and privacy · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Edge Computing Enabled Large-Scale Traffic Flow Prediction With GPT in Intelligent Autonomous Transport System for 6G Network
abstract
The Intelligent Autonomous Transport System in 6G (6G-IATS) refers to the coordination of 6G, Artificial Intelligence (AI), and intelligent transportation systems, which is expected to revolutionize future intelligent transportation systems. In 6G-IATS, large-scale traffic flow prediction, affiliated with time series prediction, holds significant value for transportation planning and urban management. As an emerging AI method, Large Language Models (LLMs) have emerged prominently in time series forecasting. Unfortunately, it is challenging to achieve accurate and efficient large-scale traffic flow prediction by LLMs in 6G-IATS, due to the two issues: a) these LLMs fail to capture the spatio-temporal correlations in a large-scale road network, leading to limited prediction accuracy, and b) they process a substantial amount of training data on the central server, which imposes low training efficiency. Jointly considering the two concerns, this paper proposes a novel LLM and edge computing-based architecture for large-scale traffic flow prediction in 6G-IATS, called Spatio-Temporal Generative Large Language Model on Edge (STGLLM-E). In this architecture, we first decompose the entire large-scale road network into several subgraphs. To capture the spatio-temporal correlations, an LLM-based method named Spatio-Temporal Generative Large Language Model (STGLLM) including Spatio-Temporal Module (STM) and Generative Large Language Model (GLLM) is proposed. Secondly, to improve the training efficiency of the STGLLM-E, an edge training strategy based on edge servers is devised. Experiments are conducted on two real-world traffic flow datasets. The experimental results illustrate that the STGLLM-E is superior to the baselines in the prediction accuracy and the efficiency of training.
Yingchi Mao, Huajun Cui, Xiaoming He 0004, Mingkai Chen 0001
IEEE Trans. Intell. Transp. Syst.3
2025 QoE-Driven Proactive Caching With DRL in Sustainable Cloud-to-Edge Continuum
abstract
Cloud-enabled edge computing scenarios can intelligently cache and update the content on a periodic basis, thereby enhancing users' overall perception of quality, which is called quality of experience (QoE). To enhance the QoE, we aim to the multi-objective optimization, which maximizes the cache hit ratio while simultaneously minimizing traffic load and time latency. To address this issue, we focus on employing an innovative algorithm named HT-PAD, which provides a complete solution for prediction and decision-making for proactive caching. First, to improve the prediction accuracy of the cached content, we use the encoding layer in hyperdimensional computing to extract the information features. Second, HD-Transformer, as the prediction part of HT-PAD, is proposed to make predictions based on user preferences, historical information, and popular information. HD-Transformer uses DNN to predict user preferences and process time series data by combining hyperdimensional computation with Transformer. Third, to avoid error in the prediction content, we employ PER-MADDPG as the decision-making part of HT-PAD, which consists of Multi-Agent Deep Deterministic Policy Gradient (MADDPG) and Prioritized Experience Replay (PER). We use MADDPG to enhance the content decision-making and utilized PER to select appropriate training samples for PER-MADDPG. Finally, our experiments have shown that our proposed approach achieves the strong performance in terms of the edge hit ratio, the latency, and the traffic load, thus improving the QoE
Xiaoming He 0004, Huajun Cui, Yinqiu Liu, Mingkai Chen 0001, Maher Guizani, Shahid Mumtaz
IEEE Trans. Mob. Comput.3
2024 Intelligent reflecting surface-aided computation offloading in UAV-enabled edge networks
Wenyu Luo, Huajun Cui, Xuefeng Xian
Wirel. Networks2
2023 CACluster: A Clustering Approach for IoT Attack Activities Based on Contextual Analysis
abstract
Attacks against IoT have shown a rapid increase in both quantity and complexity. Analysts must handle massive alerts and determine the type of attack manually. In addition, the same attack activity may present polymorphism alert sequences due to overlapping attacks, adaptive attack strategy, error alerts, etc, which poses a severe challenge for human analysis. This manual-dependent and scenario-by-scenario security model is seriously overwhelming security analysts. This paper proposes a contextual-analysis-based clustering approach, CACluster, to aggregate similar attack activities end-to-end. It embeds alert context into vector space and uses an unsupervised clustering method to find similar attack activities based on domain matching and vector distance. Experimental results demonstrate that the CACluster could accurately aggregate similar attack activities, with 0.888 purity, reducing the number of attack activities by 84.8%. It will significantly cut down analysts’ workload.
Huiran Yang, Yan Zhang 0014, Yueyue Dai, Jiyan Sun, Huajun Cui, Can Ma, Weiping Wang 0005
ICPADS5
2023 LActDet: An Automatic Network Attack Activity Detection Framework for Multi-step Attacks
abstract
With the evolution of attack tactics, cyber-attacks are presenting a sophisticated trend. The multi-step attack has become the mainstream attack form, where adversaries implement multiple attack steps to achieve their goals, which poses server challenges to attack detection. Traditional research mainly concentrates on how a particular attack step is exploited but fails to identify the whole attack activity automatically. Manual analysis is required to correlate multiple steps and determine the fine-grained type of attack activities, which is a heavy workload. In addition, the high error rate of alerts results in a negative impact on attack-activity detection performance.To address these challenges, we propose a framework, LActDet, to automatically identify attack activities from the raw alerts end-to-end. Firstly, it utilizes a document-embedding method to vectorize attack-event descriptions. Second, a seq2seq model is implemented to embed the attack-event sequence into the attack-phase sequence to represent the framework of attack activity, aiming at improving the fault tolerance for error alerts. In the end, we propose a temporal-sequence-based classifier to identify attack activities. Our experimental results demonstrate that LActDet achieves higher detection accuracy, lower artificial dependence, and less system overhead.
Huiran Yang, Jiaqi Kang, Yueyue Dai, Jiyan Sun, Yan Zhang 0014, Huajun Cui, Can Ma
TrustCom6
2022 LibHunter: An Unsupervised Approach for Third-party Library Detection without Prior Knowledge
abstract
Third-party libraries (TPLs) are a significant component of mobile apps. They provide various functionalities, and developers employ them to facilitate app development. TPL detection is a fundamental task in security research, as it can impact other security studies. TPL can act as an assistant to malware detection, privacy leakage detection, etc. Because if a TPL carries malicious code, all apps that integrate the TPL can be considered risky. However, in some studies, TPLs can also act as noise, like app traffic fingerprinting. The TPL and app traffic are mixed during app runtime, making it difficult to fingerprint the app traffic accurately. Unfortunately, all existing TPL detection studies are working with prior knowledge of TPLs, as they need a whitelist or a train on known TPLs. However, new TPLs keep emerging, and it is not feasible for existing works to identify them-especially those who have network behaviors, as they may transfer inappropriate contents in the network. To this end, we propose LibHunter - an approach to identify TPLs without prior knowledge. LibHunter inspects the HTTP(S) traffic, logs the corresponding code execution traces, extracts features from the collected data, and performs a clustering algorithm to obtain TPLs. We apply LibHunter to 3000 apps. Results demonstrate that LibHunter can identify 79 TPLs, and about 60% of them are not detected by all existing works. We perform an analysis to show how important these TPLs are; we also present the visiting graph of these TPLs. Our findings bring light to the research community that existing tools are not accurate when encountering contemporary apps.
Huajun Cui, Guozhu Meng, Yuejun Li, Yan Zhang 0014, Jiyan Sun, Dali Zhu, Weiping Wang 0005
ISCC1
2022 Automated Privacy Network Traffic Detection via Self-labeling and Learning
abstract
With the increasing popularity of mobile devices, privacy leakage has become more and more serious. The inappropriate behaviors of mobile APPs have brought substantial security risks to the public (e.g., location leakage). Existing solutions detect privacy leakage based on network traffic analysis. However, they can only detect unencrypted traffic, which leads to failures in the face of encrypted traffic. To solve this challenge, we designed an Automated Privacy Traffic Detection system (APTD). APTD can automatically generate self-labeling privacy traffic datasets, learn to identify the encrypted privacy traffic, and accurately assess the risk of privacy leakage. Due to its automation capability, APTD can directly support privacy leakage detection for newly-emerged applications without any system changes. To comprehensively evaluate APTD, we conducted an experiment on 2327 real-world mobile APPs. APTD automatically generated a labeled dataset containing 27343 real-world encrypted traffic traces. Based on the dataset, APTD identifies privacy traffic, and performs a privacy leakage risk assessment of APPs. The results show that APTD achieves 97% accuracy and 99% recall on our dataset and identifies 12 APPs that transmit high-risk privacy data.
Yuejun Li, Huajun Cui, Jiyan Sun, Yan Zhang 0014, Guozhu Meng, Weiping Wang 0005
ISCC2
2019 Towards Homograph-Confusable Domain Name Detection Using Dual-Channel CNN
Guangxi Yu, Xinghua Yang, Yan Zhang 0014, Huajun Cui, Huiran Yang, Yang Li 0192
ICICS4
2019 Mitigating Negative Impacts on DNS Caches Caused by Disposable Domain Names
abstract
DNS caches play an important role in DNS querying. However, the performance of DNS caches will be remarkably influenced by disposable domain names, which are generated by services of cloud storage, social networks, etc., and belong to a new class of misused case of DNS. In this paper, we proposed a novel solution named DC3(Domain Classification and Cascade Cache) to mitigate the negative impact. Domain Classification adopts a classifier which is based on a long short-term memory (LSTM) network to prevent disposable domains from being cached. Cascade Cache is a refined cascade LRU policy considering cache size allocation to process the remaining disposable domains. By querying the real DNS traces collected from a large ISP network, experiment results show that this solution can detect disposable domain names and mitigate their negative impacts on DNS caches effectively. Specifically, in our dataset, 67.4% of all distinct domain names are detected as disposable domain names. Correspondingly, when getting rid of them by using this solution, we can raise the cache hit rate more than double.
Guangxi Yu, Yan Zhang 0014, Huajun Cui, Xinghua Yang, Yang Li 0192
ISCC3
2018 Improving NDN forwarding engine performance by rendezvous-based caching and forwarding
Guoqiang Zhang 0004, Huajun Cui
Comput. Networks4
2016 Performance and Implications of RAN Caching in LTE Mobile Networks: A Real Traffic Analysis
abstract
Deploying caches in mobile networks, especially in the radio access network (RAN) is regarded as a promising way to improve mobile user experiences and alleviate the increasing pressure of traffic growth. However, the characteristics of mobile traffic and the performance of RAN caching still remains unclear. In this paper, we extensively analyze the traffic characteristics, the content popularity and the cache performance using a unique dataset collected from a commercial LTE network of China Mobile, from the perspective of mobile access network. The dataset spans nearly a week and consists of a collection of approximately 62.1 millions HTTP sessions, generated by more than 3200 users distributed across three base stations. Based on this realistic dataset, we observe that HTTP traffic can be reduced by 24.4% on average and the hit ratio can reach up to 42.2%, using 100GB cache size. The implications on some fundamental design issues of practical RAN caching systems, including reasonable size of RAN cache, suitable locations of cache deployment and potential benefits of collaborative RAN caching, are further presented. We believe our findings will shed light on practical RAN caching system design.
Tao Lin 0001, Hongjia Li 0002, Haiyong Xie 0001, Jiasi Chen, Huajun Cui, Guoqiang Zhang 0004, Wei An 0002, Yang Li 0017
SECON5