VLDB 2026 Research / reviewers in the wild / expert
Mingkai Tong
dblp:247/3305
· DBLP profile ↗
8ranked-venue papers
2as first author
2since 2021 · last 2025
0009-0000-7477-2925ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 3 · 1 first-authorComputer networks · 2Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Context-Aware Clustering Approach for Assisting Operators in Classifying Security AlertsabstractModern software has evolved from delivering software products to web services and applications, which need to be protected by security operation centers (SOC) against ubiquitous cyber attacks. Numerous security alerts are continuously generated every day, which have to be efficiently and correctly processed to identify potential threats. Many AIOps (artificial intelligence for IT operations) approaches have been proposed to (semi-)automate the inspection of alerts so as to reduce manual effort as much as possible. However, due to the ever-complicating attacks, a significant amount of manual work is still required in practice to ensure correct analysis results. In this paper, we propose a Context-Aware cLustering approach for cLassifying sEcurity alErts (CALLEE), which fully exploits the rich relationships among alerts in order to precisely identify similar alerts, significantly reducing the workload of SOC. Specifically, we first design a core conceptual model to capture connections among security alerts, based on which we establish corresponding heterogeneous information networks. Next, we systematically design a set of meta-paths to profile typical alert scenarios precisely, contributing to obtaining the representation of security alerts. We then cluster security alerts based on their contextual similarities, considering the tradeoff between the number of clusters and the homogeneity of each cluster. Finally, security operators only need to manually inspect a limited number of alerts within each cluster, pragmatically reducing their workload while ensuring the accuracy of alert classification. To evaluate the effectiveness of our approach, we collaborate with our industrial partner and pragmatically apply the approach to a real alert dataset. The results show that our approach can reduce the workload of SOC by 99.76%, outperforming baseline approaches. In addition, we further investigate the integration of our proposal with the real business scenario of our industrial partner. The feedback from practitioners shows that CALLEE is pragmatically applicable and helpful in industrial settings. Yu Liu 0090, Tong Li 0001, Runzi Zhang, Mingkai Tong, Wenmao Liu, Zhen Yang 0004 |
IEEE Trans. Software Eng. | 5 |
| 2022 | Context2Vector: Accelerating security event triage via context representation learning
Runzi Zhang, Wenmao Liu, Dujuan Gu, Mingkai Tong, Jianxin Xue, Huanran Wang |
Inf. Softw. Technol. | 6 |
| 2020 | Far from classification algorithm: dive into the preprocessing stage in DGA detectionabstractDomain-Flux technique has been widely used by attackers to maintain a botnet for many years and the core of it is the adoption of domain generation algorithm (DGA). To combat attackers, there are lots of works in DGA domain detection area recently. But they usually collect quite limited data and conduct experiments in a closed dataset, meaning that the DGA data and the benign data they collected can not well represent the real distribution between them. Moreover, they handle the domains roughly and use the origin data to train the classifier directly, which is also not adequate to classify these two types of domains with lots of false positives and false negatives happening during the real-world deployment. In this paper, we conduct the first large-scale DGA domain analysis in traffic level and argue that the preprocessing stage is also vital for the final classifier, which is usually ignored by the existing works. We collect the largest amount of DGA domain data than prior works and collect DNS log offered by a big company, whose DNS data covers most important industries in China. Based on this data, we analyze the distribution of DGA domains in traffic and give quantifiable results showing that NXDomain (domain not exist) is more suitable for DGA detection. Moreover, we give detailed preprocessing steps to handle the original domains. Our experiment shows that with the preprocessing stage mentioned above, classifier performs better in DGA detection task. Our research indicates that improving the classification algorithm is far from enough in DGA detection and the preprocessing stage is also the key component in bringing the DGA detection methods from lab to product. Mingkai Tong, Runzi Zhang, Jianxin Xue, Wenmao Liu, Jiahai Yang 0001 |
TrustCom | 1 |
| 2020 | CMIRGen: Automatic Signature Generation Algorithm for Malicious Network TrafficabstractAlthough machine learning (ML) based solutions are ever-evolving for the attack defending paradigm, signatures of malicious network traffic are vital resources for intrusion detection systems (IDSs) and network forensic procedure, covering the lack of interpretability and stability for ML models. However, signature extraction is still a time and labor consuming task nowadays, resulting in possible increase of the attackers' dwell time. Existing automatic solutions rely too much on sequence similarity based and heuristic based methods, encountering performance degradation in large scale and dynamic network environment. In this paper, we present a novel method, called Clustering and Model Inference-based Rule Generation (CMIRGen), automatically generating token-set based signature rules for malicious traffic payloads to be inspected. CMIRGen leverages both optimized sequence similarity based and black-box model inference based methods to extract patterns from homogeneous and heterogeneous payloads respectively. Experimental evaluations have been conducted on several datasets and show the CMIRGen framework can extract discriminative signatures, presenting high recall rate and low false positive rate at the same time for malicious content recognition. Runzi Zhang, Mingkai Tong, Jianxin Xue, Wenmao Liu |
TrustCom | 2 |
| 2019 | Network Coded Cooperative Multicast in Integrated Terrestrial-Satellite NetworksabstractWireless services have been extended from connection-centric to content-centric and this brings rapid increasing data traffic. Considering the concurrent multiple requests for popular contents, multicast is a promising delivery scheme to exploit content reuse. Satellite multicast is superior to others due to the coverage. When terrestrial base stations (BSs) are integrated with the satellite, cooperative transmission is then enabled to handle the fading satellite channels as well as maximize data throughput. Therefore, this paper proposes a co-operative multicast scheme for content delivery in the integrated terrestrial-satellite networks (ITSN) which is further enhanced by network coding. To exploit bandwidth of both terrestrial base stations (BSs) and the satellite, this cooperative multicast scheme uses two-phase transmission. The satellite multicast contents to subscribers in an opportunistic manner with real-time data rate adaptation. Then terrestrial BSs retransmit the lost packets for reliable transmission. Ground users are allocated to subgroups for the channel diversities while network coding is applied to packet loss recovery with fewer retransmissions. Sufficient numerical results demonstrate the enhancement on network throughput brought by the cooperative multicast scheme. Xinmu Wang, Hewu Li, Mingkai Tong, Kang Pan, Qian Wu 0001 |
ISCC | 3 |
| 2019 | Cooperative Network-Coded Multicast for Layered Content Delivery in D2D-Enhanced HetNetsabstractThe rapid growth in multimedia applications over cellular networks calls for multicast services. Multicast is an efficient means of delivering contents to multiple users while efficiently utilizing network resources. Layered streaming, e.g., scalable video coding (SVC) provides an excellent solution to handle channel diversities in wireless multicast. This paper presents a cooperative multicast scheme for scalable video content delivery in D2D-enabled heterogeneous cellular networks (HetNets). To extend the multicast service beyond base stations (BSs), D2D links are used to help relay content for cellular multicast. Network coding (NC), implemented through random linear network coding (RLNC) with unequal error protection (UEP) is incorporated in layered contents to enhance reliability and throughput. This paper tries to optimize the multicast scheduling procedures, aiming to assign the optimal modulation and coding schemes (MCSs) for transmissions. The constraints on cache size, backhaul capacity and channel fading are comprehensively considered. Besides, this paper also presents a interference-aware approach for D2D link selection in order to further improve the cooperative multicast. Sufficient numerical results have demonstrated the improvement brought by the proposed scheme significantly on the content delivery efficiency as well as quality of experience (QoE). Xinmu Wang, Hewu Li, Mingkai Tong, Wenbing Yao, Qian Wu 0001 |
ISCC | 3 |
| 2019 | D3N: DGA Detection with Deep-Learning Through NXDomain
Mingkai Tong, Xiaoqing Sun, Jiahai Yang 0001, Hui Zhang 0052, Shuang Zhu |
KSEM (1) | 1 |
| 2019 | HinDom: A Robust Malicious Domain Detection System based on Heterogeneous Information Network with Transductive Classification
Xiaoqing Sun, Mingkai Tong, Jiahai Yang 0001 |
RAID | 2 |