VLDB 2026 Research / reviewers in the wild / expert
Mingxin Cui
dblp:207/6586
· DBLP profile ↗
24ranked-venue papers
1as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 8 since 2021Security and privacy · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TrafficCL: Contrastive learning on network traffic for accurate, efficient and robust IP cross-regional detection
Mingxin Cui, Gaopeng Gou, Chang Liu 0049, Yong Wang 0046, Guoming Ren, Gang Xiong 0001 |
Comput. Networks | 2 |
| 2025 | Beneath the Heavens: A Thorough Measurement Study of the Starlink Terrestrial NetworkabstractThe emerging low earth orbit (LEO) satellite Internet has gained worldwide popularity. Starlink, a prominent LEO satellite Internet service provider, has attracted the most users due to its low latency, wide coverage, and strong usability. Current research on Starlink primarily focuses on the space segment and its impact on network performance. However, as an essential component, the architecture and unique features of the Starlink Terrestrial Network (SLTN) are not well-explored, which significantly affects the performance, security, and development of the whole network. In this paper, we fill this gap by conducting a thorough measurement study to profile the SLTN from various aspects and reveal its potential effects, with specific attention to the network assets, topology, and routing strategies. We developed a novel framework including active and passive measurement methods for collecting multiple network assets, tracing different route paths, and scanning active service of the SLTN. The open source intelligence was utilized for the first time to collect extensive real-user network status. Leveraging these techniques, we observed a rapid expansion of the Starlink service with the latest network assets, collected in January 2025, including over 23.5K/28 IP prefixes residing in 143 countries. We mapped the consistent topology of the global Starlink Internet access service and identified the specific IP addresses associated with the four types of network routing nodes. Multiple internal routing strategies were uncovered, which facilitate direct user-to-user interactions. Particularly, we revealed the switches of terrestrial infrastructure that users connected and the changes in routing strategies, which should be considered in future quality of service (QoS) evaluations. Our measurement also provides methods for improving users' perceptions and serves as a basis for studies like security risk evaluation and service discovery. Yanbo Wu, Mingxin Cui, Gaopeng Gou, Yuhao Wei, Gang Xiong 0001, Zhen Li 0011, Xinlei Ju |
IWQoS | 2 |
| 2024 | Tunnel User Behavior Identification Based on Self-Supervised Pre-TrainingabstractWith the widespread use of tunnel technology, the volume of encrypted tunnel traffic is rapidly increasing. Malicious users can transmit harmful information secretly through tunnels to bypass firewall censorship. Therefore, developing effective techniques to identify tunnel user behaviors is crucial. However, current efforts in tunnel traffic classification primarily focus on coarse-grained application identification and encounter the problem of insufficient extraction of tunnel traffic feature information. In this paper, we refine the previous tunnel traffic identification granularity from prevalent application identification to behavior identification, and propose TF-Net, a novel deep learning framework for fine-grained identification of tunnel user behaviors. TF-Net extracts features from raw bytes, packet length sequence, packet time interval sequence of tunnel traffic. It employs self-supervised pre-training to learn contextual distributions of raw bytes from unlabeled datasets, thereby enhancing the model’s ability to characterize packets in a tunnel flow. Moreover, we fine-tune the model based on the classification objectives of different tasks to achieve more versatile and accurate tunnel identification. Comprehensive experiments are conducted on three real-world encrypted tunnel traffic datasets, demonstrating that TF-Net achieves outstanding performance and outperforms state-of-the-art methods. Lingyun Ye, Zhishen Zhu, Gaopeng Gou, Gang Xiong 0001, Mingxin Cui |
HPCC | 5 |
| 2024 | Incremental encrypted traffic classification via contrastive prototype networks
Wei Cai 0007, Chengshang Hou, Mingxin Cui, Bingxu Wang, Gang Xiong 0001, Gaopeng Gou |
Comput. Networks | 3 |
| 2023 | Covertness Analysis of Snowflake Proxy RequestabstractSnowflake is a special proxy system against IP-based network blocking. As its IP addresses refresh frequently, faster than IP blacklist’s update, users can exploit it to access blocked websites. To block snowflake, existing methods focus on detecting snowflake proxies. But they are susceptible to various factors, for example, proxy’s location and version. In the paper, we propose a new manner to block snowflake. We observe that to adapt fast IP changes, users need to request latest proxies from proxy database before using snowflake. Thus, adversaries can block snowflake by detecting proxy request instead of proxy itself. To verify our method, we analyse covertness of snowflake proxy requests, that has been protected by imitating normal web requests. After comparing with typical web requests, we find the imitation is vulnerable in packet size, direction, time and network speed, such as, the latency time is higher than normal obviously. Using the four vulnerabilities, we train machine learning algorithm to detect snowflake proxy requests in reality. Experimental results demonstrate that proxy request can be detected accurately across different versions at the beginning of connection. In conclusion, our work paves a new way to block snowflake. Yibo Xie, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011, Mingxin Cui |
CSCWD | 5 |
| 2023 | Analysing Covertness of Tor Bridge RequestabstractTor bridges are hidden entrances of Tor network. Users can exploit bridges to hide their visits of Tor. To restrict hidden Tor visits, many attacks focus on bridge information discovery or bridge traffic detection. But these attacks are less effective because bridges' information cannot be discovered thoroughly and its traffic are often obfuscated. In the paper, we present a novel attack to stop hidden Tor visits. We observe that users need to request information of bridges from a database before visiting Tor network. Thus, attackers can stop Tor visits by detecting the process of bridge requests rather than bridge itself. To verify our attack's feasibility, we analyse covertness of the most widely-used bridge request tool, which imitates normal network request when communicating with bridge database. After comparing with five types of typical web request, we find that this tool fails to imitate in packet time, size and direction, for example, the variation of simulated packet sizes are more dynamic than normal. Based on the three imitation vulnerabilities, we train machine learning algorithms to detect bridge request. Extensive experiments demonstrate that bridge request can be identified with high accuracy and very low false-positive rates in real-world. In conclusion, our work paves a new way to block evasive Tor visits. Yibo Xie, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011, Mingxin Cui |
ICC | 5 |
| 2023 | FA-Net: More Accurate Encrypted Network Traffic Classification Based on Burst with Self-AttentionabstractEncrypted network traffic classification (ENTC) is crucial in fields including network cyberspace security, network administration and service quality. Combining the machine learning algorithms with manual-designed burst features has been studied extensively in the ENTC community. However, these features depend on professional experience heavily, which needs lots of human effort. These hand-crafted features are task-oriented and incomplete in various complex tasks. What's more, they are also affected by the potential network jitters. In this paper, we propose a novel encrypted traffic classification method FA-Net to mine burst features. We adopt two hierarchical multi-head self-attention encoders to enumerate all potential intra-burst features and inter-burst dependencies completely, and select the optimal associations automatically. For more robust against network jitter, we design an additional burst positional encoding to loose the model's sensitivity about out-of-order packets within bursts. We evaluate the FA-Net on multiple datasets, including website and mobile application classification tasks. The results show the FA-Net model outperforms other state-of-the-art methods in all the datasets, even gains more than 5% absolute improvement in accuracy. Additionally, the quantitative measurements about burst feature similarity show that the burst features learned by FA-Net exhibits more intraclass similarity and more inter-class separation. Mingxin Cui, Chengshang Hou, Wei Cai 0007, Zhen Li 0011, Gang Xiong 0001, Gaopeng Gou |
IJCNN | 2 |
| 2023 | The Potential Utility of Image Descriptions: User Identity Linkage across Social Networks Based on MultiModal Self-Attention FusionabstractThe task of user identity linkage across social networks aims to predict whether users from different social networks refer to the same person. This task plays a crucial role in cross-social network information dissemination and intelligent recommendations. However, existing user identity linkage tasks suffer from several challenges: 1) excessive reliance on social network topology, neglecting users’ visual modality information; 2) inadequate handling of noise in user feature data; and 3) ineffective fusion of users’ multimodal information. To address these issues, we investigated a method that utilizes heterogeneous multimodal posts, including user-generated text, images, and check-in messages, to achieve user identity linkage across social networks. We innovatively leveraged a pre-trained model for image-to-text conversion to further explore users’ image data and proposed an adversarial learning model based on the multimodal self-attention mechanism (AMSA). The AMSA model consists of four components: user feature extraction, user feature processing, user feature fusion, and adversarial learning. Specifically, AMSA initially employed advanced pre-trained models to extract features from multiple modalities of users, including images and text. Subsequently, it utilized multiple mechanisms, such as multi-head self-attention, to process data from each modality separately and then fused them into user representation vectors. Finally, AMSA employed adversarial learning to enhance the model’s learning capacity and mitigate semantic disparities in user information across different platforms. We conducted model performance evaluations on publicly available datasets, and experimental results demonstrated the superiority of the proposed AMSA model. Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011, Mingxin Cui |
IPCCC | 5 |
| 2023 | Zero-relabelling mobile-app identification over drifted encrypted network traffic
Mingxin Cui, Chang Liu 0049, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011 |
Comput. Networks | 2 |
| 2022 | MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question AnsweringabstractKnowledge-based visual question answering requires the ability of associating external knowledge for open-ended cross-modal scene understanding. One limitation of existing solutions is that they capture relevant knowledge from text-only knowledge bases, which merely contain facts expressed by first-order predicates or language descriptions while lacking complex but indispensable multimodal knowledge for visual understanding. How to construct vision-relevant and explainable multimodal knowledge for the VQA scenario has been less studied. In this paper, we propose MuKEA to represent multimodal knowledge by an explicit triplet to correlate visual objects and fact answers with implicit relations. To bridge the heterogeneous gap, we propose three objective losses to learn the triplet representations from complementary views: embedding structure, topological relation and semantic space. By adopting a pretraining and fine-tuning learning strategy, both basic and domain-specific multimodal knowledge are progressively accumulated for answer prediction. We outperform the state-of-the-art by 3.35% and 6.08% respectively on two challenging knowledge-required datasets: OK-VQA and KRVQA. Experimental results prove the complementary benefits of the multimodal knowledge with existing knowledge bases and the advantages of our end-to-end framework over the existing pipeline methods. The code is available at https://github.com/AndersonStra/MuKEA. Jing Yu 0007, Bang Liu 0003, Yue Hu 0002, Mingxin Cui, Qi Wu 0001 |
CVPR | 5 |
| 2022 | Accurate mobile-app fingerprinting using flow-level relationship with graph neural networks
Zhen Li 0011, Peipei Fu, Wei Cai 0007, Mingxin Cui, Gang Xiong 0001, Gaopeng Gou |
Comput. Networks | 5 |
| 2022 | Empirical Study on the Influencing Factors of WeChat Business IntegrityabstractWith the increasing popularity of the smart phones and electronic payment, WeChat shopping has become a trendy lifestyle for many people. However, the issue of WeChat business integrity has gradually appeared due to the virtuality of the Internet. This paper analyzed the development of WeChat business and its business integrity. The influencing factors of WeChat business integrity and related hypotheses had been studied based on theoretical and practical analysis. The reliability and validity of the data collected through questionnaires were tested with SPSS24.0. Empirical analysis was done to the hypotheses by using the structural equation model. The results indicated that products and service quality, after-sale guarantee, payment security and Word-of-mouth had a prominent positive effect on the integrity of WeChat business. Yangling Xiao, Bingjun Tong, Yanmei Sheng, Mingxin Cui |
J. Glob. Inf. Manag. | 4 |
| 2021 | BAPM: Block Attention Profiling Model for Multi-tab Website Fingerprinting Attacks on TorabstractWebsite fingerprinting attacks on Tor pose an security issue in anonymity privacy, in which attackers can identify websites visited by victims through passively capturing and analyzing encrypted packet traces. Although related works have been studied over a long period, most of them focus on single-tab packet traces which only contain one page tab’s data. However, users often open multiple page tabs successively when browsing the web, and multi-tab packet traces generated will corrupt common single-tab attacks. Existing multi-tab attacks still depend on an elaborate feature engineering, besides, they fail to exploit the overlapping area which contains the mixed data of two adjacent page tabs, thus suffering from the information lost or confusion. In this paper, we propose a Block Attention Profiling Model named BAPM as a new multi-tab attacking model. Specifically, BAPM fully utilizes the whole multi-tab packet trace including the overlapping area to avoid information lost. It generates a tab-aware representation from direction sequences and performs the block division to separate mixed page tabs as clearly as possible, thus relieving the information confusion. Then the attention-based profiling is used to group blocks belonging to the same page tab and finally multiple websites are simultaneously identified under a global view. We compare BAPM with state of the art multi-tab attacks, and BAPM outperforms comparison methods even with larger overlapping area. The effectiveness of model design is also validated through ablation, sensitivity and generalization analysis. Gang Xiong 0001, Gaopeng Gou, Zhen Li 0011, Mingxin Cui, Chang Liu 0049 |
ACSAC | 5 |
| 2021 | TA-GAN: GAN based Traffic Augmentation for Imbalanced Network Traffic ClassificationabstractAs the mainstream in network traffic classification (NTC), machine learning (ML) based methods suffer performance degradation due to the imbalance distribution of Internet traffic. Data augmentation methods including the traditional oversampling techniques and the Generative Adversarial Network (GAN) based generation methods are most commonly used to counter the imbalance problem in NTC. However, the former is prone to overfitting and introducing noise. The latter overcomes the above weaknesses, but the quality of the generated traffic samples is difficult to judge. Besides, these methods all divide the imbalanced traffic classification problem into two subproblems, which cannot guarantee the global optimality. In this paper, we propose a GAN based Traffic Augmentation (TA-GAN) for imbalanced traffic classification. TA-GAN is an end-to-end framework that integrates the generation of the minority traffic samples with the training of the target classifier. We design the feedback mechanism to better guide the direction of the sample generation and simultaneously indicate the quality of the synthesized samples. Moreover, the existing deep learning-based NTC methods can be easily adapted to imbalance scenarios with TA-GAN. Comprehensive experiments on the public ISCXVPN2016 dataset demonstrate that TA-GAN effectively mitigates the influence of traffic imbalance (a maximum 14.64% improvement to the minority class'$F_{1}$score) and outperforms the state-of-the-art methods. Gang Xiong 0001, Zhen Li 0011, Junzheng Shi, Mingxin Cui, Gaopeng Gou |
IJCNN | 5 |
| 2021 | Combating Imbalance in Network Traffic Classification Using GAN Based OversamplingabstractWith the proliferation of encrypted traffic, machine learning (ML) based network traffic classification (NTC) has become the mainstream method. However, most studies ignored two issues. On the one hand, Internet traffic presents a natural uneven distribution. On the other hand, machine learning algorithms generally aim to achieve the highest overall accuracy without considering class imbalance. This leads to severe performance degradation of existing ML-based NTC schemes when facing imbalanced scenarios. In this paper, we design a novel Generative Adversarial Network (GAN) architecture to generate traffic samples, in which the addition of the classifier and the pretraining module makes the generation process more stable and effective. We propose an end-to-end framework for imbalanced traffic classification, named ITCGAN, which can generate traffic samples for minority classes to adaptively rebalance the original traffic and simultaneously train the optimal classifier. We evaluate its effectiveness on the public ISCXVPN2016 dataset based on the global metrics and individual metrics. The results show that our method performs well in imbalanced NTC tasks, fully alleviating the performance degradation (a 10.27-percentage-point improvement to the precision of the most minority class). Meanwhile, it surpasses five state-of-the-art oversampling methods. Gang Xiong 0001, Zhen Li 0011, Junzheng Shi, Mingxin Cui, Gaopeng Gou |
Networking | 5 |
| 2021 | CQNet: A Clustering-Based Quadruplet Network for Decentralized Application Classification via Encrypted Traffic
Yu Wang 0134, Gang Xiong 0001, Chang Liu 0049, Zhen Li 0011, Mingxin Cui, Gaopeng Gou |
ECML/PKDD (4) | 5 |
| 2021 | TMT-RF: Tunnel Mixed Traffic Classification Based on Random Forest
Panpan Zhao, Gaopeng Gou, Chang Liu 0049, Yangyang Guan, Mingxin Cui, Gang Xiong 0001 |
SecureComm (1) | 5 |
| 2021 | Universal Website Fingerprinting Defense Based on Adversarial ExamplesabstractWebsite fingerprinting (WF) attacks pose a threat to privacy of web activity, especially on anonymity networks such as Tor. Recent studies show that the deep neural network (DNN) significantly improves the impact of website fingerprinting attacks. Especially, DNN-based attack undermines the existing defense methods which are mainly rely on the manually designed rule. In this paper, we present a novel defense that generates universal perturbation that can transform original examples to adversarial examples which is effectively defending against a specific WF model. The proposed defense is evaluated on state-of-the-art DNN attack over a public Tor traffic dataset. The experimental results show our adversarial example generation method performs better than the baseline methods. The proposed defense defeats all existing WF attacks based on deep neural networks with a low overhead. Comparing with state-of-the-art defenses such as Walkie-Talkie and WTF-PAD with a lower bound of 31% and 64% overheads, the proposed defense achieves identical defense performance with at least 50% bandwidth overhead saving. Chengshang Hou, Junzheng Shi, Mingxin Cui, Mengyan Liu, Jing Yu 0007 |
TrustCom | 3 |
| 2021 | Attack versus Attack: Toward Adversarial Example Defend Website Fingerprinting AttackabstractWebsite Fingerprinting (WF) attack is a side channel attack against encrypted tunnels which infers network activities of encrypted tunnels users. WF attack has been successfully applied to the Tor network, which poses a huge threat to the privacy of Tor visitors. A lot of countermeasures are therefore proposed to defend against such attacks. However, the newest attack successfully undermined the existing defense leveraging deep learning technique. In this paper, we propose an defense named Attack to Attack (A2A) that leverages adversarial example to attack the attacker's classifier. A2A treats website fingerprinting model as a black box. In order to find effective adversarial examples for the attacker's model, A2A manipulates traffic iteratively according to the output of a substitute model which is an elaborate model intentionally learning a similar classification boundary with the attacker's model. We evaluate the effectiveness of A2A on a public tor traffic dataset and the newest WF attack. The experimental results show that the proposed method provides effective defense with a bandwidth overhead of 2.2%, which significantly outperforms the manually designed defense (typically has a bandwidth overhead of 31%). Chengshang Hou, Junzheng Shi, Mingxin Cui, Qingya Yang |
TrustCom | 3 |
| 2021 | SiamHAN: IPv6 Address Correlation Attacks on TLS Encrypted Traffic via Siamese Heterogeneous Graph Attention Network
Tianyu Cui, Gaopeng Gou, Gang Xiong 0001, Zhen Li 0011, Mingxin Cui, Chang Liu 0049 |
USENIX Security Symposium | 5 |
| 2021 | Survey of security supervision on blockchain from the perspective of technology
Yu Wang 0134, Gaopeng Gou, Chang Liu 0049, Mingxin Cui, Zhen Li 0011, Gang Xiong 0001 |
J. Inf. Secur. Appl. | 4 |
| 2020 | MalFinder: An Ensemble Learning-based Framework For Malicious Traffic DetectionabstractMalicious events pose a significant threat to the current increasingly interconnected Internet community. Detection based on features of network traffic and machine learning algorithms is a common approach to identify malicious events. The performance of approaches is associated with the used features and algorithms. In this paper, we propose MalFinder, an ensemble learning-based framework for malicious traffic detection. Considering the trend of network traffic encryption and the complexity of decrypting traffic, we utilize statistical features and sequence features to describe network traffic. We extend the dimensions of these two types of features to enhance their capability for representing traffic data. Feature importance analysis and contrast experiments illustrate the effectiveness of our new features. Among our selected classifiers suitable for malicious traffic detection, boosting-based classifiers XGBoost and LightGBM can reduce bias, and bagging-based classifier Random Forest can reduce variance. Stacking, which is the integration method of the classification results used in our framework, can improve the generalization ability of the method. MalFinder can achieve 96.58% F-measure and 95.44% accuracy in the malicious traffic detection task on a real-world dataset, whose results are better than those of comparison methods. In terms of unseen malicious traffic discovery, MalFinder still provides good performance with 93.46% F-measure and 91.04% accuracy, which even surpasses the results in the task of known malicious traffic detection of other comparative methods. With consideration of the scarcity of public data sets used for malicious traffic detection, we have exposed our self-built dataset for more extensive researches. Candong Rong, Gaopeng Gou, Mingxin Cui, Gang Xiong 0001, Zhen Li 0011, Li Guo 0001 |
ISCC | 3 |
| 2020 | TransNet: Unseen Malware Variants Detection Using Deep Transfer Learning
Candong Rong, Gaopeng Gou, Mingxin Cui, Gang Xiong 0001, Zhen Li 0011, Li Guo 0001 |
SecureComm (2) | 3 |
| 2017 | POSTER: A Comprehensive Study of Forged Certificates in the WildabstractWith the widespread use of SSL, many issues have been exposed as well. Forged certificates used for MITM attacks or proxies can make SSL encryption useless easily, leading to privacy disclosure and property loss of careless victims. In this paper, we implement a large scale of passive measurement of SSL/TLS and analyze the forged certificates in the wild comprehensively. We measured SSL/TLS connection for 16 months on two large research networks, which provided a total of 100 Gbps bandwidth. We gathered nearly 135 million leaf certificates and studied the forged ones. Our findings reveal main reasons of signing forged certificates, and show the preference of them. Finally, we find out several suspicious servers that might be used for MITM. Mingxin Cui, Zigang Cao, Gang Xiong 0001, Junzheng Shi |
CCS | 1 |