EDBT 2026 Demo / reviewers in the wild / expert
Qingyun Liu 0001
dblp:47/9461-1
· DBLP profile ↗
100ranked-venue papers
1as first author
81since 2021 · last 2026
0000-0003-4815-3463ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 32 · 23 since 2021Security and privacy · 20 · 15 since 2021Human-computer interaction and ubiquitous computing · 18 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Databases, data management, data science and information retrieval · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Systems, architecture and hardware · 5 · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PriAgent: A Collaborative Multi-Agent Framework for Auditing Android Privacy ComplianceabstractStringent regulations like General Data Protection Regulation (GDPR) mandate that an application's code-level data handling must align with its natural-language privacy policy, creating a critical auditing challenge. However, existing methods, predominantly reliant on static analysis, suffer from a critical limitation: in their pursuit of soundness via over-approximation, they exhibit "semantic blindness"—detecting what data flows exist but not why. This leads to an overwhelming volume of false positives, rendering automated auditing impractical. To bridge this gap, we introduce PriAgent, a novel framework that approaches compliance auditing as a multi-stage, AI-driven reasoning task. Instead of a monolithic model, PriAgent deploys a team of specialized agents that execute a divide-and-conquer strategy. They systematically prune the analysis space by abstracting data flows, pinpoint semantic loci critical for inspection, and perform on-demand summarization of large code blocks to ensure scalability. PriAgent leverages Retrieval-Augmented Generation (RAG) with a curated knowledge base of Android APIs, equipping agents to discern potentially non-compliant behavior from benign functionality. By correlating code-level evidence with the app's stated privacy policy, PriAgent delivers a holistic and explainable verdict for each potential violation. Our evaluations demonstrate that PriAgent significantly reduces false positives, enabling a more scalable and precise compliance audit. Zhao Li 0010, Zhuojun Jiang, Jiangyi Yin, Jiangchao Chen, Qingyun Liu 0001 |
AAAI | 7 |
| 2026 | AdaTable: Learning to Calibrate for Robust VNF Auto-scaling under Capacity Drift
Weikang Huang, Chenkan Wang, Zhou Zhou 0007, Rong Yang 0008, Qingyun Liu 0001 |
ICIC (15) | 8 |
| 2026 | STAR: Semantic-Traffic Alignment and Retrieval for Zero-Shot HTTPS Website Fingerprinting
Yujia Zhu, Baiyang Li, Xinhao Deng 0001, Yitong Cai, Yaochen Ren, Qingyun Liu 0001 |
INFOCOM | 7 |
| 2026 | Blazer: Encrypted Video Traffic Identification for Mixed Segment Transmission Pattern based on LLMabstractDetermining the source of encrypted video traffic is an important task in network regulation. In the context of Dynamic Adaptive Streaming over HTTP (DASH), the newly emerged mixed segment transmission pattern introduces substantial difficulties for fingerprint matching, especially under adverse network conditions. To address these challenges, we propose Blazer, a DASH encrypted video traffic identification method for the mixed segment transmission pattern. First, we design a novel fingerprint that integrates video and audio segment sequences. Then, we extract the traffic fingerprint from the TLS record layer of video traffic. Finally, by observing implicit segment-mixing constraints, we design a targeted prompt and Retrieval Augmented Generation (RAG) that enables Large Language Models (LLMs) to perform fingerprint matching effectively. Across 12 network scenarios, Blazer delivers substantially better performance than the other 4 SOTA methods. Weitao Tang, Meijie Du, Die Hu 0004, Zhao Li 0010, Rong Yang 0008, Qingyun Liu 0001 |
ICMR | 7 |
| 2025 | MS-NHHO: A Swarm Intelligence Optimization Algorithm Incorporating Cognitive Science for Malicious Traffic Detection
Zhou Zhou 0007, Chengxiang Si, Qingyun Liu 0001 |
CogSci | 4 |
| 2025 | Quadruplet Fingerprinting: Onion Website Fingerprinting Through Quadruplet NetworkabstractThe onion service is designed to achieve anonymity in the Tor network. However, it is highly vulnerable to website fingerprinting (WF) attacks because websites have unique traffic patterns that can be identified. Previous server-side WF attacks have certain limitations. There is an accuracy gap of nearly 10% when comparing with the ideal modeling using server-side traces. To address this issue, we propose a remote fingerprinting attack based on quadruplet networks, named Quadruplet Fingerprinting (QF). We utilize the RP node as the remote node and construct a feature extractor with quadruplet loss. This approach enhances the generalization ability of the model. The experimental results demonstrate that higher accuracy has been achieved in both common and few-shot WF scenarios compared to the existing methods. And most importantly, the accuracy rate of the RP side has been improved by more than 10% at the client side. Can Zhao 0005, Yefeng Qin, Qingyun Liu 0001, Jinqiao Shi |
CSCWD | 5 |
| 2025 | Anya: A Novel Video Identification Attack on Media MultiplexingabstractAlthough encryption is widely employed to protect video content during transmission, protocols like DASH can still inadvertently expose critical information about the online video being watched. Attackers can potentially identify the video a user is viewing by analysing undecrypted traffic patterns. Recently, however, popular video platforms like YouTube have updated their streaming technology by utilizing audio-video multiplexing to create dynamic traffic patterns, which significantly reduce the effectiveness of previous attack methods that treat audio and video traffic as separate tracks. In this paper, we are the first to reveal the vulnerabilities about this latest streaming technology and introduce a novel attack approach named Anya. By constraining audio and video timelines, Anya constructs stable audio-video fingerprints and enhances attack accuracy and efficiency through fuzzy searching strategy. Experimental results demonstrate that Anya achieves accuracy of 0.971, 0.933 in ideal and poor network scenarios, with only one minute of traffic eavesdropping time. Finally, we propose defense strategies for streaming platform developers to protect users' privacy. Meijie Du, Lijuan Zheng, Chenyang Cui, Rong Yang 0008, Qingyun Liu 0001 |
CSCWD | 5 |
| 2025 | DygLM: Detecting Lateral Movement Threat via Continuous-time Dynamic Graph LearningabstractIn an advanced persistent threat (APT) targeting an enterprise, an adversary attempts to move through the network environment starting from their initial point, which is known as Lateral Movement (LM). Detecting LM typically uses authentication logs, which are modeled as dynamic graphs with temporal properties. Most prior research focuses on discrete snapshots, neglecting the evolving nature of the enterprise network. Recently, some researchers have proposed continuous-time graph approaches. However, they often fail to fully incorporate security domain knowledge, and suffer from detection target bias. In this paper, given the limitations of the above approaches, we propose DygLM, a framework based on continuous-time graph representation learning with suspicious LM events as the detection target. It enables self-supervised learning in an attack-free manner, adapting to the evolving network environment. By applying a well-designed encoder to the continuous-time authentication events, we encode the node-associated neighbors, edges, and time intervals to obtain topological as well as temporal representation, which is further enhanced by learning long-term correlations using Transformer. We address the detection target bias by scoring the likelihood of malicious events in the evaluation phase. Through extensive experiments on two large datasets, we demonstrate the effectiveness of DygLM in detecting LM in both transductive and inductive settings. Zhou Zhou 0007, Qingyun Liu 0001 |
CSCWD | 5 |
| 2025 | A Timestamp-Based Pacing Method for Virtual Network FunctionsabstractTraffic bursts, particularly at microsecond scale, have emerged as a significant challenge in modern networks utilizing Virtual Network Functions (VNFs). While VNFs employ various acceleration techniques to achieve hardware-comparable performance, including poll-mode APIs, batch processing, and hardware offload features, these optimizations inherently introduce microbursts into network traffic. Traditional solutions like packet pacing at host endpoints or Time-Sensitive Networking (TSN) techniques prove inadequate for VNF environments, as they either focus solely on host-side bursts or require high-precision timestamps that negate the performance benefits of packet acceleration. This paper presents an improved pacing method that maintains effectiveness while significantly reducing overhead. Our main contributions include: (1) an adaptive packet arrival time series acquisition method that adjusts sampling precision based on queue length dynamics, (2) a novel delay estimation method that minimizes forwarding latency while maintaining effective pacing, and (3) mathematical proof of our method's effectiveness. Our simulation demonstrates that the proposed solution can effectively reduce traffic bursts while introducing minimal additional latency across various deployment scenarios. Qiuwen Lu, Qingyun Liu 0001 |
CSCWD | 2 |
| 2025 | P4M3: Preventing SYN Flood Attacks on IPv6 Networks in SDN Using P4abstractSoftware-Defined Networking (SDN) centralizes control through a controller, making it a target for DDoS attacks, especially with IPv6's expansive address space, which increases vulnerability to SYN Flood attacks. P4M3(Three-layer Module architecture based on P4), a preventive solution for IPv6 SDN environments, is designed to detect and prevent SYN Flood attacks through P4 switches and the ONOS controller. The solution includes three modules: the Quick Detection and Threshold Detection modules on the data plane, and the Machine Learning Detection module on the controller. The data plane uses a danger address matching table and a Count-Min Sketch threshold method to filter suspicious SYN Flood traffic and relay packet characteristics to the controller's Machine Learning Detection module. The controller classifies packets using an ensemble learning model and issues blocking rules to the switch as needed. Results show that P4M3 reduces CPU usage by 81.92%, while maintaining low latency and enhancing SYN Flood detection accuracy, thus demonstrating its effectiveness in protecting IPv6 SDN networks from DDoS attacks. Haizhang Zhu, Qingyun Liu 0001 |
CSCWD | 6 |
| 2025 | MalTAG: Encrypted Malware Traffic Detection Framework via Graph Based Flow Interaction MiningabstractAs encrypted malware campaigns become more sophisticated, significant challenges arise in effectively detecting malicious communications when relying solely on single-stream or single-feature network traffic analysis methods. Therefore, we propose MalTAG, a graph-based flow interaction mining framework to address existing challenges. MalTAG integrates network traffic into a multi-flow correlated graph and incorpo-rates various feature types rather than relying on a single stream or single type of feature. It effectively captures the contextual correlation properties of malicious activities in both temporal and attribute dimensions. Furthermore, MalTAG achieves graph representation learning through self-supervised contrastive learning without relying on prior label knowledge. This design helps to construct more stable representations of traffic behavior to adapt to the evolving nature of malicious activities. Ultimately, it utilizes various machine learning algorithms to detect malicious traffic comprehensively. Experimental evaluations demonstrate the feasibility and good performance of MalTAG in malware detection and family classification, outperforming existing methods. Zhou Zhou 0007, Fengyuan Shi 0005, Qingyun Liu 0001 |
DSN | 5 |
| 2025 | Unraveling DoH Traces: Padding-Resilient Website Fingerprinting via HTTP/2 Key Frame Sequences
Baiyang Li, Yujia Zhu, Qingyun Liu 0001, Li Guo 0001 |
ESORICS (3) | 4 |
| 2025 | QoS-Sim: A Flexible Simulation Platform for Fine-Grained Data Center Network Monitoring
Qingyun Liu 0001 |
ICA3PP (7) | 2 |
| 2025 | GMMCL: Adaptive Concept Drift in Data Streams with Gaussian Mixture Models based on Contrastive LearningabstractClassical classification methods often fail in dynamic environments where data distributions shift over time, known as concept drift. Applications like flight delay prediction and weather forecasting require handling such dynamic data streams. Concept drift can be either virtual, affecting unconditional probability distributions, or real, affecting conditional distributions. While most research focuses on real drift, virtual drift and noise also degrade classifier performance. In this paper, we propose gaussian mixture models based on contrastive learning (GMMCL), a novel approach that integrates noise handling, contrastive learning, drift detection, and gaussian mixture models. Our method significantly enhances adaptability to noisy and drifting data streams, outperforming mainstream approaches across twelve synthetic and real-world datasets. This provides a robust solution for managing concept drift and noise in dynamic classification tasks. Hongwei Wu, Rong Yang 0008, Zhuojun Jiang, Qingyun Liu 0001 |
ICASSP | 7 |
| 2025 | COAST: Contrastive Learning with Augmented Spatio-Temporal Encoding for Next POI RecommendationabstractNext point-of-interest (POI) recommendations have garnered significant attention in industry and academia due to their crucial role in location-based social networks (LBSNs). Recent approaches have integrated sequence and geographical data to improve recommendation accuracy. However, traditional methods do not explicitly learn user similarity, which may result in suboptimal POI prediction outcomes. To address these limitations, we propose the contrastive learning with augmented spatio-temporal(COAST) model, which more effectively utilizes user check-in sequences and geographical influences. Our approach includes five techniques for augmenting check-in records and a novel Two-Head Self-Attention Encoder (THSE) to capture spatio-temporal and structural patterns. Extensive experiments on three real-world datasets demonstrate the superiority of our model compared to existing methods. Bada Xin, Zhuojun Jiang, Faqiang Liu, Rong Yang 0008, Qingyun Liu 0001 |
ICASSP | 7 |
| 2025 | APTSniffer: Detecting APT Attack Traffic Using Retrieval-Augmented Large Language ModelsabstractAdvanced Persistent Threats (APT) differ from traditional attacks by using more complex and covert strategies for long-term assaults, posing a severe threat to organizational and national security. Due to problems like the shortage of APT traffic data and encrypted traffic obfuscation, existing methods cannot accurately identify APT traffic with just a few traffic samples. To overcome the above limitation, we propose a novel encrypted APT traffic detection model, APTSniffer, which combines large language models (LLM) and retrieval-augmented technology. APTSniffer utilizes the few-shot inference and generalization abilities of large language models by converting raw traffic data into natural language inference examples understandable by the LLM. Experimental results show that, compared to other baseline models, APTSniffer exhibits SOTA performance. It achieves F1 scores above 97% on three APT datasets, making it practically applicable for APT traffic detection tasks. Chengxiang Si, Zhou Zhou 0007, Chenxu Wang 0006, Peishuai Sun, Qingyun Liu 0001 |
ICASSP | 6 |
| 2025 | Node-Centric Meta Structure Search in Heterogeneous GraphsabstractHeterogeneous graphs are increasingly used to represent complex real-world scenarios with diverse entities and interactions by meta structures. Recently, the search of meta structures is combined with graph neural architecture search to automatically extract the semantic knowledge for various tasks in heterogeneous graphs. However, prior research primarily focuses on identifying meta structures that are universally applicable across all nodes in a graph, neglecting the variations in meta structure selection that arise from the unique features and topology of individual nodes. To address this challenge, we introduce a Node-Centric approach to search Meta Structures in heterogeneous graphs (NC-MS for short). NC-MS implements a node level method that discover meaningful meta structures tailored to each node, capturing subtle differences in meta structure choices between nodes and providing nuanced identification. Additionally, NC-MS utilizes an efficient and differentiable network to enhance operational efficiency. Empirical studies across three real-world datasets validate the superiority of NC-MS, demonstrating its ability to outperform existing models in heterogeneous graph neural networks. Xiaoou Zhang, Yang Gao 0024, Yang Aron Liu, Yujia Zhu, Chuan Zhou 0001, Peng Zhang 0001, Qingyun Liu 0001, Hongyang Chen 0001 |
ICASSP | 8 |
| 2025 | MoHGNN: Enhanced Heterogeneous Graph Neural Network via Metapath OptimizationabstractIn this paper, we propose a novel heterogeneous graph neural networks (HGNNs) model that addresses two major limitations of existing metapath-based methods: (1) Defining suitable metapaths requires professional knowledge in the special domain. (2) The neighbor nodes of the target node also play crucial roles for embedding, but common methods often overlook this factor. Specifically, our model optimizes metapath selection by evaluating all potential metapaths through edge-based PageRank. Then, through aggregation of intra- and intermetapaths obtain metapath embeddings. Next, we gather neighboring nodes that are highly correlated with the target node and obtain the embeddings of node type to optimize metapath embeddings for the next step. Finally, we employ a two-layer attention architecture to acquire the embedding of target node. The experiments on three real-world datasets demonstrate that our method outperforms state-of-the-art methods in handling downstream tasks such as node classification and clustering. Taiyao Zhang, Qingyun Liu 0001 |
ICASSP | 4 |
| 2025 | Heterogeneous Graph Anomaly Detection with Graph Wavelet TransformerabstractGraph Anomaly Detection (GAD) identifies deviant patterns including anomalous nodes, edges, and subgraphs in graph data, with significant applications in social networks, cybersecurity, and financial risk control. While spectral methods have proven effective for homogeneous graph anomaly detection, their application to heterogeneous graphs remains challenging due to structural complexity and semantic richness. Existing heterogeneous graph anomaly detection methods either rely on manually designed meta-paths or decompose the graph into homogeneous subgraphs, leading to limited flexibility or loss of structural integrity. To address these limitations, we propose the Graph Wavelet Transformer (GWT), a novel spectral-based approach that integrates global graph properties and spectral analysis without requiring meta-path information. GWT employs a three-stage process: heterogeneous-to-homogeneous graph conversion, global dependency modeling via graph transformers, and spectral-aware feature enhancement focusing on frequency band components. Extensive experiments on multiple benchmarks demonstrate that GWT significantly outperforms ten baseline methods, providing a new paradigm for heterogeneous graph anomaly detection that preserves structural completeness while achieving computational efficiency. Xiaoou Zhang, Chuan Zhou 0001, Yang Aron Liu, Shuai Zhang 0007, Peng Zhang 0001, Yujia Zhu, Qingyun Liu 0001 |
ICDM | 7 |
| 2025 | PAWS: Passive Concept Drift Adaptation Based on Instance Weighting and Subspace Alignment in Data Stream
Hongwei Wu, Rong Yang 0008, Zhuojun Jiang, Qingyun Liu 0001 |
ICIC (20) | 7 |
| 2025 | Pioneer: Encrypted Video Traffic Identification for Mixed Transmission of Video-Audio SegmentsabstractThe spread of harmful content via video has made video traffic identification crucial for network regulation. In the new transmission mode, audio and video segments are mixed to combine into video chunks. However, in poor networks, such combination is unstable, and video chunks may be lost and retransmitted. To address these challenges, this paper proposes Pioneer, an encrypted video traffic identification method for mixed transmission of audio and video segments. We introduce a precise video chunk reconstruction method for video traffic encrypted by both TLS and QUIC. Additionally, we propose Pseudo-Siamese Attention-Convolutional Network (PSACN) to calculate the similarity between traffic and video, leveraging contrastive learning during training to mitigate the impact of poor networks. Pioneer significantly improves accuracy compared with state-of-the-art (SOTA) methods under various network environments. Notably, this is the first study to address this emerging new transmission mode. Weitao Tang, Taizhong Xu, Meijie Du, Die Hu 0004, Qingyun Liu 0001 |
ICME | 6 |
| 2025 | BACKTRACKER: A Novel Background Traffic Identification System For Mobile AppsabstractAs many mobile apps generate substantial background traffic without active user interaction, network operators face increasing challenges in traffic management and analysis. However, existing approaches lack systematic methods for analyzing and identifying app background traffic. This paper presents BACKTRACKER, a novel background traffic identification system for mobile apps. To establish a reliable background traffic dataset, we design a semi-automated traffic collection framework integrating Android device with network traffic interception. We propose a URL similarity algorithm based on Levenshtein distance for accurate foreground-background traffic differentiation. Furthermore, we develop a hierarchical recognition model that combines statistical stability with deep learning expressiveness, integrating data augmentation, multi-head attention and BiLSTM networks for robust feature learning. Our extensive evaluation shows that BACKTRACKER significantly outperforms baseline methods, achieving 98.36% F1-score in background traffic identification. The results demonstrate BACKTRACKER’s effectiveness in background traffic analysis under encrypted network environments. Yuyi Liu, Yitong Cai, Rong Yang 0008, Qingyun Liu 0001 |
IJCNN | 6 |
| 2025 | PORTIA: A Multi-Granularity APT Detection Model Based on Provenance GraphsabstractAdvanced Persistent Threats (APTs) have become a major cybersecurity threat due to their stealthy attack methods and long latency periods. Traditional signature-based detection struggles to detect novel attacks, and while unsupervised methods using Graph Neural Networks (GNNs) can model system behavior, they face challenges in handling large-scale provenance graphs and accurately incorporating system operational states for detection.This paper presents PORTIA, a multi-granularity APT detection model based on graph representation learning. PORTIA constructs provenance graphs by integrating temporal information from audit logs and uses a graph mask autoencoder to model normal system behavior, detecting anomalies through embedding shifts. As an unsupervised model, PORTIA can swiftly identify anomalous system states without relying on attack signatures, achieving fine-grained detection by incorporating system operational states. Evaluations on multiple datasets show that PORTIA detects both standard and APT attacks with high precision, outperforming existing detection systems. Haoqiang Wang, Zhou Zhou 0007, Chengxiang Si, Qingyun Liu 0001 |
IJCNN | 7 |
| 2025 | MOLE: Provenance Graph Generation Framework Based on LLM PromptingabstractIn the increasingly complex landscape of cyber-attacks, logs have become a critical source of data for detecting system threats. Currently, most log-based detection systems rely on converting audit logs into provenance graphs during the process of attack investigation. However, this construction process is still heavily dependent on manually written code with regular expressions tailored to each specific log type. In this paper, we propose MOLE, a provenance graph generation framework based on prompting with large language models (LLMs). Unlike traditional approaches, MOLE does not rely on prior knowledge and is adaptable to diverse types of log data. The framework automatically generates provenance graph extraction templates through instruction generation and parses logs locally to produce the final provenance graph.MOLE leverages the log patterns and structures learned by LLMs from large-scale data during training. As a result, tasks that previously required several days of manual coding to generate a provenance graph can now be completed in just a few minutes. Furthermore, when processing 50 million log entries, the entire provenance graph generation process consumed only 20k tokens. Haoqiang Wang, Zhou Zhou 0007, Chengxiang Si, Qingyun Liu 0001 |
IJCNN | 6 |
| 2025 | FlowMiner: A Powerful Model Based on Flow Correlation Mining for Encrypted Traffic Classification
Chengxiang Si, Zhenyu Cheng 0001, Chenxu Wang 0006, Jiang Xie 0004, Peishuai Sun, Qingyun Liu 0001 |
INFOCOM | 8 |
| 2025 | FCC: Fast Fair Congestion Control in Data Center Networks with RDMAabstractCongestion control (CC) is a critical technology for improving the performance of data center networks with Remote Direct Memory Access (RDMA). Existing end-to-end CCs find appropriate congestion signals and calculate the congestion window size based on the Additive Increase Multiplicative Decrease (AIMD) model. Custom host protocol stacks and Inband Network Telemetry (INT) enable CCs to leverage congestion signals that reflect more comprehensive network status. However, these CCs still face a long time for draining up queues, link underutilization, and frequently tuning several hyperparameters. We attribute these issues to the limitations of the AIMD model and the reliance on time-varying congestion signals. An ideal CC should drain queues and fully utilize link capacity at each decision point. To overcome these shortcomings, we propose a novel congestion control algorithm called FCC, which achieves per-flow fairness and efficiency within one control loop without tuning any hyperparameter. FCC collects per-hop flow information through INT, optimizes link utilization based on a fair bandwidth allocation model, and drains the queue by a time-invariant congestion signal. Our evaluations show that FCC reduces the flow completion time (FCT) of short flows by up to 86 %. Qingyun Liu 0001 |
IPCCC | 2 |
| 2025 | DarkDC:Analysis of Scanning Behavior Subject Portrait Based on Deep Clustering
Xinyi Ji, Yixiang Zhao, Zhou Zhou 0007, Qingyun Liu 0001 |
ISCC | 5 |
| 2025 | MV-TFNet: Malware Traffic Detection Based on Multi-View Flow Sequence and Time-Frequency Feature FusionabstractThe proliferation of malware’s command-andcontrol (C2) communications poses significant challenges to traditional network security defence as adversaries increasingly leverage advanced evasion tactics such as dynamic infrastructure rotation and traffic obfuscation. However, previous works suffer from inadequate temporal modeling, session-completion assumption, and inflexible data augmentation. To address these limitations, this paper proposes MV-TFNet, a novel C2 traffic detection framework that synergizes multi-view flow sequence analysis with temporal-frequency feature fusion. By decomposing host-server interactions into upstream, downstream, and bidirectional views, we employ a dual-path encoder to capture localized temporal dependencies via dilated convolutions and global frequencydomain patterns through spectral analysis. A contrastive learning paradigm, enhanced by label-guided adaptive data augmentation, further optimizes robustness against adversarial noise and distribution shifts. Moreover, MV-TFNet has a protocol-agnostic design, eliminating the dependency on plaintext information or protocol-specific parsing. Extensive evaluations on the MTA and DoHBrw datasets demonstrate MV-TFNet’s superiority over state-of-the-art methods, achieving an F1-score range from 0.94 to 0.99 and a false positive rate below 0.05. Zhou Zhou 0007, Qingyun Liu 0001 |
ISCC | 5 |
| 2025 | FLASK-Sketch: Identifying Sparse Superspreaders in High Speed NetworkabstractA sparse superspreader is a host that establishes connections to a large number of distinct destinations while transmitting only a small number of packets. This phenomenon is frequently observed in various network activities, including network scanning, the early propagation of worm viruses, and spam sending. However, existing methods often fail to simultaneously capture the high spread and low-frequency characteristics of these hosts. In this paper, we propose FLASK-Sketch, a realtime approach for detecting sparse superspreaders. The core idea of FLASK-Sketch is to track both the frequency of a host’s appearances and the number of its distinct connections, and then to integrate their ratio into a scoring mechanism. We evaluate FLASK-Sketch against two baseline methods, SpreadSketch+CM and ExtendedSketch+CM. Experimental results show that FLASK-Sketch improves the F1 score by 28 % and 50 %, reduces the Average Relative Error (ARE) by 54 % and 60 %, and achieves throughput gains of 20 % and 39 % compared to these strawman solutions. Rong Yang 0008, Qingyun Liu 0001 |
ISCC | 5 |
| 2025 | Optimized DFA-Based URL Filtering for P4 Programmable Switches
Yike Zhao, Kedong Liu, Qingyun Liu 0001 |
KSEM (4) | 7 |
| 2025 | Leveraging Cross-Layer Network Probing to Detect Stealth ServicesabstractStealth services have become increasingly popular due to growing demand for privacy protection. To avoid detection by active probing, many have adopted probe-resistant strategies. In this work, we design a suite of carefully crafted probes to expose their hidden vulnerabilities. Through detailed analysis of the corresponding responses, we find that despite the defensive measures implemented, certain implicit information, such as protocol stack fingerprints, can still serve as strong indicators. We present SSChecker, a detection system that combines cross-layer probing with a classification model inspired by Information Bottleneck Theory to address the data sparsity issue inherent in active probing. Our experiments on real-world datasets show that SSChecker outperforms existing methods, including leading industrial detection engines. Our findings demonstrate that current probe-resistant strategies of stealth services remain insufficient. To strengthen privacy protection, we further propose mitigation strategies that help stealth services enhance their ability. Jiangyi Yin, Chenxu Wang 0006, Zhao Li 0010, Zhuojun Jiang, Jiangchao Chen, Dongfang Hao, Qingyun Liu 0001 |
TrustCom | 7 |
| 2025 | Broken Chains: An Empirical Analysis of DNS Resolution in IPv6-only EnvironmentsabstractThe global transition to IPv6 is impeded by failures within the Domain Name System (DNS), where the mere presence of an AAAA record does not guarantee a domain’s resolvability in IPv6-only environments. In this paper, we presents a comprehensive measurement study analyzing the complete DNS dependency chain—including parental, delegation, and alias dependencies—to reveal the true state of IPv6 resolvability. We introduce 6ChainChecker, a lightweight tool developed for this analysis. Our analysis of Tranco top domains reveals a critical discrepancy: 7.68% of domains with published AAAA records are nevertheless unresolvable from a strict IPv6-only stack due to structural failures in their dependency chains. To understand the broader landscape of failures, we find that while the absence of an AAAA record is the most common reason for unresolvability (58.0%), a substantial portion of failures stem from broken dependency paths, including delegation (17.1%) and alias chain (24.8%) issues. Crucially, we also identify significant dependency concentration, where the non-compliance of a few critical infrastructure zones creates cascading failures for hundreds of their dependent domains. These findings demonstrate that upstream infrastructure, not just endpoint configuration, is a significant impediment to the IPv6 transition. Our tool and dataset are publicly available to foster further research. Yujia Zhu, Baiyang Li, Qingyun Liu 0001 |
TrustCom | 6 |
| 2025 | HOLMES & WATSON: A Robust and Lightweight HTTPS Website Fingerprinting through HTTP Version ParallelismabstractWebsite Fingerprinting (WF) is a traffic analysis technique that aims to identify websites visited by users through the analysis of encrypted traffic patterns.Existing approaches often exhibit limited robustness against network variability and concept drift, resulting in significant performance degradation under real-world HTTPS conditions.Moreover, these methods typically require large-scale training datasets and substantial computational resources, which further increases the complexity of deployment.In this paper, we propose HOLMES, a novel approach that exploits HTTP version parallelism to extract enhanced application-layer features.These features, including the number of web resources transmitting in various HTTP versions, expose up to 4.28 bits of information-surpassing 98% of previously reported features and demonstrate increased stability across varying network conditions.Complementary to this, we introduce WATSON, a lightweight classification method based on lazy learning, which substantially reduces the dependency on large training datasets.To further enhance the identification accuracy, we incorporate two fingerprint-specific distance metrics that ensure high intra-class similarity.Our experimental evaluation demonstrates that HOLMES & WATSON significantly enhance both robustness and efficiency, achieving an average accuracy of 87.7% with only a single sample per website, marking an improvement of over 15% compared to state-of-the-art methods. Yujia Zhu, Baiyang Li, Peishuai Sun, Xinhao Deng 0001, Qingyun Liu 0001 |
WWW | 7 |
| 2024 | 6GAI: Active IPv6 Address Generation via Adversarial Training with Leaked InformationabstractGlobal IPv6 scanning has always been a challenge for researchers because of the limited network speed and computational power. In this paper, we introduce 6GAI to implement more efficient target address generation. 6GAI is built with Generative Adversarial Net (GAN) integrated with Convolutional Bottleneck Attention Module (CBAM). 6GAI allows the discriminative net to leak generated address’s high-level features extracted by CBAM to the generative net, while the generative net incorporates such informative signals into all generation steps through an additional Manger module, which takes the extracted features of current generated address nybbles and outputs a latent vector to guide the Worker module for active IPv6 address generation. This work outperformed the state-of-the-art target generation algorithms on two datasets including one public dataset and one independently collected dataset. Liang Jiao, Yujia Zhu, Wen-Xiu Zhang, Qingyun Liu 0001 |
CSCWD | 6 |
| 2024 | PFTB: A Prediction-Based Fair Token Bucket Algorithm based on CRDTabstractIn today’s rapidly evolving network landscape, an increasing number of applications are finding deployment on cloud-based computing platforms. With network traffic growing at an accelerated pace, the rational control of bandwidth utilization by these applications has emerged as a formidable technical challenge. Existing distributed rate limiting algorithms, while capable of enforcing stringent rate limits, often come at the cost of significant bandwidth wastage and lack comprehensive discussions on the global fairness of applications. In response to these challenges, we introduce an innovative distributed rate limiting algorithm termed the Prediction-based Fair Token Bucket, and it achieves equitable rate limiting among applications while optimizing the utilization of the entire network’s capacity. We introduce a TK-CRDT module based on a fair token bucket mechanism, which is integrated into our rate limiting algorithm. Through a comparative analysis against state-of-the-art rate limiting schemes, our algorithm enhances the excessive rate limiting metric by 91% and increases the Jain’s fairness index by 39%. Luting Zhang, Qingyun Liu 0001, Rong Yang 0008 |
CSCWD | 4 |
| 2024 | From Fingerprint to Footprint: Characterizing the Dependencies in Encrypted DNS Infrastructures
Baiyang Li, Yujia Zhu, Qingyun Liu 0001, Li Guo 0001 |
ESORICS (2) | 6 |
| 2024 | ProxyKiller: An Anonymous Proxy Traffic Attack Model Based on Traffic Behavior Graphs
Zhenyu Cheng 0001, Chenxu Wang 0006, Peishuai Sun, Jiang Xie 0004, Qingyun Liu 0001 |
ESORICS (2) | 7 |
| 2024 | AlterCell Attack: Exploiting a Logic Vulnerability in Tor Cell Integrity ValidationabstractThe hidden service is used to protect the anonymity of receivers in the Tor network, but it is often exploited for malicious purposes and becomes a breeding ground for crime. To protect the hidden service from abuse, we deeply analysed the Tor specifications and summarised a finite-state machine for Tor cell transmission protocol, deriving a new logical vulnerability of the Tor cell integrity validation. This vulnerability allowed the plaintext field of cell_command in the cell to be arbitrarily modified. Therefore, we proposed the AlterCell attack against potentially malicious hidden services, by introducing malformed cells whose fields of cell_command were modified. In the experiments, when our node acted as the guard for the target hidden service, the AlterCell attack was able to de-anonymise the hidden service if the responsible HSDir was under our control. Otherwise, our attack could block the descriptor publishing process or the client access process, to carry out a denial-of-service attack on the hidden service. Can Zhao 0005, Baiwei Duan, Qingyun Liu 0001, Jinqiao Shi |
HPCC | 5 |
| 2024 | A Targeted Adversarial Attack Method for Multi-Classification Malicious Traffic DetectionabstractLeveraging deep learning to detect malicious network traffic is a crucial technology in network management and network security. However, deep learning security has raised concerns among scholars. In this work, we explore executing targeted adversarial attacks for multi-classification malicious traffic detection with limited interactions. Specifically, we constrain the number of interactions with detection and employ a hop-skip-jump attack (HSJA) to generate a small number of adversarial samples. These adversarial samples are then heuristically used to train a generative adversarial network (GAN) to generate a substantial quantity of adversarial samples. Experiments demonstrate that our method is more adversarial and displays a certain degree of generalization compared with other methods. Peishuai Sun, Chengxiang Si, Zhenyu Cheng 0001, Qingyun Liu 0001 |
ICASSP | 6 |
| 2024 | Meta Structure Search for Link Weight Prediction in Heterogeneous GraphsabstractRecently link weight prediction has attracted an increasing research interest due to its merits in quantifying the strength between nodes within a graph. Nonetheless, current link weight prediction methods focus solely on graph topology, disregarding node feature information embedded in graphs. In real-world applications, we often collect heterogeneous graph data where multiple types of nodes linked by multiple types of edges are available for analysis, and it is essential and challenging to quantify the proximity of different types of nodes. To solve this challenge, we present a new model for Heterogeneous Graph Link Weight Prediction (HLWP for short). In HLWP, message passing in heterogeneous graph neural networks is described as a meta structure, which can be effectively designed by Differentiable Neural Architecture Search (DARTS) algorithms. Thus, HLWP can enhance the message passing in heterogeneous graphs by DARTS. In addition, HLWP employs a perturbation-based algorithm to enhance stability and precision. Through empirical experiments conducted on three real-world datasets, we demonstrate that HLWP achieves accurate predictions of link weights. Our results highlight the superiority of HLWP over existing methods for link weight prediction and baseline GNN models in terms of accurately predicting link weights within heterogeneous graphs. Xiaoou Zhang, Yang Gao 0024, Yang Aron Liu, Yujia Zhu, Peng Zhang 0001, Chuan Zhou 0001, Qingyun Liu 0001, Hongyang Chen 0001 |
ICASSP | 7 |
| 2024 | Failed Yet Stored: A First Look at DNS Negative CachingabstractCaching is a critical method for enhancing the efficiency and the security of the Domain Name System (DNS). Initially, only successful domain name resolution results were cached. To mitigate failures in DNS transactions (e.g. NXDomain), the IETF proposed standards, further developed into RFC 9520 as of December 2023. In addition to the basic implementation, RFC 9520 standardizes more sophisticated forms of negative caching. This new standard aims to reduce redundant query retries in DNS traffic and protect resolvers from Denial of Service (DoS) attacks.In this study, we present a comprehensive examination of the specific implementations of Negative Caching in resolvers. We designed and validated a method for measuring negative caching and conducted experiments on 44 public resolvers, including their Do53, DoH, and DoT interfaces. Our findings indicate that while public resolvers generally implement various types of negative caching, some exhibit unexpected cache handling behaviors when encountering specific negative responses. Additionally, we discovered that most public resolvers modify the TTL value of negative responses before returning them to clients. Despite the lack of explicit TTL values for newly specified negative responses, we devised a method to approximate the default TTL values used by public resolvers. Meng Zeng, Yujia Zhu, Baiyang Li, Qingyun Liu 0001, Binxing Fang |
IPCCC | 5 |
| 2024 | P4-FILB: Stateless Load Balancing Mechanism in Firewall and IPv6 Environment with P4abstractLoad balancers are critical infrastructure in modern distributed systems, and their main function is to evenly distribute client traffic to multiple servers for high availability and scalability. However, current load balancers face challenges in balanced resource utilization. To ensure per-connection consistency, load balancers typically assign subsequent requests from the same client to the same server. While this strategy simplifies session management and reduces the overhead of state synchronization, it also leads to uneven resource utilization.In this paper, we propose P4-FILB, a stateless load balancing scheme that achieves uniform load distribution among servers so that the resources of each server can be optimally balanced when receiving a large number of data streams. The key idea behind the implementation of P4-FILB is that a load balancer between the client and the server keeps track of the size of the packets that are currently being processed by the different servers. It also uses segment routing in IPv6, which uses an ordered list called "segments" to direct packets to servers that are currently utilizing a relatively small amount of resources. And this paper also proposes a two-tier load policy in firewall environment. According to the intelligent routing policy of the firewall, the packets are distributed to different groups of servers, and then the load balancer carries out further request allocation to improve the allocation of network resources. After evaluating the performance of P4-FILB, it is shown that the load balancing scheme performs better than the existing studies in terms of load balancing state between servers. It also reduces the overhead caused by extra packets by means of packets carrying connection information. Haizhang Zhu, Zhou Zhou 0007, Chengwei Peng, Rong Yang 0008, Qingyun Liu 0001 |
IPCCC | 8 |
| 2024 | Measuring Encrypted DNS Service with TLS1.3 Support over IPv6abstractThe Encrypted Domain Name System (DNS) and Encrypted Server Name Indication (ESNI) are recently proposed to enhance network security and privacy protection; we refer to these schemes collectively as domain name encryption technologies. Previous research has shown that the destination IP address accessed by the user cannot be associated with common web services such as websites because a large number of websites are hosted through cloud or CDN over IPv6. However, encrypted DNS, as an internet infrastructure service, is typically deployed independently by the service provider rather than hosted through cloud or CDN. In this paper, we propose a method to discover the unique service provider of encrypted DNS resolvers on a large-scale encrypted traffic with TLS1.3 support over IPv6. The model utilizes a Siamese Network to determine whether two IPv6 resolver addresses belong to the same service provider of encrypted DNS, even if the DNS query is protected by ESNI. Through a comprehensive analysis of two real-world datasets, which include encrypted DNS data and common web data, we find that the implementation of TLS1.3, especially ESNI, does not impact the association of encrypted DNS server addresses. Our model achieves an accuracy rate of 95.29%. Liang Jiao, Wen-Xiu Zhang, Tianyu Cui, Yujia Zhu, Qingyun Liu 0001 |
ISCC | 7 |
| 2024 | NFVDC: A flexible and efficient NFV platform for dynamic service chainingabstractThis paper presents the design and implementation of a novel Network Function Virtualization (NFV) platform, NFVDC, specifically for dynamic service chaining. Unlike traditional service chaining methods, NFVDC supports real-time, reconfigurable service chains that accommodate service functions dynamically joining or leaving the chain. By integrating an advanced traffic steering method, optimized CPU scheduling, and a causal message-passing mechanism, NFVDC addresses challenges in deploying dynamic service chaining, such as traffic steering, resource scheduling, and session state management. Our evaluation demonstrates that NFVDC’s forwarding performance nearly doubles compared to existing solutions when extending service chain lengths to six. Qiuwen Lu, Qingyun Liu 0001, Binxing Fang |
ISPA | 2 |
| 2024 | No Source Code? No Problem! Demystifying and Detecting Mask Apps in iOS
Lingjing Yu, Qingyun Liu 0001, Bo Luo |
ICPC | 4 |
| 2024 | Zenith: Real-time Identification of DASH Encrypted Video Traffic with DistortionabstractSome video traffic carries harmful content, such as hate speech and child abuse, primarily encrypted and transmitted through Dynamic Adaptive Streaming over HTTP (DASH). Promptly identifying and intercepting traffic of harmful videos is crucial in network regulation. However, QUIC is becoming another DASH transport protocol in addition to TCP. On the other hand, complex network environments and diverse playback modes lead to significant distortions in traffic. The issues above have not been effectively addressed. This paper proposes a real-time identification method for DASH encrypted video traffic with distortion, named Zenith. We extract stable video segment sequences under various itags as video fingerprints to tackle resolution changes and propose a method of traffic fingerprint extraction under QUIC and VPN. Subsequently, simulating the sequence matching problem as a natural language problem, we propose Traffic Language Model (TLM), which can effectively address video data loss and retransmission. Finally, we propose a frequency dictionary to accelerate Zenith's speed further. Zenith significantly improves accuracy and speed compared to other SOTA methods in various complex scenarios, especially in QUIC, VPN, automatic resolution, and low bandwidth. Zenith requires traffic for just half a minute of video content to achieve precise identification, demonstrating its real-time effectiveness. Weitao Tang, Meijie Du, Die Hu 0004, Qingyun Liu 0001 |
ACM Multimedia | 5 |
| 2024 | LightRL-AD: A Lightweight Online Reinforcement Learning Approach for Autonomous Defense against Network AttacksabstractWith the rapid growth of the Internet, network structure has become increasingly complex, leading to more diverse and impactful network attacks. Traditional methods of detecting and defending against network attacks struggle with increasingly complex situations due to human decision-making processes. Recent research has started exploring autonomous defense mechanisms for network attacks within software-defined network (SDN) environments. However, these methods typically employ complex reinforcement learning techniques, making them challenging to implement in online deployment environments. In this paper, we propose LightRL-AD, a lightweight online reinforcement learning approach for autonomous defense against network attacks in SDN. LightRL-AD integrates a machine learning-based Intrusion Detection System (IDS), a reinforcement learning-based Intrusion Prevention System (IPS), and a Moving Target Defense (MTD) mechanism. The ML-based IDS classifies network flows into categories such as malicious or benign, while the RL-based IPS utilizes the SARSA algorithm to determine and execute appropriate defensive actions, ensuring robust network security. We employ specific hardware and software to establish a simulated SDN network for our experiments. And we implement LightRL-AD in the network and evaluate its performance. Experimental results demonstrate that LightRL-AD performs better to defend against slow-rate DDoS attacks autonomously. Fengyuan Shi 0005, Zhou Zhou 0007, Qingyun Liu 0001, Xiuguo Bao |
TrustCom | 7 |
| 2024 | SCENE: Shape-based Clustering for Enhanced Noise-resilient Encrypted Traffic ClassificationabstractNetwork traffic classification is critical in network management, quality of service optimization, and security monitoring. However, most existing methods for encrypted traffic classification rely heavily on supervised learning, requiring large amounts of labeled data, and struggle to perform effectively in complex and dynamic network environments. To address these limitations, we propose a novel unsupervised method for encrypted traffic classification, which analyzes byte rate variations to capture traffic behavior patterns. Our approach does not require prior knowledge or large volumes of labeled data, enabling adaptive processing of encrypted traffic in complex network conditions. Specifically, we introduce a noise-resilient shape-line extraction method that preserves core behavioral characteristics of traffic; we design a multidimensional feature extraction strategy that analyzes both uplink and downlink features; and we propose an unsupervised classification algorithm that combines shape-based density clustering with a feature assignment strategy. This algorithm overcomes the limitations of traditional methods, such as the need for predefined cluster numbers, and can classify unknown traffic patterns. We validate our method on five real-world traffic datasets with differing levels of openness, demonstrating its remarkable robustness and accuracy in encrypted traffic classification tasks, thereby greatly enhancing the precision and stability of service classification. Meijie Du, Mingqi Hu, Zhao Li 0010, Qingyun Liu 0001 |
TrustCom | 5 |
| 2024 | LayyerX: Unveiling the Hidden Layers of DoH Server via Differential FingerprintingabstractAs a rapidly developing DNS security enhancement technology, DoH(DNS over HTTPS) is gaining popularity among people. It allows users to quickly set up a DoH server by combining several components which create a layered structure. However, the multi-layer setup, which involves both HTTPS and DNS protocols, makes internal structural details more difficult to be detected. To address this issue, we propose a method that utilizes cross-protocol fingerprinting and analysis techniques, which is capable of identifying various components of multi-layer DoH servers with a focus on the underlying differences within protocol. Using this approach, we developed LayyerX, a system for detecting multi-layer DoH servers. Finally, through experiments and large-scale measurements in the wild, we showcased LayyerX’s outstanding capabilities and presented a meaningful structural overview of multi-layer DoH server. Yunyang Qin, Yujia Zhu, Linkang Zhang, Baiyang Li, Qingyun Liu 0001 |
TrustCom | 6 |
| 2024 | Knock-Knock: De-Anonymise Hidden Services by Exploiting Service Answer VulnerabilityabstractThe hidden service, devised by the Tor Project, serves to protect receiver anonymity. However, to address the potential abuse of hidden services, this paper introduces the “Knock-Knock attack,” a novel de-anonymization method that utilizes watermark to enable an attacker to conduct an attack with control over only the client and the guard relays. The key of our approach is manipulating the number of RELAY_COMMAND_BEGIN cells and RELAY_COMMAND_CONNECTED cells to construct flexible and robust watermarks during the Tor protocol handshake, thus facilitating parallel de-anonymization of hidden services while remaining insensitive to the network state. Empirical experiments demonstrate that our attack boasts a 100% true positive rate and 0% false positive rate. Additionally, we propose a theoretical framework to guide the optimal encoding form of watermark, leading to a notable 3.6 times improvement in speed compared to prior works. Lastly, we present a method to mitigate watermark attacks and report the design flaw to the Tor Project. Muqian Chen, Can Zhao 0005, Qingyun Liu 0001, Jinqiao Shi |
WCNC | 5 |
| 2024 | Identifying VPN Servers through Graph-Represented BehaviorsabstractIdentifying VPN servers is a crucial task in various situations, such as geo-fraud detection, bot traffic analysis and network attack identification. Although numerous studies that focus on network traffic detection have achieved excellent performance in closed-world scenarios, particularly those methods based on deep learning, they may exhibit significant performance degradation due to changes in network environment. To mitigate this issue, a few studies have attempted to use methods based on active probing to detect VPN servers. However, these methods still have two limitations. They cannot handle situations without probing responses and are limited in applicability due to their focus on specific VPNs. In this work, we propose VPNChecker, which utilizes the graph-represented behaviors to detect VPN servers in real-world scenarios. VPNChecker outperforms existing methods in four offline datasets. The results from our datasets, containing multiple different VPNs, indicate that VPNChecker has better applicability. Furthermore, we deploy VPNChecker in an Internet Service Provider's (ISP) environment to evaluate its effectiveness. The results show that VPNChecker can improve the coverage of sophisticated detection engines and serve as a complement to existing methods. Chenxu Wang 0006, Jiangyi Yin, Zhao Li 0010, Qingyun Liu 0001 |
WWW | 6 |
| 2024 | HSDirSniper: A New Attack Exploiting Vulnerabilities in Tor's Hidden Service Directories
Zhiyang Teng, Yue Gao 0003, Qingyun Liu 0001, Jinqiao Shi |
WWW | 5 |
| 2023 | DualDNSMiner: A Dual-Stack Resolver Discovery Method Based on Alias Resolution
Dingkang Han, Yujia Zhu, Liang Jiao, Dikai Mo, Qingyun Liu 0001 |
CollaborateCom (3) | 7 |
| 2023 | Long-Short Terms Frequency: A Method for Encrypted Video Streaming IdentificationabstractNowadays, with the vigorous development of self-media services, more and more individual users upload videos freely. While bringing goodness, it also inevitably brings evil. Therefore, it is particularly necessary to identify and supervise illegal videos through network stream. However, many video streaming services, such as YouTube, have applied encryption to protect users’ privacy, which makes it more difficult to analyze network stream. Many researches show that DASH (Dynamic Adaptive Streaming over HTTP) will leak information about video segmentation, which is related to the video content. Consequently, it is possible to analyze the content of encrypted video stream without decryption. Previous studies have proposed a series of encrypted video identification methods based on this. However, most of them need to wait for a long video playback time, such as more than 10s, or even wait for the entire video playback to complete the identification. In this paper, we propose a fast, lightweight, and accurate method named Long-Short Terms Frequency(LSTF) for online encrypted video identification. Experiments have proved that compared with the state-of-the-art, our method has advantages in both speed and accuracy, and even if CDN switching occurs during the video playback, it still has a high identification accuracy. Meijie Du, Minchao Xu, Kedong Liu, Weitao Tang, Lijuan Zheng, Qingyun Liu 0001 |
CSCWD | 6 |
| 2023 | A Robust and Accurate Encrypted Video Traffic Identification Method via Graph Neural NetworkabstractThe explosive growth of video traffic has brought major challenges for network providers to improve user experience. On account of traffic encryption, network providers need to identify encrypted video traffic first before adopting optimization approaches to them. Traditional encrypted video traffic identification methods try to reveal the pattern of video traffic by using statistical features, which are not robust enough in different network environments. Some sophisticated graph-based methods recently have shown their advantages for encrypted traffic identification. However, these works lack optimization when it comes to the video streaming scenario. Inspired by these works, we propose GraphV, a GNN-based approach for identifying encrypted video traffic. Specifically, we construct an information-rich graph structure enhanced by unique features of video transmission. Then the embedding representation of each graph can be obtained through a Bi-LSTM layer added to all the sequential nodes embedding on this graph. The experiments on a well-known dataset and two open-world datasets from different network environments we collected show that GraphV outperforms the existing methods, especially on the generalization ability of the model. Zhao Li 0010, Jiangchao Chen, Xiaoqing Ma, Meijie Du, Qingyun Liu 0001 |
CSCWD | 6 |
| 2023 | SpoofingGuard: A Content-agnostic Framework for Email Spoofing Detection via Delivery GraphabstractEmail spoofing is an effective attack vector for infiltrating companies and organizations. Traditional detectors are primarily based on the content of emails, but they ignore the frequent contextual changes. The blacklist-based solutions commonly used in the industry suffers from latency issues. Additionally, there are protocol-based solutions, such as SPF, DKIM, etc., but their adoption rates are unsatisfactory. To address these issues, this work presents a new framework named SpoofingGuard that detects email spoofing based on graph representation learning. As SpoofingGuard extracts important delivery path information related to the email service infrastructure from email headers, it is completely content-agnostic, and is expected to be more robust in the face of complex content variations. Finally, the evaluation results on two public datasets show that SpoofingGuard can achieve 99.51% precision and less than 0.5% false positive rate, demonstrating its effectiveness and advancement. Yujia Zhu, Xiaoou Zhang, Zhen Jie, Qingyun Liu 0001 |
CSCWD | 5 |
| 2023 | Detecting Fake-Normal Pornographic and Gambling Websites through one Multi-Attention HGNNabstractThe rapid development of pornographic and gambling websites, fueled by the widespread abuse of information technology, has become a growing concern. They pose a serious threat to the physical and mental health of children and can also endanger personal property. Therefore, it is necessary to detect them. However, pornographic and gambling websites become more and more tricky, which shows fake-normal to evade censorship and challenges traditional content-based detection methods. Therefore, it is essential to rely on information about relationships between websites.We propose HMAN, one Multi-Attention Heterogeneous Graph Neural Network (HGNN) model to detect pornographic and gambling websites by integrating content features and structural information, even if they present fake-normal. By one multi-attention mechanism consisting of explicit weight, self-attention and attention mechanism, content features can be selectively utilized with the assistance of structural information. The experimental results show that our method achieves the best 95.1% Macro-Avg-F1 and outperforms all baselines. We also illustrate that all extracted metapaths do contribute to the detection, where the hyperlink, title/meta terms and IP address are relatively important. Xiaoqing Ma, Chao Zheng 0001, Zhao Li 0010, Jiangyi Yin, Qingyun Liu 0001, Xunxun Chen |
CSCWD | 5 |
| 2023 | AHIP: An Adaptive IP Hopping Method for Moving Target Defense to Thwart Network AttacksabstractIn a static network, attackers can easily launch network attacks on target hosts which have long-term constant IP addresses. In order to defend against attackers effectively, many defense approaches use IP hopping to dynamically transform IP configuration. However, these approaches usually focus on one type of network attacks, scanning attacks or Denial of Service (DoS) attacks, and cannot sense network situations. This paper proposes AHIP, an adaptive IP hopping method for moving target defense (MTD) to defend against different network attacks. We use a trained lightweight one-dimensional convolutional neural network (1D-CNN) detector to judge whether there are no attacks, scanning attacks or DoS attacks in the network, which can adaptively trigger corresponding IP hopping strategy. We use specific hardware and software to create the software defined network (SDN) environment for experiments. The experiments prove that AHIP performs better to thwart network attacks and has lower system overhead. Fengyuan Shi 0005, Zhou Zhou 0007, Qingyun Liu 0001, Xiuguo Bao |
CSCWD | 5 |
| 2023 | DarkFT: Automatic Scanning Behavior Analysis with FastText in Darknet TrafficabstractNetwork telescopes (Darknets) collect and record unsolicited Internet-wide traffic destined for a routed but unused address space, which provides a global perspective on Inter-net scanning behavior. However, it’s very difficult to extract meaningful information from Darknet traffic including a large number of unlabeled packets. In recent years, some work has used NLP techniques for self-supervised learning to generate embeddings as a rich representation of the Darknet. However, we found that previous resulting embeddings are not general enough. This paper proposed a new traffic representation model called Darknet Traffic Representations using FastText (DarkFT) which trains contextualized representation from large-scale unlabeled data. The embedding features can be applied to semi-supervised and unsupervised tasks. We conduct experiments on a dataset collected by network telescopes located in Japan and get better accuracy compared with state-of-the-art NLP-based approaches, especially for unknown senders. Yixiang Zhao, Zhou Zhou 0007, Qingyun Liu 0001 |
CSCWD | 6 |
| 2023 | IDTracker: Discovering Illicit Website Communities via Third-party Service IDsabstractIllicit websites are restricted by governments and application marketplaces due to their detrimental impact on society. Third-party web services play a crucial role in enabling illicit webmasters to establish websites rapidly and evade detection. In this paper, we discover that third-party services usually assign unique credentials to website developers as their identifications (IDs). Websites using the same services with identical IDs are likely to be hosted on shared infrastructures and have textually similar domain names. This observation sparks the idea of building a community of illicit websites by leveraging third-party service IDs. Therefore, we design IDTracker, a novel system for detecting illicit website communities based on domain name semantic and infrastructure relationship features, which empower classification algorithms to achieve a high F1 score of 0.8968. Furthermore, we deploy IDTracker on an Internet Service Provider's (ISP) environment for three months and identify 6,830 illicit communities containing 165,378 illicit websites. Many of these illicit websites can not be identified by the most sophisticated engines, such as Symantec and Baidu, because of the cloaking tactics. In addition, we conduct a large-scale and long-term measurement on the network infrastructures and third-party services of illicit communities, revealing new phenomena. Our findings can help security communities to thwart illicit websites more effectively. Chenxu Wang 0006, Zhao Li 0010, Jiangyi Yin, Zhenni Liu, Qingyun Liu 0001 |
DSN | 6 |
| 2023 | Unveiling Flawed Cache Structures in DNS Infrastructure via Record WatermarkingabstractThe Domain Name System (DNS) is an essential component of the internet, providing name resolution services to navigate clients to various resources on the network. Caches are critical to the efficient operation of DNS, both in terms of service quality and security. Major service providers maintain complex DNS infrastructure with multiple cache layers to handle client queries. Unfortunately, access to these caches is not available, making it difficult to understand the cache structure. In this study, we propose methodologies for identifying hidden cache structures in DNS infrastructure using watermark records. We further applied our methods to conduct global measurements, utilizing open resolvers as vantage points. Our measurement results indicated that flawed DNS cache structures exist, leaving the DNS infrastructure vulnerable and inefficient. Thousands of client networks suffer from severe cache fragmentation. Furthermore, a large number of recursive resolvers rely on fragile or poor cache structures. Dikai Mo, Yujia Zhu, Zhen Jie, Qingyun Liu 0001, Binxing Fang |
GLOBECOM | 5 |
| 2023 | Shrink: Identification of Encrypted Video Traffic Based on QUICabstractWith the increasing prevalence of network videos, video traffic has become a significant portion of overall network traffic. Due to the presence of harmful content such as pornography and violence in network videos, network monitoring is necessary. However, the encryption of videos poses challenges for network monitoring. More and more video service providers are adopting QUIC as the default video transmission protocol to accelerate data transfer speeds. However, the existing methods for identifying encrypted video traffic do not apply to QUIC. Video service providers typically employ Content Delivery Network (CDN) technology to enhance user experience, which can result in missing video chunks for side-channel identification. Additionally, fluctuations in network conditions can lead to the retransmission of video chunks. This paper proposes Shrink, a QUIC-based encrypted video traffic identification method. It effectively extracts video chunks from online QUIC encrypted video traffic and proposes a bucket structure and global-local match to alleviate the issues of video chunks retransmission and loss. Furthermore, a bucket word dictionary is designed to enhance the method’s running speed. Experimental results demonstrate that Shrink performs well in real network environments, exhibiting superior accuracy and speed compared to existing state-of-the-art methods. Weitao Tang, Meijie Du, Zhao Li 0010, Zhou Zhou 0007, Qingyun Liu 0001 |
IPCCC | 6 |
| 2023 | CCSv6: A Detection Model for DNS-over-HTTPS Tunnel Using Attention Mechanism over IPv6abstractIn this paper, we first show DNS-over-HTTPS (DoH) tunneling detection methods verified to be effective over IPv4 can be applied to IPv6, and then propose a new model called CCSv6, using attention-based convolution neural network to build classifiers with flow-based features to detect DoH tunneling over IPv6, achieve 99.99% accuracy on the IPv6 dataset. In addition, we discuss the influence of various factors such as locations or DoH resolvers on the detection results in detail over IPv6. All the more important, our model shows better transfer learning ability, which can achieve the F1-score of 96% when trained on the IPv6 dataset and tested on the IPv4 dataset. Liang Jiao, Yujia Zhu, Fenglin Qin, Qingyun Liu 0001 |
ISCC | 6 |
| 2023 | GoGDDoS: A Multi-Classifier for DDoS Attacks Using Graph Neural NetworksabstractDistributed Denial of Service (DDoS) attacks are rising, evolving and growing sophistication. Multi-vector which leverages more than one methods is prevalent recently. To cope with multi-vector DDoS attack, it is necessary to classify DDoS attacks for taking robust measures. However, existing ML-based approaches for DDoS traffic multi-classification barely leverage relationships between packets and flows, which are crucial information that can significantly improve multi-classification performance. This paper proposes GoGDDoS, a multi-classifier for DDoS attacks. Concretely, we construct GoG traffic graph to clearly compress relationships between packets and flows. It merges relationship graphs of packets and flows by using graph of graph. Then, we build a two-level Graph Neural Network model to mine potential attack patterns from GoG traffic graph. The experiments with well-known datasets show that GoGDDoS performs better than its counterparts. Zhou Zhou 0007, Fengyuan Shi 0005, Qingyun Liu 0001 |
ISCC | 6 |
| 2023 | Hunting for Hidden RDP-MITM: Analyzing and Detecting RDP MITM Tools Based on Network FeaturesabstractRemote Desktop Protocol (RDP) is commonly used for remote access to windows computers. As more and more people work remotely, the number of users of RDP is increasing, making RDP a growing concern in cybersecurity. The latest way to threaten RDP security is RDP man-in-the-middle (MITM) tools which realize the MITM function in an RDP connection and automate the MITM attack process, significantly reducing the difficulty of network attacks. At the same time, RDP MITM tools can be used for high-interaction RDP honeypots. In order to mitigate this risk, we present the first in-depth study of RDP MITM tools in this paper. By analysis and experiment, we identify network features that can be used to detect RDP MITM tools effectively. Based on packet latency and TLS handshake, we propose a machine learning classifier that can detect RDP MITM tools for securing RDP connections. Finally, we analyze the deployment of RDP MITM tools in the wild and effectively measure the RDP MITM tools using our proposed detection approach. Zhou Zhou 0007, Fengyuan Shi 0005, Qingyun Liu 0001 |
ISCC | 7 |
| 2023 | Before Toasters Rise Up: A View into the Emerging DoH Resolver's Deployment RiskabstractAs an encryption protocol for DNS queries, DNS-over-HTTPS (DoH) is becoming increasingly popular, and it mainly addresses the last-mile privacy protection problem. However, the security of DoH is in urgent need of measurement and analysis due to its reliance on certificates and upstream servers. In this paper, we focus on the DoH ecosystem and conduct a one-month measurement to analyze the current deployment of DoH resolvers. Our findings indicate that some of these resolvers use invalid certificates, which can compromise the security and privacy advantages of the protocol. Furthermore, we found that many providers are at risk of certificate outages, which could cause significant disruptions to the DoH ecosystem. Additionally, we observed that the centralization of DoH resolvers and upstream DNS servers is a potential issue that needs addressing to ensure the stability of the ecosystem. Yuqi Qiu, Baiyang Li, Zhiqian Li, Liang Jiao, Yujia Zhu, Qingyun Liu 0001 |
ISCC | 6 |
| 2023 | A Comprehensive Evaluation of the Impact on Tor Network Anonymity Caused by ShadowRelayabstractAs a distributed anonymous network run by volunteers, Tor relays are often manipulated by operators to achieve their goals. Our work reveals that some relays, named ShadowRelay, are bound to hidden nodes and actively forward user traffic to the next-hop relay or target without the user's knowledge. To detect ShadowRelays, we developed HiddenSniffer based on client and Tor relay collusion, and found 162 hidden nodes distributed across 22 countries, along with 85 Shadow Relays which account for 2.08% of the total relay bandwidth. Additionally, there exists a family relationship among the Shadow Relays, with the largest family containing 24 members. The experimental results indicate that ShadowRelays have increased the number of ASes capable of sniffing user traffic by 27.6%, and improved the ability of 14.7% of attackers to launch traffic confirmation attacks. Furthermore, ShadowRelays adversely impact the Tor network's availability by introducing increased transmission delay within the circuits. Muqian Chen, Qingyun Liu 0001, Jinqiao Shi |
ISCC | 5 |
| 2023 | SIFAST: An Efficient Unix Shell Embedding Framework for Malicious Detection
Songyue Chen, Rong Yang 0008, Hongwei Wu, Yanqin Zheng, Qingyun Liu 0001 |
ISC | 7 |
| 2023 | Measuring DNS-over-Encryption Performance Over IPv6abstractIn recent years, encrypted DNS such as DNS-over-HTTPS (DoH) and DNS-over-TLS (DoT) has gained significant traction as privacy-preserving alternative to conventional DNS. While several studies have measured the performance of encrypted DNS relative to conventional DNS, they are only performed over IPv4, little has been done to understand their status over IPv6. Besides, previous studies can not obtain the absolute query latency due to lack of control over vantage points.This paper performs by far the fist end-to-end performance measurements on encrypted DNS over IPv6. By analyzing measurement results, we have gained several insights. In general, the quality of service for encrypted DNS is satisfying. Over IPv6, encrypted DNS performance varies across resolvers, and is affected by the location issuing DNS queries, the type of encrypted DNS protocol used and the latency to resolvers. Compared with IPv4, the performance of encrypted DNS of different resolvers over IPv6 is improved to some extent. In addition, we also find other problems such as the quality of service of resolver Ahadns is significantly low both over IPv6 and IPv4, as well as the performance of encrypted DNS for resolver Alidns significantly deteriorates when switching from IPv4 to IPv6. Based on our observations, we provide recommendations and discuss situations in which switching to IPv6 may be beneficial. We hope that our tools developed for performing measurements can help people in different regions to choose to the right recursive resolver and network environment, and that our findings can contribute to improve IPv6 Internet infrastructure and inform continuing encrypted DNS deployment over IPv6. Liang Jiao, Yujia Zhu, Baiyang Li, Qingyun Liu 0001 |
TrustCom | 4 |
| 2022 | HinPage: Illegal and Harmful Webpage Identification Using Transductive Classification
Lingjing Yu, Qingyun Liu 0001 |
Inscrypt | 3 |
| 2022 | GraphDDoS: Effective DDoS Attack Detection Using Graph Neural NetworksabstractDistributed Denial of Service (DDoS) attacks have occurred frequently in recent years, causing massive damage. It is critical to detect DDoS attacks fast and accurately. Previous Deep Learning (DL) methods for detecting DDoS attacks barely leverage the relationships between packets and between flows in traffic, which are crucial information that can significantly improve detection performance. This paper proposes GraphDDoS, a GNN-based approach for detecting DDoS attacks using endpoint traffic graphs. Concretely, we convert traffic into endpoint traffic graphs, containing information of packets’ relationships (structure of a single flow) and flows’ relationships (burst information and periodic information of multiple flows). Then, converted endpoint traffic graphs are sent to the GNN classifier to learn DDoS attack patterns accurately. The experiments with well-known datasets show that GraphDDoS outperforms the state-of-the-art DL-based approaches. The effectiveness is mainly introduced by the capability of GraphDDoS to learn patterns of attacks structured as graphs. Zhou Zhou 0007, Meijie Du, Qingyun Liu 0001 |
CSCWD | 7 |
| 2022 | Accelerate State Sharing of Network Function with RDMAabstractState sharing can help network functions (NFs) provide support for clustered deployment and minimize the jitter caused by elastic scaling, one of the central futures of NFV. But the network overhead of remote access and the dynamically changing workload restrict the performance of the state sharing framework. The state-of-the-art state sharing frameworks mainly use the affinity strategy to migrate the target state from another node to itself. This strategy is unsuitable for asymmetric routing scenario because multiple nodes will access the same state. This paper presents RedKV, a fast and flexible Remote Direct Memory Access (RDMA) Enhanced Distributed Key-Value store that supports low-latency state sharing both in symmetric and asymmetric routing scenarios. RedKV has two unconventional designs: First, it scatters states over cluster nodes without affinity strategy and uses a unified addressing (UA) space to hide the complexity of state access. Second, it accelerates the operation with RDMA and optimizes the process according to the characteristics of RDMA and NF to mitigate the cost of remote access. Our evaluation shows that RedKV can help network functions share states with an added latency overhead of 2.26µs, outperforming state-of-art solutions by 4.6 times. The pause time caused by the scaling event has also been reduced by 81.5%. Chenming Chang, Chao Zheng 0001, Liang Zhan, Qingyun Liu 0001 |
GLOBECOM | 6 |
| 2022 | P4-NSAF: defending IPv6 networks against ICMPv6 DoS and DDoS attacks with P4abstractInternet Protocol Version 6 (IPv6) is expected for widespread deployment worldwide. Such rapid development of IPv6 may lead to safety problems. The main threats in IPv6 networks are denial of service (DoS) attacks and distributed DoS (DDoS) attacks. In addition to the similar threats in Internet Protocol Version 4 (IPv4), IPv6 has introduced new potential vulnerabilities, which are DoS and DDoS attacks based on Internet Control Message Protocol version 6 (ICMPv6). We divide such new attacks into two categories: pure flooding attacks and source address spoofing attacks. We propose P4-NSAF, a scheme to defend against the above two IPv6 DoS and DDoS attacks in the programmable data plane. P4-NSAF uses Count-Min Sketch to defend against flooding attacks and records information about IPv6 agents into match tables to prevent source address spoofing attacks. We implement a prototype of P4-NSAF with P4 and evaluate it in the programmable data plane. The result suggests that P4-NSAF can effectively protect IPv6 networks from DoS and DDoS attacks based on ICMPv6. Zhou Zhou 0007, Qingyun Liu 0001, Zhao Li 0010 |
ICC | 4 |
| 2022 | Fighting Against Piracy: An Approach to Detect Pirated Video Websites Enhanced by Third-party ServicesabstractAlong with the development of video streaming, the increasing number of pirated video websites has caused unprecedented damage to copyright holders and potential security risks to their users. Though many efforts have been made to take down pirated video websites, they are still emerging by utilizing evading approaches like Fast-Flux domains and Cybercrime-as-a-Service(CaaS) tools. In this paper, to detect pirated video websites, we propose a Third-party Enhanced Pirated Video Website Classification Network (TEP-Net), which integrates both semantic features and relationship information between websites and their third-party services. More specifically, we apply CNN-BiLSTM-Attention to explore both character-level and domain-level textual embedding and utilize relationship information by constructing statistical features in classification. The experiment shows that TEP-Net achieves a significant performance compared with existing methods. Furthermore, we perform an in-depth analysis of the CaaS behind pirated video websites. Our research can help the security community fight against video piracy more precisely and effectively. Zhao Li 0010, Jiangyi Yin, Meijie Du, Qingyun Liu 0001 |
ISCC | 6 |
| 2022 | Detection of DoH Tunnels with Dual-Tier ClassifierabstractDNS over HTTPS (DoH) has been deployed to provide confidentiality in the DNS resolution process. However, encryption is a double-edged sword in providing security while increasing the risk of data tunneling attacks. Current approaches for plaintext DNS tunnel detection are disabled. Due to the diversity of tunneling tool variations and the low proportion of tunneled traffic in real situations, detecting malicious behaviors is becoming more and more challenging. In this paper, we propose a novel behavior-based model with Dual-Tier Tunnel Classifier (DTC) for tool-level DoH tunneling detection. The major advantage of DTC is that it can not only capture existing tunneling tools but also explore unknown ones in the wild. In particular, DTC considers data imbalance, which improves robustness of the model in the open environment. Our method has been proven successful in both closed and open scenarios, achieving 99.99 % accuracy in detecting known malicious DoH traffic, 96.93% accuracy in unknown and 95.31 % accuracy in identifying malicious DoH tunnel tools. Yuqi Qiu, Baiyang Li, Liang Jiao, Yujia Zhu, Qingyun Liu 0001 |
MSN | 5 |
| 2022 | A Lightweight Graph-based Method to Detect Pornographic and Gambling Websites with Imperfect DatasetsabstractWith the widespread abuse of information technology, pornographic and gambling websites develop rapidly. They affect the physical and mental health of children and endanger personal property. Therefore, it is necessary to detect them. However, the existing detection methods ignored that imperfect datasets are common in the scenario of pornographic and gambling websites which are hence adverse to the detection. Those imperfections specifically include sparse samples, mismatch and imbalanced datasets. In addition, over-reliance on visual features incurred high overhead.To overcome these shortcomings, we innovatively propose a lightweight graph-based method to detect pornographic and gambling websites through semi-supervised learning of textual content. The semi-supervised learning is to solve sparse samples and mismatch datasets, while the graph-based approach can combine the semi-supervised part with community discovery to deal with imbalanced datasets. Specifically, we perform the detection process with the utilization of modified TF-IDF and Louvain during the iteration and updating by the EM algorithm. The experimental results show that our method achieves the best 92.01% Macro-Avg-F1 with the shortest CPU time and outperforms all baselines. We also illustrate that the designed components in our model do contribute to the detection. Xiaoqing Ma, Chao Zheng 0001, Zhao Li 0010, Jiangyi Yin, Qingyun Liu 0001, Xunxun Chen |
TrustCom | 5 |
| 2021 | Attributed Heterogeneous Graph Neural Network for Malicious Domain DetectionabstractMalicious activities on the Internet are one of the most dangerous threats to users and organizations. Because of the flexibility a nd accessibility of domains, cyber criminals often utilize them to launch cyber attacks such as phishing or malware. Most of traditional malicious domain detection methods rely on feature engineering to learn the patterns of malicious domains. However, these methods can be easily evaded by some sophisticated evasion techniques such as Domain-flux or Fast-flux. Some recent studies utilized graph-based models to infer malicious domains and achieved better performance, yet without a fine-grained modeling of DNS scenarios. In this paper, we propose an attributed heterogeneous graph neural network model, GAMD, to detect malicious domains in a semi-supervised learning paradigm. Concretely, we utilize attributed heterogeneous information network to model the DNS scenarios with different types of nodes including domain, host, resolved-IP and different types of relation, such as request and resolution relations. We then design a fine-grained node type-aware feature transformation and edge type-aware aggregation mechanism to fuse the node attributes and structure information simultaneously and complete the inference over DNS graphs. In the experiments, we evaluate the performance of our model on a large-scale realworld passive DNS data and show that the proposed method outperforms the state-of-the-art in most evaluation metrics. Shuai Zhang 0007, Zhou Zhou 0007, Da Li 0002, Youbing Zhong, Qingyun Liu 0001 |
CSCWD | 5 |
| 2021 | A P4-Based Packet Scheduling Approach for Clustered Deep Packet Inspection AppliancesabstractPacket scheduling approach enables clustered deep packet inspection (DPI) appliances to achieve stateful packet processing through efficient cooperation. In this paper, we propose a novel methodology of packet scheduling, called P4CLUS, for clustering DPI appliances efficiently. P4CLUS eliminates unnecessary packet transmission between peer nodes via the idea of packet path prediction. It performs sketch-based measurement on traffic handling nodes to record their flow distributions and then compresses the results to a compact packet forwarding table. Our approach guarantees per-flow consistency with lower bandwidth overhead and is effective in the practice of eliminating asymmetric traffic in stateful DPI. We construct the P4CLUS prototype with P4 on Protocol-Independent Switch Architecture (PISA). Our evaluation shows that P4CLUS can reduce the bandwidth overhead by up to 73.75% than the hash-based method, and only requires 2MB SRAM and 80KB TCAM resources at most. Qingyun Liu 0001, Chao Zheng 0001 |
ICCCN | 3 |
| 2021 | CDNFinder: Detecting CDN-hosted Nodes by Graph-Based Semi-Supervised ClassificationabstractAs a crucial internet infrastructure, Content Delivery Network (CDN) is widely deployed. Detecting CDN-hosted nodes from network traffic is important for Quality of Service (QoS), malware detection and firewall rule-sets. Current researches use hand-crafted rules, classification or clustering methods. However, those methods relying on plaintext are limited by the invisibility of plaintext due to encryption, as well as the limitations of DNS Resource Records, such as unreliability. Besides, those methods don't dig the structural information of domains and IPs. To overcome those shortcomings, we present CDNFinder, a novel method to detect CDN-hosted nodes by graph-based semi-supervised classification. Based on the active datasets collected in 10 vantage points, we construct the graph and extract innovative attributes. By modifying Graph Neural Network (GNN), CDNFinder outperforms classical machine learning methods, especially in recall rate (around 98%). Meanwhile, CDNFinder shortens the runtime of classical GNN algorithm by about 31% with no loss in metrics. Xiaoqing Ma, Chao Zheng 0001, Zhao Li 0010, Qingyun Liu 0001, Xunxun Chen |
ISCC | 4 |
| 2021 | Neighbours and Kinsmen: Hateful Users Detection with Graph Neural Network
Nayyar Abbas Zaidi, Qingyun Liu 0001, Gang Li 0009 |
PAKDD (1) | 3 |
| 2021 | Peek Inside the Encrypted World: Autoencoder-Based Detection of DoH ResolversabstractDNS-over-HTTPS (DoH), as a rising star to improve DNS security and privacy, has developed rapidly in recent years. It mixes with HTTP features, shares ports with other web services and provides API with URI templates. The unique characteristics of DoH, as well as its fast growth, bring both promising prospects and new risks, e.g. botnet communication, name abuse and data exfiltration. It is essential for network operators to learn about adoption and usage of DoH resolvers. Active scanning may be a possible way. However, it is considered to incur significantly additional overhead, which can be inefficient and aggressive. In this paper, we present DOHUNTER, a system for automati-cally discovering DoH resolvers. DOHUNTER: (i) picks DoH flow from miscellaneous HTTPS traffic, (ii)confirms DoH resolvers based on the detected DoH flow, (iii)mines other related DoH resolvers from the known ones. Our real-world experiments demonstrate the effectiveness of DOHUNTER in detecting DoH flow and finding DoH resolvers. Utilizing DOHUNTER, we witness an alarming increase in DoH adoption. Additionally, we also reveal oblivious growing trends of DoH, which may provide advice for both users and network operators. Jiating Wu, Yujia Zhu, Baiyang Li, Qingyun Liu 0001, Binxing Fang |
TrustCom | 4 |
| 2020 | BPA: The Optimal Placement of Interdependent VNFs in Many-Core System
Youbing Zhong, Zhou Zhou 0007, Xuan Liu 0006, Da Li 0002, Meijun Guo, Shuai Zhang 0007, Qingyun Liu 0001, Li Guo 0001 |
CollaborateCom (2) | 7 |
| 2020 | Predicting User Influence in the Propagation of Toxic Information
Yishuo Zhang, Penghui Jiang, Zhao Li 0010, Qingyun Liu 0001 |
KSEM (1) | 6 |
| 2020 | MAAN: A Multiple Attribute Association Network for Mobile Encrypted Traffic Classification
Fengzhao Shi, Chao Zheng 0001, Qingyun Liu 0001 |
SecureComm (1) | 4 |
| 2020 | You Are What You Broadcast: Identification of Mobile and IoT Devices from (Public) WiFi
Lingjing Yu, Bo Luo, Zhaoyu Zhou, Qingyun Liu 0001 |
USENIX Security Symposium | 5 |
| 2019 | NTS: A Scalable Virtual Testbed Architecture with Dynamic Scheduling and Backpressure
Youbing Zhong, Zhou Zhou 0007, Da Li 0002, Wenliang He, Chao Zheng 0001, Qingyun Liu 0001, Li Guo 0001 |
CollaborateCom | 6 |
| 2019 | Hunting for Invisible SmartCam: Characterizing and Detecting Smart Camera Based on Netflow AnalysisabstractNowadays, the rapid growth of cloud computing and IoT enabled services among multiple organizations brings both promising prospects and security & privacy challenges. IP cameras have become a top target for hackers because of their relatively high computing power and throughput. To understand the risks of these threats requires learning about IP cameras-where are they, how many are there? Active scanning is considered to be an effective way, like SHODAN. However, deployment of smart cameras in the network address translation (NAT) environments with dynamic locations is usually desired. To find these Invisible Cameras, CamHunter: (i) introduces three statements of smart cameras when they are online, (ii) concludes the most popular smart cameras in China have very similar communication patterns, (iii) proposes a model to detect smart cameras in a passive way constructed by nineteen feature sets, and (iv) raises alarms for IoT manufacturers. Our real-world experiments demonstrate the effectiveness of CamHunter in finding smart cameras even if they are behind NATs and using encrypted connections like SSL/TLS or private protocols. We argue that CamHunter represents an important view of IoT security and privacy, and it can guide the effort of designing and protecting smart cameras. Baiyang Li, Yujia Zhu, Qingyun Liu 0001, Zhou Zhou 0007, Li Guo 0001 |
ICC | 3 |
| 2019 | Tear Off Your Disguise: Phishing Website Detection Using Visual and Network Identities
Zhaoyu Zhou, Lingjing Yu, Qingyun Liu 0001, Yang Aron Liu, Bo Luo |
ICICS | 3 |
| 2019 | IDNS: A High-Performance Model for Identification of DNS Infrastructures on Large-scale TrafficabstractDomain Name System (DNS) is indispensable in a large number of network applications. Identifying DNS infrastructures into different roles hierarchically is highly desired for a variety of purposes such as network management and threat evaluation. However, traditional measurements almost all depend on active scanning without considering dynamic packet-level features of different DNS infrastructures.In this paper, we propose a high-performance model IDNS (Identifying DNS) based on passive measurement. IDNS: (i) extracts single-packet field features (SFF) and multi-packet statistical features (MSF) from DNS traffic, (ii) utilizes an estimation algorithm to calculate MSF for satisfying online processing speed, and (iii) applies several classifiers in Ensemble Learning and Incremental Learning. We perform an extensive evaluation based on a large volume of DNS queries and responses collected from one ISP. The evaluation results demonstrate that the best classifier in Ensemble Learning can reach 90% accuracy rate while the classifier in Incremental Learning can reach 80% with the highest scalability. Caiyun Huang, Yujia Zhu, Qingyun Liu 0001, Binxing Fang |
ISCC | 4 |
| 2018 | SASD: A Self-Adaptive Stateful Decompression ArchitectureabstractDue to the increasing threats in the current network environment, many researchers have shifted their interests to network content audit, which combines deep packet inspection and natural language processing. However, the performance of network content audit systems is becoming the bottle-neck because of the demand on processing fast growing compressed traffic. While compressed traffic is often split into multiple out-of-order packets for transmission, stateful decompression ensures that the compressed data are processed in a timely manner without waiting for all the compressed traffic to arrive before decompressing. In the meanwhile, hardware innovations lead to new type of devices being invented, which shows promise to fully handle the offloaded traffic for complex calculations at higher throughput than software-based solutions. We consider both software-based and hardware-based solutions for decompressing traffic from network content audit systems and study the workload. We notice that the performance is data-dependant: hardware-based decompression solutions perform better for longer compressed data than software method. On the contrary, software-based decompressing methods are more preferred for the short content in terms of the processing speed. So there is no one-size-fits-all solution. In this paper, we combine the advantages of hardware and software and propose a novel self-adaptive stateful decompression architecture to support fast decompression in accordance with the traffic status and system state. Experiments on real-world traffic show that our proposed architecture can achieve about three times of the data decompression efficiency, compared to the best pure software and hardware algorithm, which can significantly improve the detection efficiency of many network content audit systems. Zhou Zhou 0007, Qingyun Liu 0001, Yujia Zhu, Da Li 0002, Li Guo 0001 |
GLOBECOM | 3 |
| 2018 | Hashing Incomplete and Unordered Network Streams
Chao Zheng 0001, Qingyun Liu 0001, Binxing Fang |
IFIP Int. Conf. Digital Forensics | 3 |
| 2018 | WDMTI: Wireless Device Manufacturer and Type Identification Using Hierarchical Dirichlet ProcessabstractWireless devices have been widely adopted across all domains. With the convenience brought by wireless communication technology, increasing number of conventional (wired) devices are evolving to become wireless. However, significant security issues arise with the popularity of wireless devices. To start an attack, the adversary usually performs a network reconnaissance to discover exposed devices, identify device manufacturers and types, and then scan for vulnerabilities. From the defense side, network administrators are expected to identify the potential vulnerabilities/risks and enforce Network Access Control (or Network Admission Control, NAC) on all the connecting devices. To do this, it is essential to accurately identify the make/model/type of each device that attempts to connect to the network, e.g., MacBooks, Samsung smart phones (Android), Amazon kindles, DLink surveillance cameras, TP-Link smart plugs, etc. In this paper, we present a novel approach, namely WDMTI, for the identification of wireless device manufacturer and type. We tackle the challenge from two aspects: the features and the classification model. First, we claim that it is critical to discover the device manufacturer and type as soon as the device requests to join the WLAN, and it is unrealistic to make other assumptions on the status of the device, e.g., assuming that the device is booting up or initializing a new connection to corresponding servers/clouds. We primarily depend on the features extracted from the network connection phase, while features from device booting are considered "bonus". In particular, we propose to utilize features from the raw HDCP packets, which is shown to be sufficient for device manufacturer and type recognition with high accuracy. Meanwhile, in the WDMTI system, we employ the Hierarchical Dirichlet Process (HDP), which is a nonparametric Bayesian model for grouped data. HDP allows new groups to be introduced with new data being added, i.e. previously unknown devices connect to the network and the extracted features receive new labels. The WDMTI mechanism is dynamically retrained on-line, instead of requiring a time-consuming off-line retraining process. Our experiments show that WDMTI identifies known types of devices with average accuracy of 0.89, and new types of devices with average accuracy of 0.96, both of which is higher than the state-of-art approaches. In summary, we present a wireless device manufacturer and type identification (WDMTI) system that is both scalable and accurate, and capable of adapting to unknown types of devices on-the-fly. Lingjing Yu, Zhaoyu Zhou, Yujia Zhu, Qingyun Liu 0001, Jianlong Tan |
MASS | 5 |
| 2018 | My Friend Leaks My Privacy: Modeling and Analyzing Privacy in Social NetworksabstractWith the dramatically increasing participation in online social networks (OSNs), huge amount of private information becomes available on such sites. It is critical to preserve users' privacy without preventing them from socialization and sharing. Unfortunately, existing solutions fall short meeting such requirements. We argue that the key component of OSN privacy protection is protecting (sensitive) content -- privacy as having the ability to control information dissemination. We follow the concepts of private information boundaries and restricted access and limited control to introduce a social circle model. We articulate the formal constructs of this model and the desired properties for privacy protection in the model. We show that the social circle model is efficient yet practical, which provides certain level of privacy protection capabilities to users, while still facilitates socialization. We then utilize this model to analyze the most popular social network platforms on the Internet (Facebook, Google+, WeChat, etc), and demonstrate the potential privacy vulnerabilities in some social networks. Finally, we discuss the implications of the analysis, and possible future directions. Lingjing Yu, Sri Mounica Motipalli, Dongwon Lee 0001, Peng Liu 0005, Qingyun Liu 0001, Jianlong Tan, Bo Luo |
SACMAT | 6 |
| 2017 | ProxyDetector: A Guided Approach to Finding Web ProxiesabstractWith the Internet becoming the dominant channel for business and life, web proxies are also increasingly used for illegal purposes such as propagating malware, impersonate phishing pages to steal sensitive data or redirect victims to other malicious targets. In this paper, using thousands of web proxy URLs crawled, we performed a large-scale study on the DOM (Document Object Model) structure features. Our study reveals the existence of the dedicated web proxy DOM among hosts that play orchestrating roles in proxy activities. Motivated by their distinctive features in DOM and URL, we developed an automatical stepping-stone detection system-ProxyDetector. Specially, we explored the potential benefits of considering DOM-based features, which improved 25% recall rate than before. We extensively evaluated ProxyDetector with four methods on a diverse spectrum of corpora with 2,068 web proxy sites and 26,066 legitimate sites. Capable of achieving over 95% precision of web proxy sites with a high recall rate of 96.5% on average, our ProxyDetector has been demonstrated to be an effective solution of detecting the web proxy sites. Peng Zhang 0001, Qingyun Liu 0001 |
LCN | 3 |
| 2017 | WiFi fingerprint releasing for indoor localization based on differential privacyabstractWiFi fingerprint-based localization is regarded as one of the most promising techniques for indoor localization. However, this raises serious privacy concerns. Current approaches to mitigate the privacy concerns rely on the encryption with large calculation consumption. In this paper, we propose a data obfuscation mechanism based on the generalized version of differential privacy. We extend the standard definition to the indoor WiFi fingerprint data for spatial counting where the inputs belong to multiple dimensions of numerical data in a limited range. With a given privacy budget, the proposed method generalizes the original dataset, and then specializes it using differential privacy. As the designed novel scheme expand the range for specialization, the data set released by the proposed algorithm can yield better mining results. Furthermore, experimental results give out comparisons between nonuniform and uniform ε selection scheme, and find uniform ε selection scheme can fully use the privacy budget in our situation. Yujia Zhu, Qingyun Liu 0001, Yang Aron Liu, Peng Zhang 0001 |
PIMRC | 3 |
| 2016 | CookieMiner: Towards real-time reconstruction of web-downloading chains from network tracesabstractNetwork traces are one of the most exhaustive data sources for the forensic investigation of computer security incidents. Recent advances in capturing the network traces techniques have facilitated the forensic processing, including the reconstruction. Unfortunately, off-line Web-downloading chain reconstruction could not meet the demands for real time processing. Furthermore, the packets in prior studies are mainly captured on the client or server side, which is difficult to monitor the network traffic in the real-time. Consequently, how to online reconstruct from the packets captured by the gateway is getting challenging. In this paper, based on the packets from the gateway, we propose a novel system, CookieMiner. CookieMiner first identifies the web-downloading resources, and then use the cookies to reconstruct the web-downloading chains reversely. All cookies would be firstly split into a series of tokens by semicolon and a threshold value will be set by the SET_K algorithm according to the number of tokens. Next, the HTTP packets whose tokens' number is greater than the threshold will be sorted by timestamp and then inserted into the corresponding web-downloading chain. Finally, the most frequent URLs are extracted as entry points from all the chains with the same web-downloading resources. In addition, by a user study involving 6 pair-wise downloading applications, CookieMiner can reconstruct the web-downloading chains and find their entry points with high precision and low false positive rate. Peng Zhang 0001, Chao Zheng 0001, Qingyun Liu 0001 |
ICC | 4 |
| 2015 | Limited Dictionary Builder: An approach to select representative tokens for malicious URLs detectionabstractCybercriminals use Malicious Uniform Resource Locators (URLs) as the entry to implement a variety of web attacks, such as phishing, spamming, and malware distribution, which may lead to huge finance and data loss. Thus, malicious URLs should be detected as accurately and quickly as possible. Heuristic-based detection approaches are one of the most popular methods to achieve the above goals. The detection results come from the usage of many heuristic features in this approach. However, tremendous new pages and meaningless tokens lead to the explosion of feature sets, and exhaust memory space finally. In this paper, we try to address the problem by selecting some representative members from the initial feature set, which should have the best predictive ability among the same number of selected features. For each feature, we give an evaluation method of O(1) complexity to measure its predictive ability. Then we make the selection based on all the measured values with linear complexity. Experimental results show that our approach can achieve almost the same false negative rate using only 8.3% features for malicious URLs detection, comparing with prior approaches. Moreover, our approach may work efficiently in the big data era, as it can handle 20 thousand URLs per second in our experiments on average. Hongzhou Sha, Zhou Zhou 0007, Qingyun Liu 0001, Tingwen Liu, Chao Zheng 0001 |
ICC | 3 |
| 2015 | Data-oriented multi-index hashingabstractMulti-index hashing (MIH) is the state-of-the-art method for indexing binary codes, as it divides long codes into substrings and builds multiple hash tables. However, MIH is based on the dataset codes uniform distribution assumption, and will lose efficiency in dealing with non-uniformly distributed codes. Besides, there are lots of results sharing the same Hamming distance to a query, which makes the distance measure ambiguous. In this paper, we propose a data-oriented multi-index hashing method. We first compute the covariance matrix of bits and learn adaptive projection vector for each binary substring. Instead of using substrings as direct indices into hash tables, we project them with corresponding projection vectors to generate new indices. With adaptive projection, the indices in each hash table are near uniformly distributed. Then with covariance matrix, we propose a ranking method for the binary codes. By assigning different bit-level weights to different bits, the returned binary codes are ranked at a finer-grained binary code level. Experiments conducted on reference large scale datasets show that compared to MIH the time performance of our method can be improved by 36.9%-87.4%, and the search accuracy can be improved by 22.2%. Qingyun Liu 0001, Hongtao Xie 0001, Li Guo 0001 |
ICME | 1 |
| 2014 | GuidedTracker: Track the victims with access logs to finding malicious web pagesabstractMalicious web pages have become a malignant tumour for the Internet, which spread malicious code, steal people's private information, and deliver spamming advertisements. And how to distinguish them from the huge number of normal web pages effectively remains a huge challenge in the era of big data. To detect malicious pages, one needs to first collect candidate web pages that are live on the web; then filter massive legitimate pages using fast filters and finally examine the remaining pages using precisely but slow analyzer. However, there are new challenges recently for these conventional techniques, including large scale, imbalance data and the usage of cloaking techniques. To cope with these challenges, the malicious URL detection system should perform more efficiently. In this paper, we propose a system, named GuidedTracker, to search for suspicious malicious pages. GuidedTracker starts from the seed set which includes known malicious pages. Then, it automatically figures out those victims based on the seed set and the visit relation database. Finally, the access records of these victims are used to identify other malicious pages. In this way, GuidedTracker increase the percentages of malicious URLs in the input URL stream submitted to the precisely analyzer. To our best knowledge, GuidedTracker is the first to introduce visit relations to tackle the malicious URL detection problem. The introduction of visit relations limits the scope of URL inspection and enables this approach to have the ability of self-learning. Experimental results show that the overall "toxicity" can be improved by 6.97%-50.38% compared with full inspection of access logs. Hongzhou Sha, Qingyun Liu 0001, Zhou Zhou 0007, Chao Zheng 0001 |
GLOBECOM | 2 |
| 2014 | A factor-searching-based multiple string matching algorithm for intrusion detectionabstractMultiple string matching plays a fundamental role in network intrusion detection systems. Automata-based multiple string matching algorithms like AC, SBDM and SBOM are widely used in practice, but the huge memory usage of automata prevents them from being applied to a large-scale pattern set. Meanwhile, poor cache locality of huge automata degrades the matching speed of algorithms. Here we propose a space-efficient multiple string matching algorithm BVM, which makes use of bit-vector and succinct hash table to replace the automata used in factor-searching-based algorithms. Space complexity of the proposed algorithm is O(rm2+ ΣpϵP|p|), that is more space-efficient than the classic automata-based algorithms. Experiments on datasets including Snort, ClamAV, URL blacklist and synthetic rules show that the proposed algorithm significantly reduces memory usage and still runs at a fast matching speed. Above all, BVM costs less than 0.75% of the memory usage of AC, and is capable of matching millions of patterns efficiently. Yanbing Liu 0007, Qingyun Liu 0001, Ping Liu 0001, Jianlong Tan, Li Guo 0001 |
ICC | 2 |