VLDB 2026 Research / reviewers in the wild / expert
Xiaoqing Sun
dblp:37/9540
· DBLP profile ↗
29ranked-venue papers
9as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 17 · 3 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Security and privacy · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling LLM Agent Tool Access at Cloud ScaleabstractLLM agents increasingly rely on tool calling, and the Model Context Protocol (MCP) standardizes it between agents and tool providers, reducing integration cost and driving rapid growth in tool scale. Yet a standardized interface does not make tool access work at production scale: legacy services are not MCP-callable, fast protocol evolution creates compatibility cost, large tool sets exhaust the context window, and stateful sessions complicate load balancing. We solve these with a shared control point, a centralized MCP Gateway System that makes MCP operational at cloud scale. The gateway breaks the direct-connect data plane and consolidates legacy API integration, protocol bridging, access control, and session-aware routing, while scaling out elastically at low per-call overhead. It scales agent tool access to thousands of cloud operations. Enge Song, Yueshang Zuo, Rong Wen, Jing Tie, Zhou Shao, Qiang Fu 0011, Xiaobo Xue, Luyao Zhong, Shaokai Zhang, Jiangu Zhao, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Changgang Zheng, Tian Pan 0001, Yang Song 0031, Xing Li 0007, Biao Lyu, Meng Li 0010, Haipeng Dai 0001, Guihai Chen, Shunmin Zhu |
APNet | 17 |
| 2026 | ATRU: A Stage-based Framework for Designing Ethology-Inspired Social RobotsabstractAnimal behavior (ethology) has emerged as a promising source of inspiration for social robot design. However, existing efforts have commonly resulted in isolated design instances. Our high-level understanding of the design processes for integrating ethological insights into social robot design and evaluation remains limited. To address this gap, we conducted a two-step investigation. First, we developed a stage-based framework through a systematic review, identifying six core design stages along with their descriptive dimensions. Using this framework as an analytic lens, we then analyzed design cases drawn from academic, commercial, and public contexts, deriving stage-specific considerations and actionable strategies to support designers in navigating the process. Our findings provide a conceptual scaffold for operationalizing ethology as a design resource, enabling more systematic, reflective, and transferable practices, while also surfacing new opportunities for future social robot interaction design. Xiaoqing Sun, Yanheng Li 0002, Xipei Ren |
CHI | 1 |
| 2026 | SOPSmith: Forging Executable SOPs for LLM-Driven GPU Cluster Network Diagnosis
Guoyao Yu, Xiaoqing Sun, Yangyang Shi, Yang Song 0031, Xing Li 0007, Biao Lyu, Zhenguang Liu, Qinming He |
IWQoS | 2 |
| 2026 | Bifrost: Alibaba's Next-Generation VPC Network with High-Performance Multipath Reliable Transport
Xing Li 0007, Bo Jiang 0003, Yilong Lv, Yuke Hong, Yinian Zhou, Junnan Cai, Jiayue Xu, Yunrui Hu, Zhao Gao, Enge Song, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Changgang Zheng, Yang Song 0031, Biao Lyu, Rong Wen, Zhigang Zong, Shunmin Zhu |
NSDI | 18 |
| 2026 | CStar Gateway: Augmenting Public Cloud Infrastructure for Heterogeneous Network Function Virtualization
Tian Pan 0001, Jin Ke 0005, Baohai Hu, Changgang Zheng, Enge Song, Donglin Lai, Yisong Qiao, Bengbeng Xue, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Yang Song 0031, Xionglie Wei, Biao Lyu, Rong Wen, Zhigang Zong, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu |
NSDI | 14 |
| 2026 | ZooRoute: Enhancing Cloud-Scale Network Reliability via Candidate Path Provisioning and Overlay Proactive Rerouting
Xiaoqing Sun, Xing Li 0007, Xionglie Wei, Tian Pan 0001, Yi Wang 0004, Chenhao Jia, Zhanlong Zhang, Xiaobo Xue, Jianyuan Lu, Shize Zhang, Enge Song, Yang Song 0031, Rong Wen, Biao Lyu, Yang Xu 0010, Shunmin Zhu |
NSDI | 1 |
| 2026 | ZooWear: Animal-Inspired Head-Mounted Haptic Interfaces to Augment the Zoo ExperienceabstractZoos play a crucial role in public education and wildlife engagement, yet traditional visits often lack interactive elements that foster meaningful connections between visitors and animals. In this paper, we introduce ZooWear, a head-mounted wearable featuring cartoon-style animal ears that provides haptic feedback patterns simulating an animal’s reactions to other species in the food chain. With ZooWear, we focused on examining its effectiveness in affording perspective-taking during human-animal encounters, promoting embodied knowledge retention, and enriching the zoo experiences. We first conducted a between-subject experiment in lab-based virtual zoo visits to evaluate its effectiveness in creating effective learning experiences and enhancing connections to wildlife. This was followed by a real-world zoo experiment, which showed that ZooWear promoted nature connectedness and enabled more emotionally and socially engaging experiences. Our findings highlight the potential of integrating perspective-taking into zoo experiences through animal-inspired wearables, embodied sensory feedback, and narrative-driven experiences. Pingting Chen, Bin Yu 0004, Xiaoqing Sun, Jiangnan Xia, Xipei Ren |
Int. J. Hum. Comput. Interact. | 3 |
| 2026 | Alert2Vec: Eliminating Alert Fatigue by Embedding Security Alerts Through Subgraph Learning
Songyun Wu, Xiaoqing Sun, Enhuan Dong, Jiahai Yang 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | Understanding the Long Tail Latency of TCP in Large-Scale Cloud Networks
Enge Song, Bo Jiang 0003, Yang Song 0031, Yuke Hong, Yilong Lv, Yinian Zhou, Junnan Cai, Chao Wang 0128, Yi Wang 0004, Yehao Feng, Shize Zhang, Xiaoqing Sun, Jianyuan Lu, Xing Li 0007, Biao Lyu, Zhigang Zong, Shunmin Zhu |
APNet | 16 |
| 2025 | Dense SAE Latents Are Features, Not BugsabstractSparse autoencoders (SAEs) are designed to extract interpretable features from language models by enforcing a sparsity constraint. Ideally, training an SAE would yield latents that are both sparse and semantically meaningful. However, many SAE latents activate frequently (i.e., are *dense*), raising concerns that they may be undesirable artifacts of the training procedure. In this work, we systematically investigate the geometry, function, and origin of dense latents and show that they are not only persistent but often reflect meaningful model representations. We first demonstrate that dense latents tend to form antipodal pairs that reconstruct specific directions in the residual stream, and that ablating their subspace suppresses the emergence of new dense features in retrained SAEs---suggesting that high density features are an intrinsic property of the residual space. We then introduce a taxonomy of dense latents, identifying classes tied to position tracking, context binding, entropy regulation, letter-specific output signals, part-of-speech, and principal component reconstruction. Finally, we analyze how these features evolve across layers, revealing a shift from structural features in early layers, to semantic features in mid layers, and final to output-oriented signals in the last layers of the model. Our findings indicate that dense latents serve functional roles in language model computation and should not be dismissed as training noise. Xiaoqing Sun, Alessandro Stolfo, Joshua Engels, Ben Wu 0001, Senthooran Rajamanoharan, Mrinmaya Sachan, Max Tegmark |
NeurIPS | 1 |
| 2025 | Hermes: Enhancing Layer-7 Cloud Load Balancers with Userspace-Directed I/O Event NotificationabstractLayer-7 load balancers (L7 LBs) improve service performance, availability, and scalability in public clouds. They rely on I/O event notification mechanisms such as epoll to dispatch connections from the kernel to userspace workers. However, early epoll versions suffered from the thundering herd problem. Epoll exclusive (available since Linux 4.5) mitigates this but introduces LIFO wakeups, causing connection concentration on a few workers. Reuseport (Linux 3.9) hashes connections across workers but suffers from hash collisions and lacks awareness of worker load. Since each worker serves multi-tenant traffic, inter-worker load balancing is critical to avoid worker overload and preserve tenant performance isolation. Tian Pan 0001, Enge Song, Yueshang Zuo, Shaokai Zhang, Yang Song 0031, Jiangu Zhao, Wengang Hou, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Jiao Zhang 0002, Tao Huang 0005, Biao Lyu, Xing Li 0007, Rong Wen, Zhigang Zong, Shunmin Zhu |
SIGCOMM | 9 |
| 2025 | Nezha: SmartNIC-based Virtual Switch Load SharingabstractCloud providers use SmartNIC-accelerated virtual switches (vSwitches) to offer rich network functions (NFs) for tenant VMs. Constrained by limited SmartNIC resources, it is a challenge to provide sufficient network performance for high-demand VMs. Meanwhile, we observed a significant number of idle vSwitches in the data center, which led us to consider leveraging them to build a remote resource pool for high-demand virtual NICs (vNICs). In this work, we propose Nezha, a distributed vSwitch load sharing system. Nezha reuses the existing idle SmartNICs to handle the excess load from the local SmartNIC without adding new devices. Nezha offloads stateless rule/flow tables to the remote, while keeping states locally. This eliminates the need for state synchronization, facilitating load sharing and failover. The deployment cost of Nezha is only a small fraction of that required to deploy new devices. Data collected from production show that our CPS capability bottleneck has shifted from the vSwitch to the VM kernel stack, with #concurrent flows and #vNICs increased by up to 50.4x and 40x, respectively. Xing Li 0007, Enge Song, Tian Pan 0001, Qiang Fu 0011, Yang Song 0031, Yilong Lv, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Rong Wen, Xionglie Wei, Biao Lyu, Zhigang Zong, Qinming He, Shunmin Zhu |
SIGCOMM | 12 |
| 2025 | Albatross: A Containerized Cloud Gateway Platform with FPGA-accelerated Packet-level Load BalancingabstractAlibaba Cloud's centralized gateways relied heavily on high-capacity switching ASICs, but the abrupt halt of Tofino chip evolution in Jan 2023 forced us to seek alternatives that can meet the requirements of performance, supply-chain security, code reuse, and resource efficiency. After evaluating multiple options, we developed Albatross, our 3rd gen cloud gateway based on FPGA and x86 CPUs. Albatross delivers FPGA-based packet-level load balancing to the host CPUs to prevent CPU core overload, manages large reorder buffers under high-latency jitters (100μs) during complex cloud service processing, and resolves head-of-line (HOL) blocking from packet losses or software exceptions in CPUs. To avoid being overloaded by heavy hitters due to anomalies or attacks, it also implements a two-stage rate limiter for millions of tenants with only 2MB of FPGA memory. To maximize resource utilization, Albatross uses containerization to host multiple gateway instances and designs a BGP proxy to lessen the BGP peering overhead on uplink switches caused by high-density container deployments. After hundreds of man-months of development, a single Albatross node can process 80~120Mpps of cloud network traffic with an average latency of 20μs, reducing gateway and sandbox infra costs by 50%. Jianyuan Lu, Shunmin Zhu, Tian Pan 0001, Yisong Qiao, Yang Song 0031, Wenqiang Su, Yanqiang Li, Enge Song, Shize Zhang, Xiaoqing Sun, Rong Wen, Xionglie Wei, Biao Lyu, Xing Li 0007 |
SIGCOMM | 13 |
| 2025 | ZooRoute: Enhancing Cloud-Scale Network Reliability via Overlay Proactive ReroutingabstractThis paper presents ZooRoute, a tenant-transparent, fast failure recovery service that requires no modifications to physical devices. ZooRoute leverages the overlay layer and enables traffic flows to bypass failures by altering source ports (srcPorts) in packet headers during encapsulation. To enable deployment in large-scale cloud networks, ZooRoute proposes: 1) On-demand probing to efficiently monitor a vast number of hosts while minimizing telemetry costs. 2) Table compression to record the states of numerous paths with limited on-chip resources. 3) A device-sensing mechanism to prevent unnecessary reconnections in stateful forwarding. Deployed in Alibaba Cloud for 18 months, ZooRoute has significantly improved network reliability, reducing cumulative outage time by 92.71%. Xiaoqing Sun, Xionglie Wei, Xing Li 0007, Yi Wang 0004, Chenhao Jia, Zhanlong Zhang, Jianyuan Lu, Shize Zhang, Enge Song, Yang Song 0031, Tian Pan 0001, Rong Wen, Biao Lyu, Yang Xu 0010, Shunmin Zhu |
SIGCOMM | 1 |
| 2025 | InfecBlock: Investigating the Effects of a Tower-Defense Serious Game for Increasing Epidemic-Related Health LiteracyabstractSerious game can potentially improve social awareness and health literacy related to the epidemic, where interactivity, such as strategic game elements, could play a crucial role in increasing learning engagement and motivation. In this paper, we present the user study of InfecBlock, a tower-defense game designed to facilitate individual users to acquire public health knowledge related to coronavirus disease and epidemic prevention. We employed a between-subject experiment design and collected a variety of quantitative data to examine players’ learning outcomes, engagement, and emotional responses. Our results confirmed the effectiveness of InfecBlock in improving learning performance and highlighted its potential to facilitate a more engaged and enjoyable learning experience. We discussed a set of implications for designing tower-defense serious games for supporting the improvements of public health literacy. Xiaoqing Sun, Kexin Miao, Mengchi Zhang, Xipei Ren |
Int. J. Hum. Comput. Interact. | 1 |
| 2024 | Hicclip: Sonification of Augmented Eating Sounds to Intervene Snacking BehaviorsabstractIn this paper, we present a field study on using sonification of augmented eating sounds to intervene snacking behaviors in daily routines. The sonic feedback achieved through a snack storing device named Hicclip for verifying snacking behaviors and producing augmented eating sounds. The study was conducted with nine participants who were commonly addicted to snacking. The effectiveness of the sonification was examined by comparing snack-related data and questionnaire over the three study weeks: a baseline week, a Hicclip intervention week, and a post-intervention week. We also analyzed interview results to understand user experiences and opportunities for future research. Quantitative results showed that the snacking pattern has been improved due to reduced eating duration and snack consumptions. Qualitative results suggested that Hicclip may benefit self-regulation, afford easy adoption, and support data acquisition. We discuss design implications for embodiment of augmented eating sounds for healthy snacking. Xipei Ren, Xinrui Ren, Xiaoqing Sun |
Conference on Designing Interactive Systems | 4 |
| 2024 | Understanding Network Startup for Secure Containers in Multi-Tenant Clouds: Performance, Bottleneck and OptimizationabstractIn this paper, we use empirical measurements to show that container network startup is a key factor that contributes to the slow startup of secure containers in multi-tenant clouds, especially in the scenario of serverless computing, where the issue is pronounced by high-volume concurrent container invocations. We conduct extensive and detailed analysis on existing Container Network Interface (CNI) plugins and show that even the fastest one doubles the startup time from the no-network scenario. We show that the major cause of the blowup in total startup time is that enabling networking significantly increases the contention among different startup stages, particularly for global Linux kernel locks, including the Routing Table NetLink (RTNL) mutex lock and various spin locks. We reveal that contending for these locks hinders startup performance in three ways, including directly increasing stage time, causing poor pipeline overlap and wasting CPU resources. To mitigate such kernel lock contention, we propose a multi-stage concurrency control mechanism based on Bayesian optimization to limit the concurrency of each contended stage. Our results show that this lightweight mechanism can effectively reduce the end-to-end container startup time by 18.8% with negligible extra overhead. Yunzhuo Liu, Junchen Guo, Bo Jiang 0003, Xiaoqing Sun, Yang Song 0031, Zhiyuan Hou, Biao Lyu, Rong Wen, Shunmin Zhu, Xinbing Wang |
IMC | 5 |
| 2024 | CloudPlanner: Minimizing Upgrade Risk of Virtual Network Devices for Large-Scale Cloud NetworksabstractCloud networks continuously upgrade softwarized virtual network devices (VNDs) to meet evolving tenant demands. However, such upgrades may result in unexpected failures. An intuitive idea to prevent upgrade failures is to resolve all compatibility issues before deployment, but it is impractical to replicate all deployed VND cases and test them with lots of replayed real traffic for the VND developers. As a result, the operations team takes upgrade risk to test upgrades by gradually deploying them. Although careful upgrade schedule planning is the most common method to minimize upgrade risk, to the best of our knowledge, no VND upgrade schedule planning scheme has been adequately studied for large-scale cloud networks. To fill this gap, we propose CloudPlanner, the first VND upgrade schedule planning scheme aiming to minimize the VND upgrade risk for large-scale cloud networks. CloudPlanner prioritizes upgrading VNDs that are more likely to trigger failures based on expert knowledge and historical failure-trigger VND properties and limits the number of tenants associated with simultaneously upgraded VNDs. We also propose a heuristic solver which can quickly and greedily plan schedules. Using real-world data from production environments, we demonstrate the benefits of CloudPlanner through extensive experiments. Enhuan Dong, Jiahai Yang 0001, Shize Zhang, Zejie Wang, Xiaoqing Sun, Enge Song, Jianyuan Lu, Biao Lyu, Shunmin Zhu |
INFOCOM | 9 |
| 2024 | POSEIDON: A Consolidated Virtual Network Controller that Manages Millions of Tenants via Config Tree
Biao Lyu, Enge Song, Tian Pan 0001, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Chenxiao Wang, Xiuheng Chen, Yandong Duan, Weisheng Wang, Jinpeng Long, Kunpeng Zhou, Zhigang Zong, Xing Li 0007, Guangwang Li, Peng Cheng 0001, Jiming Chen 0001, Shunmin Zhu |
NSDI | 6 |
| 2024 | A DenseNet-based feature weighting convolutional network recognition model and its application in industrial part classificationabstractAbstract Traditional warehousing typically needs machine learning or manual tagging to classify objects. However, this method is less robust and consumes a lot of labour and material resources. Based on DenseNet, this work proposes a feature weighting convolutional network recognition model and designs a set of software and hardware for data acquisition, which is applied to the efficient classification of industrial parts in warehouse management. Firstly, this work modifies DenseNet by embedding SE‐Block, and replaces the cross‐entropy loss function with the focus loss function to optimize the model structure. Secondly, a multi‐view hardware and software acquisition system is designed to complete the functions of part image acquisition, image preprocessing, model training and part recognition. Finally, an industrial parts sorting experiment was designed. Compared with the original DenseNet model, the proposed weighted convolutional network identification model showed that the accuracy of the modified model was increased by 3.09% and the convergence rate was significantly improved. The modified model proposed in this work aims to improve the recognition accuracy of industrial parts in modern warehouse management, so as to modify the classification efficiency of warehouse parts in production. Xiaoqing Sun, Yaqing Song, Yebin Lu, Qianqian Shangguan |
IET Image Process. | 2 |
| 2024 | Proactive Telemetry in Large-Scale Multi-Tenant Cloud Overlay NetworksabstractAt present, public clouds have served millions of tenants. To provide reliable services, cloud vendors need to perceive health status of the cloud network by building a telemetry system to detect possible network failures. While telemetry systems for physical networks have been extensively studied, research on telemetry systems for virtual networks is still insufficient. Different from physical networks, we conclude that building a virtual network telemetry system faces new challenges of feasibility, efficiency, and effectiveness. Specifically, we need to 1) protect privacy of tenants and adapt to heterogeneous middleboxes at the data plane; 2) handle frequent virtual network topology updates and compress large-scale measurement paths for millions of tenants at the control plane; 3) analyze telemetry results to locate network failures at the analysis plane. To address these challenges, we present Zoonet, a proactive virtual network telemetry system for multi-tenant clouds. At the data plane, Zoonet uses host agent and arp-ping to protect tenants’ privacy and defines an elegant generalization of ping and traceroute, which can work on heterogeneous middleboxes. At the control plane, Zoonet conducts update batch processing and substantial probing path pruning to lessen the overhead. At the analysis plane, Zoonet reduces noises and aggregates alerts based on temporal and spatial correlation and conducts the hop-by-hop telemetry mode to locate failures. Zoonet has been deployed in Alibaba Cloud for over two years, covering tens of cloud regions, hundreds of thousands of servers. We become increasingly reliant on Zoonet as it reduces 86% of the personnel engaged in troubleshooting. Shunmin Zhu, Jianyuan Lu, Biao Lyu, Tian Pan 0001, Shize Zhang, Xiaoqing Sun, Chenhao Jia, Xin Cheng 0022, Daxiang Kang, Yilong Lv, Fukun Yang, Xiaobo Xue, Xihui Yang, Jiahai Yang 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2024 | CloudSentry: Two-Stage Heavy Hitter Detection for Cloud-Scale Gateway Overload ProtectionabstractThe cloud vendors provide sharing resources for millions of tenants across the world to achieve economies of scale. At the same time, the cloud network keeps the performance isolation between different tenants as if they use their private dedicated resources. However, heavy hitters caused by a single tenant at cloud gateways will break such isolation, undermining the predictable performance expected by other cloud tenants. To prevent it, heavy hitter detection becomes a key concern at the performance-critical cloud gateways but faces the dilemma between fine granularity and low overhead. In this work, we presentCloudSentry, a scalable two-stage heavy hitter detection system dedicated to multi-tenant cloud gateways against such a dilemma. CloudSentry uses CPU utilization as an indicator of heavy hitters and conducts a lightweight coarse-grained detection running 24/7 to detect such CPU spikes. Then it invokes a fine-grained detection to precisely dump and analyze the potential heavy-hitter packets at the CPU spikes. After that, a more comprehensive analysis is conducted to associate heavy hitters with the cloud service scenarios and invoke a corresponding backpressure procedure. CloudSentry significantly reduces memory, computation and storage overhead compared with existing approaches. In a gateway cluster under an average traffic throughput of 251 Gbps, CloudSentry consumes only a fraction of 2%–5% CPU utilization with 8 KB run-time memory, producing only 10 MB heavy hitter logs during one month. Additionally, as it has been deployed in Alibaba Cloud for over two years, we share case studies and a lot of deployment experiences in this article. Jianyuan Lu, Tian Pan 0001, Mao Miao, Guangzhe Zhou, Yining Qi, Shize Zhang, Enge Song, Xiaoqing Sun, Huaiyi Zhao, Biao Lyu, Shunmin Zhu |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2023 | GraphIoT: Accurate IoT Identification based on Heterogeneous GraphabstractIoT devices deployed on campus and enterprise networks facilitate people's lives and work. However, these devices also bring serious network asset management and security management problems. IoT device identification is the premise to solve these problems. Although current IoT identification methods can identify devices with relatively high accuracy in ideal environments, it is difficult to accurately identify devices in real-world complex environments (e.g., campus networks, enterprise networks). Therefore, we propose to use exact features. To solve the problem of different dimensions of exact features, we creatively model the IoT identification problem as a heterogeneous graph representation learning problem and design a new representation learning algorithm. We are the first to propose an approach to accurately identify IoT devices in real-world complex environments and solve this problem through heterogeneous graphs. The evaluation shows that GraphIoT's macro F1 is on average 13.58% and 12.77% higher than the other methods on two public datasets. Linna Fan, Lin He 0004, Xiaoqing Sun, Enhuan Dong, Jiahai Yang 0001, Jinlei Lin, Guanglei Song |
IWQoS | 3 |
| 2022 | HetGLM: Lateral Movement Detection by Discovering Anomalous Links with Heterogeneous Graph Neural NetworkabstractAs a critical stage in the Advanced Persistent Threat (APT) lifecycle, lateral movement (LM) has become a major concern in cybersecurity due to its stealthy nature. Recent authentication graph-based LM detection systems have achieved promising results. However, these methods have some unpractical requirements on data collection and model deployment, which severely affects their performance in real-world scenarios. In this paper, we propose HetGLM, a more accurate and practical LM detection system. Specifically, to fully explore the scenario, HetGLM constructs a heterogeneous graph with various network entities like users, devices, processes, etc. On this basis, we design MADR, a Graph neural network (GNN)-based anomaly link detection algorithm, to spot lateral movements. With the metapath-based sampling strategy, attention mechanism, the dual-decoder structure, and a mutual information regularization term, MADR can detect anomaly links on heterogeneous graphs, requiring neither labeled or purely benign training datasets nor manually preset thresholds. We implement a prototype of HetGLM and evaluate its performance via comprehensive experiments over public datasets. Comparison results show that HetGLM outperforms the state-of-the-art approaches in accuracy and practicality. Xiaoqing Sun, Jiahai Yang 0001 |
IPCCC | 1 |
| 2020 | Domain-Embeddings Based DGA Detection with Incremental Training MethodabstractDGA-based botnet, which uses Domain Generation Algorithms (DGAs) to evade supervision, has become a part of the most destructive threats to network security. Over the past decades, a wealth of defense mechanisms focusing on domain features have emerged to address the problem. Nonetheless, DGA detection remains a daunting and challenging task due to the big data nature of Internet traffic and the potential fact that the linguistic features extracted only from the domain names are insufficient and the enemies could easily forge them to disturb detection. In this paper, we propose a novel DGA detection system which employs an incremental word-embeddings method to capture the interactions between end hosts and domains, characterize time-series patterns of DNS queries for each IP address and therefore explore temporal similarities between domains. We carefully modify the Word2Vec algorithm and leverage it to automatically learn dynamic and discriminative feature representations for over 1.9 million domains, and develop an simple classifier for distinguishing malicious domains from the benign. Given the ability to identify temporal patterns of domains and update models incrementally, the proposed scheme makes the progress towards adapting to the changing and evolving strategies of DGA domains. Our system is evaluated and compared with the state-of-art system FANCI and two deep-learning methods CNN and LSTM, with data from a large university’s network named TUNET. The results suggest that our system outperforms the strong competitors by a large margin on multiple metrics and meanwhile achieves a remarkable speed-up on model updating. Xiaoqing Sun, Jiahai Yang 0001 |
ISCC | 2 |
| 2020 | HGDom: Heterogeneous Graph Convolutional Networks for Malicious Domain DetectionabstractAs a fundamental component of the Internet, Domain Name System (DNS) is widely abused by attackers in various cybercrimes, making malicious domain detection an essential task in network defenses. However, some well-crafted attacks with tricky techniques can not only bypass blacklists but also make some machine learning-based detection systems infeasible. In this paper, we design HGDom, an accurate and robust malicious domain detection system based on a heterogeneous graph convolutional network method. First, we jointly analyze domain features as well as the complex relations among domains, clients, and IP addresses. To capture richer information, we introduce a Heterogeneous Information Network (HIN) to model the DNS scene. Then, we propose a novel representation method named MAGCN. With a meta-path-based attention mechanism, it can handle node features and the graph structure in HIN at the same time. To our best knowledge, this is the first work to apply GCN in cyber security analysis. Comprehensive experiments over DNS data from TUNET and CERNET2 are conducted to validate the effectiveness and superiority of our proposed methods. The comparison results show that HGDom outperforms state-of-the-art approaches with promising performance. Besides, the system is decided to be deployed in production to assist with network security management for CERNET2. Xiaoqing Sun, Jiahai Yang 0001 |
NOMS | 1 |
| 2020 | Deepdom: Malicious domain detection with scalable and heterogeneous graph convolutional networks
Xiaoqing Sun, Jiahai Yang 0001 |
Comput. Secur. | 1 |
| 2019 | D3N: DGA Detection with Deep-Learning Through NXDomain
Mingkai Tong, Xiaoqing Sun, Jiahai Yang 0001, Hui Zhang 0052, Shuang Zhu |
KSEM (1) | 2 |
| 2019 | HinDom: A Robust Malicious Domain Detection System based on Heterogeneous Information Network with Transductive Classification
Xiaoqing Sun, Mingkai Tong, Jiahai Yang 0001 |
RAID | 1 |