VLDB 2026 Research / reviewers in the wild / expert
Guanglei Song
dblp:75/9109
· DBLP profile ↗
32ranked-venue papers
6as first author
30since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 21 · 5 first-author · 19 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Understanding the IPv6 Address Usage Strategies of Top Internet Services
Lin He 0004, Zedong Jia, Daguo Cheng, Jinlong E, Yuhan Du, Guanglei Song, Ying Liu 0024, Xingang Shi, Shenglin Zhang, Jiahai Yang 0001, Mingwei Xu 0001 |
ICC | 6 |
| 2026 | LGTSR: General Textual Semantic Representation and Topology-Aware Learning via Large Language Models
Jinfeng Fang, Guanglei Song, Haiting Tan |
ICIC (23) | 4 |
| 2026 | Breaking the Seed Barrier: Discovering Active IPv6 Addresses in Seedless Scenarios
Wenjian Zhang, Guanglei Song, Binkai Ma, Lin He 0004, Songyun Wu, Jiahai Yang 0001 |
INFOCOM | 2 |
| 2026 | ISP or Customer? Inferring the Ownership of Public IPs of Non-Cooperative Satellite Internet via Internet Measurements
Enhuan Dong, Jiahai Yang 0001, Wenjian Zhang, Guanglei Song, Kexin Qiang, Hui Zhang 0141, Xiaowen Quan |
IWQoS | 5 |
| 2026 | MSRNet: Multi-scale Spatiotemporal Retention Network for Motor Imagery Recognition
Shasha Mo, Guanglei Song, Shuo Tan |
KSEM (3) | 3 |
| 2026 | AddrProbe: An Internet-Wide Active IPv6 Address Probing System With Limited SeedsabstractWith the large-scale deployment of IPv6, it is becoming more and more important to probe active IPv6 addresses on the global Internet. However, the vast address space and the random distribution of active addresses make the probing process full of challenges, especially for the probing of IPv6 prefixes without seed addresses. Furthermore, the widespread existence of IPv6 aliased prefixes also causes significant trouble for probing. In this paper, we presentAddrProbe, an active IPv6 address probing system, which dynamically probes all global routing prefixes based on learned fine-grained address patterns from limited seed addresses and quickly detects aliased prefixes during probing. The evaluation results show thatAddrProbeachieves a hit rate of 23%-45% with all routing prefixes announced by the BGP system, which is 6.6-13× that of current state-of-the-art approaches (no more than 4%). Moreover, we find 1.2×1033aliased addresses characterized by the detected aliased prefixes, covering 6,412 routing prefixes, which is a 107× and 5.9× improvement over existing methods, respectively. Finally, an IPv6 Hitlist is constructed based on the long-term probing results, which contains 562M addresses covering 190K routing prefixes and 29K ASes. These widely distributed addresses are meaningful for analyzing IPv6 address assignments and some other IPv6 measurement activities. Daguo Cheng, Lin He 0004, Qilei Yin, Guangxing Han, Boran Jin, Ying Liu 0024, Guanglei Song, Jinlong E, Tiankai Yang 0001, Jiahai Yang 0001 |
IEEE Trans. Netw. | 8 |
| 2025 | ScannerGrouper: A Generalizable and Effective Scanning Organization Identification System Toward the Open WorldabstractIn recent years, many scanning organizations deploy large numbers of scanners to actively probe the Internet. Identifying the organizations behind these scanners is of significant value. The problem of analyzing the sources of scanners has been investigated in various studies. However, as far as we know, the problem of effectively and generally identifying scanner organizations in real-world scenarios remains unsolved. Enhuan Dong, Jiyuan Han, Hui Zhang 0141, Lianyi Sun, Supei Zhang, Guanglei Song, Xiaowen Quan, Jiahai Yang 0001 |
CCS | 10 |
| 2025 | TopoMiner: Efficient IPv6 Topology DiscoveryabstractTopology discovery can be used to obtain the connectivity and operational status of network devices by discovered interfaces. By performing topology discovery on networks, administrators can better understand and manage networks and improve network reliability and security. However, the large IPv6 address space, sparse address distribution, and unknown address assignment policies make it infeasible to simply and roughly probe the IPv6 network topology. To this end, we design an efficient IPv6 interface-level topology discovery method called TopoMiner. It performs multiple rounds of probing the IPv6 topology with the given set of IPv6 prefixes and locates the high-density interface regions based on the results of each round of topology discovery. TopoMiner then performs prefix expansion and address generation for the high-density interface address regions. At the same time, TopoMiner also uses some of the prefixes that have been probed to regenerate the target addresses. In this way, TopoMiner can dynamically adjust the breadth and depth of topology discovery and thus complete multiple rounds of topology discovery within the packetsending budget. In real-word probing, TopoMiner achieves a$3 \sim 7 \times$enhancement in probing efficiency compared to state-of-the-art topology discovery methods. Hongwei Li 0021, Lin He 0004, Guanglei Song, Daguo Cheng, Jiahai Yang 0001, Ying Liu 0024 |
ICC | 5 |
| 2025 | SubRecon: Efficient Internet-Wide IPv6 Subnet Discovery and Its ApplicationsabstractThe vastness of the IPv6 address space has led to the common practice of allocating prefixes to end users rather than individual addresses. Users can assign these prefixes as a single subnet or divide them into multiple subnets for different purposes. Allocation strategies vary significantly in terms of prefix granularity, and identifying the actual granularity of subnet assignments is crucial for improving measurement efficiency, accuracy, and for better IPv6 network management. However, no existing method can discover IPv6 subnets at an Internet-Wide scale.To this end, we propose SubRecon, an Internet-Wide IPv6 subnet discovery system. SubRecon consists of two key phases: subnet delimitation and target expansion. In the subnet delimitation phase, we perform a systematic scan across the entire IPv6 address space without relying on any existing seed dataset. This phase adopts a top-down approach, probing prefixes from the shortest to the longest in a hierarchical manner. At each level, we recursively refine prefixes and discard sub-prefixes that do not meet the convergence condition. This pruning strategy eliminates redundant probes in unallocated regions, significantly reducing the search space and improving probing efficiency. To further improve coverage, the target expansion phase leverages the active address dataset as an auxiliary input. It identifies active addresses not covered by previously discovered subnets, expands them into new candidate target prefixes, and performs another round of subnet delimitation. This helps enhance the completeness and coverage of the final discovered subnet set. Experimental results show that SubRecon discovers 8,381,974 IPv6 subnets across 14,147 autonomous systems, and the resulting subnet list can serve as high-quality input for topology discovery. Additionally, during the subnet discovery process, SubRecon identifies a large number of last-hop router interfaces, discovering approximately 10 million more than the current state-of-the-art methods. Ying Liu 0024, Lin He 0004, Yifan Yang 0009, Xiaoyi Shi, Daguo Cheng, Chentian Wei, Yun Fan, Guanglei Song |
ICNP | 9 |
| 2025 | Poster: TopoHunter: Enabling Efficient and High-Coverage Active IPv6 Topology DiscoveryabstractWe introduce TopoHunter, an efficient IPv6 Internet topology discovery system. The central concept of TopoHunter is to allocate more probing resources to target prefix spaces that yield greater topological benefits, as well as to their surrounding areas. To achieve this, we design a feedback-based target generation module comprised of a Target Prefix Probing Value Forest that maintains the estimated probing values of hierarchical target prefix spaces. Our system has successfully discovered the most extensive and complete IPv6 topology map to date, comprising over 144 million router interfaces and 251 million edges, covering 72.83% of autonomous systems and 43.36% of routing prefixes announced by the BGP system. Lin He 0004, Hongwei Li 0021, Guanglei Song, Wentong Wang, Daguo Cheng, Enhuan Dong, Chenglong Li 0006, Hui Zhang 0141, Jinlong E, Ying Liu 0024, Jiahai Yang 0001 |
IMC | 4 |
| 2025 | 6RV: Incremental Learning-Based Continuous Identification of IPv6 Router VendorsabstractThe growth of IPv6 networks has led to an expanding number of network routers, but there is not enough research on vendors of these devices. Existing router vendor identification algorithms are based on IPv4, and these static analysis algorithms cannot adapt to dynamically changing IPv6 networks. In this paper, we develop 6RV, a framework for IPv6 router vendors continuous identification based on incremental learning. First, we count the addresses of newly discovered router interfaces every month and obtain router device fingerprints through active probing, and then analyze these data fingerprints through an incremental learning approach to identify the router vendors of new nodes. We validate our identification framework on the ITDK dataset over a period of 8 months, obtaining more than 500K router vendor labels with 94% correctness. Finally, we also analyze the IPv6 router vendor dataset from different perspectives and draw some interesting conclusions. Shenao Li, Jiahai Yang 0001, Enhuan Dong, Chenglong Li 0006, Lin He 0004, Hui Zhang 0052, Guanglei Song |
NOMS | 8 |
| 2025 | IPdb: A High-Precision IP Level Industry Categorization of Web ServicesabstractIP addresses with web services are crucial in the Internet ecosystem. Classifying these addresses by industry and organization offers valuable insights into the entities utilizing them, enabling more efficient network management and enhanced security. Previous work in website classification and Internet management struggles to offer an IP-level perspective of the industries of web services due to their limited industry categories or potential industry inconsistencies between IP address owners and AS owners. To this end, we present IPdb, an IP-level industry categorization dataset. To construct the dataset, we developed LLMIC, a Large Language Model-based Industry Categorization framework with a precision of nearly 96%. IPdb serves as a labeled database for future endeavors in developing IP-level industry classifiers, encompassing over 200 million IP addresses. Furthermore, our study indicates that 30% ~ 50% of organizations within critical infrastructure industries deploy web servers across multiple ASes. Our study also validates the problem of mismatched granularity in industry categorization at the AS level with 87.83% ASes in IPv4 and 72.96% ASes in IPv6 containing IP addresses from different industries. Guanglei Song, Jiahai Yang 0001, Songyun Wu, Jinlei Lin, Lin He 0004, Chenglong Li 0006 |
WWW | 2 |
| 2025 | HGExplainer: Heterogeneous Graph Explainer for IoT Device IdentificationabstractIoT device identification is vital for network asset and security management. However, existing methods use statistical features that can not identify IoT devices accurately in complex network environments.GraphIoTproposes using non-statistical features and building a heterogeneous graph neural network to identify IoT devices accurately. However, heterogeneous graph neural networks lack interpretability, which reduces trust in the model. Besides, it is difficult to deploy on resource-constrained devices, limiting the broad application of IoT device identification. To make IoT device identification interpretable, easy to deploy, and with high accuracy, we get the interpretation results ofGraphIoTthrough interpretability and further build the rule set based on the interpretation results. Considering there is no suitable interpreter forGraphIoTwith many nodes and edges, we proposeHGExplainer, which reduces the time complexity by splitting the interpretation target into important relation solving and edge solving and uses a novel solution method, ExpandTree. Then, we also designed a rule extractor, which can build rule sets based on the interpretation results. Experimental results on Yourthings and UNSW datasets show thatHGExplainercan build high fidelity, concise sample-level explanations in less than 3 seconds, and the established rule set can precisely identify IoT devices. Linna Fan, Xuan Shen, Guanglei Song, Chaocan Xiang, Duohe Ma, Yongfeng Huang 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | 6Vision: Image-Encoding-Based IPv6 Target Generation in Few-Seed ScenariosabstractEfficient global Internet scanning is crucial for network measurement and security analysis. While existing target generation algorithms verify remarkable performance in largescale detection, their efficiency notably diminishes in few-seed scenarios. This decline is primarily attributed to the intricate configuration rules and sampling bias of seed addresses. Moreover, instances where BGP prefixes have few seed addresses are widespread, constituting$63.65 \%$of occurrences. We introduce 6 Vision to tackle this challenge by introducing a novel approach to encoding IPv6 addresses into images, facilitating comprehensive analysis of intricate configuration rules. Through feature stitching, 6 Vision not only improves the learnable features but also amalgamates addresses associated with configuration patterns for enhanced learning. Moreover, it integrates an environmental feedback mechanism to refine model parameters based on identified active addresses, thereby alleviating the sampling bias inherent in seed addresses. As a result, 6Vision achieves high-accuracy detection even in few-seed scenarios. The HitRate of 6 Vision is improved by$181 \% \sim 2,490 \%$compared to existing algorithms, while the CoverNum is$1.18 \sim 11.20$times that of them. Additionally, 6Vision can function as a preliminary detection module for existing algorithms, yielding a conversion gain (CG) ranging from$242 \% \sim 2,081 \%$. Ultimately, we achieve a conversion rate (CR) of$28.97 \%$for few-seed scenarios. We enrich the IPv6 hitlist, not only enhancing current target generation algorithms for large-scale address detection in few-seed scenarios but also effectively supporting IPv6 network measurement and security analysis. Wenjian Zhang, Guanglei Song, Lin He 0004, Jinlei Lin, Songyun Wu, Chenglong Li 0006, Jiahai Yang 0001 |
ICNP | 2 |
| 2024 | IoTa: Fine-Grained Traffic Monitoring for IoT Devices via Fully Packet-Level ModelsabstractWith Internet-of-Things (IoT) devices gaining popularity, dedicated monitoring systems which accurately detect intrusion traffic for them are in high demand. Existing methods mainly use statistical spatial-temporal traffic features and machine learning models. Their practicality has been limited due to the lack of detection ability for stealthy and tricky attacks, diagnostic utility and long-term performance. To address these problems and motivated by the simplicity of mini IoT devices, we propose to construct fully packet-level models to profile traffic patterns for IoT devices by constructing automaton for short flow and long flow, where the length and direction of each packet are the representative features. We apply these fine-grained models to design and develop a traffic monitoring system, namelyIoTa, to detect intrusion traffic for IoT devices.IoTamatches the ongoing traffic with patterns extracted from normal traffic traces. With visible and interactive traffic profiles,IoTacan generate interpretable alerts and is available for long-term use under reasonable human efforts. Evaluations on dozens of common IoT devices show thatIoTacan achieve excellent detection accuracy (nearly perfect recalls and always over 0.999 precisions) for various intrusion traffic covering the complete kill chains. Incorrect detection results can be compensated for by error recovery mechanisms and the understandable alert context can be used by the operator to enhance the system. The diagnostic utility and little alert weariness are recognized by the experienced operators. Chenxin Duan, Sainan Li, Guanglei Song, Chenglong Li 0006, Jiahai Yang 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2024 | ProbeGeo: A Comprehensive Landmark Mining Framework Based on Web ContentabstractIP geolocation is essential for various location-aware Internet applications. High-quality IP geolocation landmarks play a decisive role in IP geolocation accuracy. However, the previous research works focusing on mining landmarks from the Internet are hampered by limited quantity, poor coverage, and insufficient landmark quality. In this paper, we present a new framework called ProbeGeo to mine high-quality landmarks automatically. We divide landmarks into common landmarks and probe landmarks, providing systematic mining methods based on online retrieval and web content. ProbeGeo expands traditional common landmarks by taking advantage of the exposure of multiple IoT (Internet of Things) devices on the Internet, mining them based on search engines and webpage contents. Common landmarks, consisting of multi-type devices, significantly improve landmark quantity and coverage. Furthermore, ProbeGeo establishes a methodology for acquiring new probe landmarks from Internet VPs (Vantage Points) webpages, extracting geographical locations from heterogeneous webpages and utilizing active probe functions. Probe landmarks enhance landmark quality and functions, bringing new geolocation frameworks and breaking through the geolocation accuracy bottleneck. We develop the ProbeGeo as a continuously running system and conduct real-world experiments to validate its efficacy. Our results show that ProbeGeo can detect 89,849 high-quality landmarks, including 6,874 probe landmarks and 82,975 common landmarks. ProbeGeo landmarks are about 10x more than existing work, distributed in 181 countries and 7,094 cities. ProbeGeo landmarks cover more than 8 types of devices, and more than 60% of them remain stable over one month. Moreover, the landmark accuracy of more than 58% of ProbeGeo landmarks is above street level, which has not been achieved in previous works. ProbeGeo can provide geolocation services with higher landmark accuracy and broader coverage by correlating a large scale of landmarks. Jinlei Lin, Chenglong Li 0006, Guanglei Song, Linna Fan, Jiahai Yang 0001 |
IEEE/ACM Trans. Netw. | 4 |
| 2024 | PMap: Reinforcement Learning-Based Internet-Wide Port ScanningabstractInternet-wide scanning is a commonly used research technique in various network surveys, such as measuring service deployment and security vulnerabilities. However, these network surveys are limited to the given port set, not comprehensively obtaining the real network landscape, and even misleading survey conclusions. In this work, we introduce PMap, a port scanning tool that efficiently discovers the most open ports from all 65K ports in the whole network. PMap uses the correlation of ports to build an open port correlation graph of each network, using a reinforcement learning framework to update the correlation graph based on feedback results and dynamically adjust the order of port scanning. Compared to current port scanning methods, PMap performs better on hit rate, coverage, and intrusiveness. Our experiments over real networks show that PMap can find 90% open ports by only scanning 125 ports (90%@125) to each address, which is 99.3% less than the state-of-the-art port scanning methods. It reduces the number of scanned ports to decrease the intrusive nature of port scanning. In addition, PMap is highly parallel and lightweight. It scans 500 networks in parallel, achieving a port recommendation rate of up to 18 million per second, consuming only 7GB of memory. PMap is the first effective practice for scanning open ports using reinforcement learning. It bridges the gap of existing scanning tools and effectively supports subsequent service discovery and security research. Guanglei Song, Lin He 0004, Jinlei Lin, Linna Fan, Jiahai Yang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | AddrMiner: A Fast, Efficient, and Comprehensive Global Active IPv6 Address Detection SystemabstractFast Internet-wide scanning is essential for network situational awareness and asset evaluation. However, the vast IPv6 address space makes brute-force scanning infeasible. Despite advancements in state-of-the-art methods, they do not work in seedless regions and suffer low detection efficiency and speed in regions with known active IPv6 addresses (i.e., seed addresses). Moreover, the collected active address list (i.e., IPv6 hitlist) with low coverage cannot truly represent the active IPv6 address landscape of the Internet. This paper introduces AddrMiner, a fast, efficient, and comprehensive global active IPv6 address detection system. We design a systematic active IPv6 address detection strategy that divides the IPv6 space into two detection scenarios based on the presence or absence of seed addresses to discover active IPv6 addresses from scratch and from few to many. In the seedless regions, we present AddrMiner-N, leveraging a multi-level association policy to probe active addresses. It fills the gap of address detection in seedless regions and successfully discovers active addresses in 39,899 BGP prefixes without seed addresses, with a$1.03\times $higher hit rate,$30\sim 911\times $higher speed, and$2.7\times $broader coverage, compared to existing solutions. In the regions with seed addresses, our method AddrMiner-S dynamically generates target addresses using reinforcement learning. Compared to state-of-the-art methods, AddrMiner-S achieves an impressive 56.3% hit rate and a discovery speed of 839.0/s, which is$1.9\sim 2153\times $and$1.5\sim 755\times $of existing works, respectively. Finally, we deploy AddrMiner and discover 2.1B active IPv6 addresses, including 1.7B de-aliased active addresses and 0.4B aliased addresses, through continuous probing for three years. Guanglei Song, Lin He 0004, Feiyu Zhu 0002, Jinlei Lin, Wenjian Zhang, Linna Fan, Chenglong Li 0006, Jiahai Yang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | GraphIoT: Accurate IoT Identification based on Heterogeneous GraphabstractIoT devices deployed on campus and enterprise networks facilitate people's lives and work. However, these devices also bring serious network asset management and security management problems. IoT device identification is the premise to solve these problems. Although current IoT identification methods can identify devices with relatively high accuracy in ideal environments, it is difficult to accurately identify devices in real-world complex environments (e.g., campus networks, enterprise networks). Therefore, we propose to use exact features. To solve the problem of different dimensions of exact features, we creatively model the IoT identification problem as a heterogeneous graph representation learning problem and design a new representation learning algorithm. We are the first to propose an approach to accurately identify IoT devices in real-world complex environments and solve this problem through heterogeneous graphs. The evaluation shows that GraphIoT's macro F1 is on average 13.58% and 12.77% higher than the other methods on two public datasets. Linna Fan, Lin He 0004, Xiaoqing Sun, Enhuan Dong, Jiahai Yang 0001, Jinlei Lin, Guanglei Song |
IWQoS | 8 |
| 2023 | Which Doors Are Open: Reinforcement Learning-based Internet-wide Port ScanningabstractInternet-wide scanning is a commonly used research technique in various network surveys, such as measuring service deployment and security vulnerabilities. However, these network surveys are limited to the given port set, not comprehensively obtaining the real network landscape, and even misleading survey conclusions. In this work, we introduce PMap, a port scanning tool that efficiently discovers the majority of open ports from all 65K ports in the whole network. PMap uses the correlation of ports to build an open port correlation graph of each network, using a reinforcement learning framework to update the correlation graph based on feedback results and dynamically adjust the order of port scanning. Compared to current port scanning methods, PMap achieves better performance on hit rate, coverage, and intrusiveness. Our experiments over real-world networks show that PMap can find 90% open ports by only scanning 125 ports (90% @125) to each active address with 136× less than the state-of-the-art port probing methods. PMap reduces the number of scanned ports to decrease the intrusive nature of port scanning. PMap is the first effective practice for scanning open ports using reinforcement learning. It bridges the gap of existing scanning tools and effectively supports subsequent service discovery and security research. Guanglei Song, Lin He 0004, Tianyun Zhao, Yirui Luo, Yichao Wu, Linna Fan, Chenglong Li 0006, Jiahai Yang 0001 |
IWQoS | 1 |
| 2023 | Your Router is My Prober: Measuring IPv6 Networks via ICMP Rate Limiting Side Channels
Long Pan, Jiahai Yang 0001, Lin He 0004, Leyao Nie, Guanglei Song, Yaozhong Liu |
NDSS | 6 |
| 2022 | PerfTrace: A New Multi-metric Network Performance Monitoring ToolabstractWe present PerfTrace, an end-to-end tool for efficient, real-time, and multi-metric network performance monitoring. PerfTrace provides a high integration of different existing measurement functions, supporting the measurement of essential metrics such as latency, jitter, packet loss, and available bandwidth. More importantly, innovative schemes and algorithms are proposed to address the weaknesses of existing tools.After conducting comprehensive evaluations, we find that (i) PerfTrace measures one-way and two-way latency, jitter, and packet loss ∼9.4× faster and ∼3.6× more data-efficiently; (ii) PerfTrace measures available bandwidth in our testbed with minimal mean relative error (5.22%), outperforming all the tools compared (ranging from 8.17% to 37.24%). Meanwhile, PerfTrace consumes a more constant percentage of bandwidth resources than other tools when monitoring available bandwidth. PerfTrace’s data overhead is always only about 1/600 of the total bandwidth for a measurement frequency once per minute. Yaozhong Liu, Long Pan, Chenglong Li 0006, Lin He 0004, Yirui Luo, Guanglei Song, Jiahai Yang 0001 |
CNSM | 6 |
| 2022 | Towards a Behavioral and Privacy Analysis of ECS for IPv6 DNS ResolversabstractThe Domain Name System (DNS) is critical to Internet communications. EDNS Client Subnet (ECS), a DNS extension, allows recursive resolvers to include client subnet information in DNS queries to improve CDN end-user mapping, extending the visibility of client information to a broader range. Major content delivery network (CDN) vendors, content providers (CP), and public DNS service providers (PDNS) are accelerating their IPv6 infrastructure development. With the increasing deployment of IPv6-enabled services and DNS being the most foundational system of the Internet, it becomes important to analyze the behavioral and privacy status of IPv6 resolvers. However, there is a lack of research on ECS for IPv6 DNS resolvers.In this paper, we study the ECS deployment and compliance status of IPv6 resolvers. Our measurement shows that 11.12% IPv6 open resolvers implement ECS. We discuss abnormal noncompliant scenarios that exist in both IPv6 and IPv4 that raise privacy and performance issues. Additionally, we measured if the sacrifice of clients’ privacy can enhance IPv6 CDN performance. We find that in some cases ECS helps end-user mapping but with an unnecessary privacy loss. And even worse, the exposure of client address information can sometimes backfire, which deserves attention from both Internet users and PDNSes. Leyao Nie, Lin He 0004, Guanglei Song, Chenglong Li 0006, Jiahai Yang 0001 |
CNSM | 3 |
| 2022 | Both Efficient and Accurate: A Large-scale One-way Delay Measurement SchemeabstractOne-way delay (OWD) is one of the essential network performance metrics. In large-scale resilient overlay networks (RONs), OWD measurements can be used for shortest path selection and troubleshooting. However, OWD measurements remain difficult because of the need for precise time synchronization. Especially in large-scale networks, clock synchronization of all nodes has always been a considerable challenge. Therefore, in many cases, people use half of the round-trip time (RTT/2) as a rough substitute for the OWD. This paper presents an efficient and easy-to-deploy scheme for large-scale OWD measurements with the algorithm ClockConverger at its core. The scheme consists of three steps: Firstly, we perform low-precision time synchronization for all the measured nodes relying on network time protocol daemons (ntpd); Then, we use the open-source tool OWPing to perform OWD measurements; Finally, we correct the errors of the measured OWDs with our proposed ClockConverger. The theory and experiments show that our scheme's accuracy is significantly better than RTT/2. Meanwhile, the complexity of ClockConverger is$O(n^{2})$, which is much lower than the exponential complexity of the existing Maximum-Entropy algorithm. Yaozhong Liu, Jiahai Yang 0001, Long Pan, Lin He 0004, Jinlei Lin, Guanglei Song, Chenglong Li 0006 |
GLOBECOM | 7 |
| 2022 | What Causes Delay Asymmetry: A Large-scale One-way Delay Measurement and Empirical StudyabstractIn global communications, severe one-way delay (OWD) asymmetry often occurs. Due to the difficulties of OWD measurement (need to control both ends and synchronize their clocks), now RTT/2 is commonly used to estimate OWD. However, OWD asymmetry can lead to large errors in the halving RTT method, which in turn affects the end-to-end quality of service (QoS) guarantees. In this paper, we investigate OWD asymmetry through large-scale OWD measurements on a global scale. The measurements show that more than 11% of network paths have OWDs with a relative difference of more than 10% compared to RTT/2. By analyzing the measurement results in depth, we try to explain why the delay asymmetry occurs. We find that 67% is caused by hop inflation or a significant increase in propagation distance, and 33% is caused by variable queuing delays. We also find AS-level paths between node pairs with significant delay asymmetry are much more likely (~ 10 ×) to violate the well-known valley-free rule. Yaozhong Liu, Jiahai Yang 0001, Long Pan, Lin He 0004, Jinlei Lin, Guanglei Song, Chenglong Li 0006 |
GLOBECOM | 7 |
| 2022 | Monitoring Smart Home Traffic under Differential PrivacyabstractRecent years have witnessed the proliferation of smart home ecosystems. Well-characterized traffic generated by smart home devices has promoted the development of security enhancing techniques for smart homes but exposes users to the privacy disclosure risk at the same time. Malicious eavesdroppers can infer working states of smart home devices and user activities based on spatial-temporal traffic characteristics. Existing countermeasures towards this kind of side channel attack ignore the utility of smart home traffic profiles and signatures for network management and attempt to completely eliminate them. In this paper, we give a comprehensive study on the trade offs between the usability of smart home traffic for security monitoring and its privacy threat. We propose to monitor the smart homes under differential privacy. Based on our solution, decoy traffic can be generated in a controlled manner so as to confound the attackers without disturbing the running monitoring systems. We prototyped our proposal and demonstrate its effectiveness empirically. An interview study is also conducted to learn the user acceptance of the proposed privacy preserving mechanism. Chenxin Duan, Guanglei Song, Jiahai Yang 0001 |
NOMS | 4 |
| 2022 | AddrMiner: A Comprehensive Global Active IPv6 Address Discovery System
Guanglei Song, Jiahai Yang 0001, Lin He 0004, Chenxin Duan, Yaozhong Liu, Zhongxiang Sun |
USENIX ATC | 1 |
| 2022 | ByteIoT: A Practical IoT Device Identification System Based on Packet Length DistributionabstractA tremendous amount of Internet-of-Things (IoT) devices have been deployed in recent years, bringing new challenges for network management and cyber security. It is important for network managers to know what types of IoT devices are connecting to the network. Despite much research efforts, previous works place more emphasis on accuracy but ignore some other performance indicators also in high demand, like efficiency, robustness, adaptability to special scenarios and extensibility for new devices. In this paper, we propose a practical IoT device identification system, namely ByteIoT, based on a simple but well-organized traffic feature, i.e., the frequency distribution of bidirectional packet lengths. ByteIoT applies k-nearest neighbors algorithm as the classifier to gain extensibility and adaptability. We evaluate ByteIoT on several datasets and the results show that ByteIoT can outperform other state-of-the-art methods in the aspects of accuracy, efficiency, extensibility and adaptability. Chenxin Duan, Guanglei Song, Jiahai Yang 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2022 | DET: Enabling Efficient Probing of IPv6 Active AddressesabstractFast IPv4 scanning significantly improves network measurement and security research. Nevertheless, it is infeasible to perform brute-force scanning of the IPv6 address space. Alternatively, one can find active IPv6 addresses through scanning the candidate addresses generated by state-of-the-art algorithms. However, the probing efficiency of such algorithms is often very low. In this paper, our objective is to improve the probing efficiency of IPv6 addresses. We first perform a longitudinal active measurement study and build a high-quality dataset, hitlist, including more than 1.95B IPv6 addresses distributed in 58.2K BGP prefixes and collected over 17 months period. Different from the previous works, we probe the announced BGP prefixes using a pattern-based algorithm. This results in a dataset without uneven address distribution and low active rates. Further, we propose an efficient address generation algorithm, DET, which builds a density space tree to learn high-density address regions of the seed addresses with linear time complexity and improves the active addresses’ probing efficiency. We then compare our algorithm DET against state-of-the-art algorithms on the public hitlist and our hitlist by scanning 50M addresses. Our analysis shows that DET increases the de-aliased active address ratio and active address (including aliased addresses) ratio by 10%, and 14%, respectively. Furthermore, we develop a fingerprint-based method to detect aliased prefixes. The proposed method for the first time directly verifies whether the prefix is aliased or not. Our method finds that 10.64% of the public aliased prefixes are false positive. Guanglei Song, Jiahai Yang 0001, Lin He 0004, Jinlei Lin, Long Pan, Chenxin Duan, Xiaowen Quan |
IEEE/ACM Trans. Netw. | 1 |
| 2021 | Deception Maze: A Stackelberg Game-Theoretic Defense Mechanism for Intranet ThreatsabstractThe intranets in modern organizations are facing severe data breaches and critical resource misuses. By reusing user credentials from compromised systems, Advanced Persistent Threat (APT) attackers can move laterally within the internal network. A promising new approach called deception technology makes the network administrator (i.e., defender) able to deploy decoys to deceive the attacker in the intranet and trap him into a honeypot. Then the defender ought to reasonably allocate decoys to potentially insecure hosts. Unfortunately, existing APT-related defense resource allocation models are infeasible because of the neglect of many realistic factors.In this paper, we make the decoy deployment strategy feasible by proposing a game-theoretic model called the APT Deception Game to describe interactions between the defender and the attacker. More specifically, we decompose the decoy deployment problem into two subproblems and make the problem solvable. Considering the best response of the attacker who is aware of the defender’s deployment strategy, we provide an elitist reservation genetic algorithm to solve this game. Simulation results demonstrate the effectiveness of our deployment strategy compared with other heuristic strategies. Jieling Liu, Jiahai Yang 0001, Bo Wang 0066, Lin He 0004, Guanglei Song |
ICC | 6 |
| 2020 | Towards the Construction of Global IPv6 Hitlist and Efficient Probing of IPv6 Address SpaceabstractFast IPv4 scanning has made sufficient progress in network measurement and security research. However, it is infeasible to perform brute-force scanning of the IPv6 address space. We can find active IPv6 addresses through scanning candidate addresses generated by the state-of-the-art algorithms, whose probing efficiency of active IPv6 addresses, however, is still very low. In this paper, we aim to improve the probing efficiency of IPv6 addresses in two ways. Firstly, we perform a longitudinal active measurement study over four months, building a high-quality dataset called hitlist with more than 1.3 billion IPv6 addresses distributed in 45.2k BGP prefixes. Different from previous work, we probe the announced BGP prefixes using a pattern-based algorithm, which makes our dataset overcome the problems of uneven address distribution and low active rate. Secondly, we propose an efficient address generation algorithm DET, which builds a density space tree to learn high-density address regions of the seed addresses in linear time and improves the probing efficiency of active addresses. On the public hitlist and our hitlist, we compare our algorithm DET against state-of-the-art algorithms and find that DET increases the de-aliased active address ratio by 10%, and active address (including aliased addresses) ratio by 14%, by scanning 50 million addresses. Guanglei Song, Lin He 0004, Jiahai Yang 0001, Jieling Liu |
IWQoS | 1 |
| 2019 | Measurement and Analysis of Adult Websites in IPv6 NetworksabstractThe Internet is in the transition from IPv4 to IPv6. At present, researches on IPv6 networks mainly focus on architectural issues, such as routing, addressing, and security; there are few studies on the operational issues of IPv6 networks. Our preliminary observation shows that there are a large amount adult websites and traffic in IPv6 networks. Adult websites can damage health of teenagers and bring operational issues in IPv6 networks. This paper conducts a comprehensive measurement and analysis of the adult websites and traffic in IPv6 networks to help solve these operational issues. The data used in this paper is the raw packet traffic from CNGI-CERNET2 which is a pure IPv6 academic network in China. The duration of the data is from July 2017 to January 2018 and the total amount is 40+ terabytes. We detected about 3000 adult websites in the global IPv6 network. This paper analyzes these adult websites and traffic from the perspectives of websites, users and ISPs respectively. We find that adult websites are still in the developing stage in IPv6 networks and only 30% adult websites with full resources can be accessed in IPv6-only networks. But due to the IPv6-first policy in RFC 4038, adult traffic will continue to migrate to IPv6 networks from IPv4 networks. On the other hand, we find that CDN vendors promote the development of adult websites in IPv6 networks and many adult website owners use muti-domain policies to escape ISPs restricting. Our findings may help ISPs effectively understand adult websites and enhance the restriction of adult content in IPv6 networks. Shize Zhang, Hui Zhang 0052, Jiahai Yang 0001, Guanglei Song |
APNOMS | 4 |