VLDB 2026 Research / reviewers in the wild / expert
Chenglong Li 0006
dblp:83/7820-6
· DBLP profile ↗
24ranked-venue papers
0as first author
24since 2021 · last 2026
0000-0003-4300-678XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 11 since 2021Security and privacy · 6 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FASST-LLM: Implementing Service Scanning Tools Automatically with LLM
Lichao Qin, Yirui Luo, Chenglong Li 0006, Jiahai Yang 0001 |
INFOCOM | 3 |
| 2026 | F2D: Detection of resolver DNS hijacking based on filtration funnel strategyabstractAbstract In recent years, DNS hijacking represents a significant security threat to the infrastructure of the Domain Name System (DNS). A prevalent form of DNS hijacking involves exploiting open resolvers to manipulate DNS records. Such attacks undermine the availability and confidentiality of network services, posing serious risks to legitimate users. Current DNS hijacking detection methods tend to focus on specific domains, leveraging the unique characteristics of domain-specific hijacking to identify attacks. Consequently, these methods are often limited in applicability and may lack accuracy when dealing with diverse hijacking scenarios. Additionally, many existing approaches face challenges related to efficiency, making them less effective for long-term monitoring of hijacking activities. To address these challenges, this paper introduces an efficient detection method F2D tailored for general DNS hijacking. First, F2D uses an accurate and efficient filtration funnel strategy for targeted resolver hijacking detection. Second, two optimized detection algorithms are proposed for comprehensive filtration. Third, the method includes an efficient mechanism for identifying CDN domains, enabling the filtration of a large number of content replication servers and enhancing overall detection efficiency. During the validation phase, we monitor around 36k domains and around 600 resolvers over a one-month period. The effectiveness of our method is validated using manually labeled sample data. Experimental results demonstrate that our method can improve the F1 performance by 10% with the same false alert level, and time efficiency by 39% compared to the state-of-the-arts. Furthermore, we conduct an in-depth analysis of the captured hijacking incidents and deduce the motivation of the hijacking. Cong Dong, Haoran Jiao, Jiahai Yang 0001, Chenglong Li 0006, Xia Yin 0001 |
Cybersecur. | 5 |
| 2026 | HINHJ: Hierarchical Attention-Based Heterogeneous Graph Neural Network for DNS Hijacking DetectionabstractThe Domain Name System (DNS) is a critical internet infrastructure that translates human-readable domain names into machine-routable IP addresses. However, DNS is inherently vulnerable to manipulation, with hijacking attacks growing in both frequency and sophistication. Existing detection methods primarily rely on traffic analysis at specific network points. However, they suffer from limited coverage and low accuracy in complex environments, such as when CDN is employed. While recent approaches employ graph-based techniques, they still suffer from detection inaccuracy issues due to their failure to account for the complex interdependencies among multiple types of nodes. To address these limitations, we propose a novel heterogeneous graph-based detection framework. Based on the collected DNS records from distributed scanners, our method extracts activity and security features and constructs a heterogeneous graph to capture resolution patterns and cross-entity relationships. We further design a time-decay graph neural network TNHAN that enhances traditional Heterogeneous Graph Attention Networks (HAN) by dynamically weighting recent records. This network improves adaptability to legitimate DNS changes. For evaluation, we conduct experiments on real-world resolvers and domain datasets. Experiment results demonstrate the effectiveness of our method. Our method can achieve an F1-score of 0.96, outperforming the best baseline by 0.057 on average, and up to 0.113 under low label proportion. Moreover, we conduct several case studies on detected incidents, including cases related to geopolitical conflicts, censorship-related hijacking, and manipulation by malicious resolvers. These cases demonstrate the method’s effectiveness in identifying diverse hijacking behaviors in practice. Haoran Jiao, Cong Dong, Chenglong Li 0006, Jiahai Yang 0001, Leyao Nie, Changzhi Zhao, Xia Yin 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | ZVDetector: State-Guided Vulnerability Detection System for Zigbee DevicesabstractNowadays, Zigbee devices are widely used in smart home, smart agriculture and other industries. However, there are many vulnerabilities in Zigbee devices that could compromise their normal functionality. Existing research either analyzes firmware or fuzzes devices through Zigbee networks to discover potential vulnerabilities. However, they overlook the impact of device state and protocol state on firmware or explore only a limited state space. Thus, they fail to identify many vulnerabilities caused by hidden states within each of the two states, especially vulnerabilities triggered by the combination of these two states. In this paper, we design a state-guided fuzzing system, named ZVDetector, aimed at uncovering firmware vulnerabilities caused by hidden and combined states. Specifically, we design two state-aware modules that explore richer unknown protocol state transitions based on message relationships and gain a more complete understanding of the intrinsic device state attributes. We develop a fuzzing algorithm that incorporates message semantics awareness and correlation state analysis. By integrating the perceived state information, it can explore the combined state space more efficiently. We validate the performance of ZVDetector on 10 Zigbee devices and find 25 vulnerabilities (19 zero-day). Our experiments also demonstrate the ability to explore more device state attributes and discover more message relationships related to unknown protocol states. Chenglong Li 0006, Jiahai Yang 0001 |
CCS | 2 |
| 2025 | Poster: TopoHunter: Enabling Efficient and High-Coverage Active IPv6 Topology DiscoveryabstractWe introduce TopoHunter, an efficient IPv6 Internet topology discovery system. The central concept of TopoHunter is to allocate more probing resources to target prefix spaces that yield greater topological benefits, as well as to their surrounding areas. To achieve this, we design a feedback-based target generation module comprised of a Target Prefix Probing Value Forest that maintains the estimated probing values of hierarchical target prefix spaces. Our system has successfully discovered the most extensive and complete IPv6 topology map to date, comprising over 144 million router interfaces and 251 million edges, covering 72.83% of autonomous systems and 43.36% of routing prefixes announced by the BGP system. Lin He 0004, Hongwei Li 0021, Guanglei Song, Wentong Wang, Daguo Cheng, Enhuan Dong, Chenglong Li 0006, Hui Zhang 0141, Jinlong E, Ying Liu 0024, Jiahai Yang 0001 |
IMC | 8 |
| 2025 | 6RV: Incremental Learning-Based Continuous Identification of IPv6 Router VendorsabstractThe growth of IPv6 networks has led to an expanding number of network routers, but there is not enough research on vendors of these devices. Existing router vendor identification algorithms are based on IPv4, and these static analysis algorithms cannot adapt to dynamically changing IPv6 networks. In this paper, we develop 6RV, a framework for IPv6 router vendors continuous identification based on incremental learning. First, we count the addresses of newly discovered router interfaces every month and obtain router device fingerprints through active probing, and then analyze these data fingerprints through an incremental learning approach to identify the router vendors of new nodes. We validate our identification framework on the ITDK dataset over a period of 8 months, obtaining more than 500K router vendor labels with 94% correctness. Finally, we also analyze the IPv6 router vendor dataset from different perspectives and draw some interesting conclusions. Shenao Li, Jiahai Yang 0001, Enhuan Dong, Chenglong Li 0006, Lin He 0004, Hui Zhang 0052, Guanglei Song |
NOMS | 5 |
| 2025 | LMGeo6: A Comprehensive IPv6 Landmark Mining Methodology to Facilitate IPv6 GeolocationabstractIP geolocation is essential for various location-aware Internet applications. With the rapid development of the IPv6 protocol, IPv6 geolocation is becoming increasingly important. High-accuracy IPv6 geolocation relies heavily on high-quality landmarks. However, existing landmark mining methods mainly focus on IPv4 landmarks, and cannot mine IPv6 landmarks or are extremely inefficient, making IPv6 geolocation accuracy hard to improve. In this paper, we present a novel IPv6 geolocation framework, LMGeo6, proposing a comprehensive IPv6 landmark mining methodology for high-accuracy IPv6 geolocation. LMGeo6 migrates IPv4 landmark mining methods and designs specific IPv6 landmark mining methods, improving the number and coverage of IPv6 landmarks significantly. Based on these designs, we implement the LMGeo6 system and conduct real-world experiments to validate its efficacy. The experiment results show that LMGeo6 can mine 62,697 IPv6 landmarks, improving 15× over the total of other methods. LMGeo6 landmarks cover 165 countries, 5,805 cities, and 4,030 ASes, improving over 300% and reaching the same scale as IPv4 landmarks. Jinlei Lin, Chenglong Li 0006, Hui Zhang 0052, Wentong Wang, Jiahai Yang 0001 |
NOMS | 2 |
| 2025 | Post-Standardization Analysis of DoQ: Deployment, Certificates Ecosystem and ImplementationabstractTo address the security issues caused by traditional plaintext DNS transmission, encrypted protocols were introduced to protect DNS traffic. DNS over QUIC (DoQ) is the most recent DNS encryption protocol standardized in 2022. While earlier protocols like DoT and DoH have been extensively studied, research on DoQ remains limited, focusing primarily on basic deployment and performance. There is a lack of comprehensive research on the DoQ ecosystem after its standardization, and the compliance and security of its deployment remain unclear. This paper presents the first in-depth measurement of DoQ deployment across IPv4, IPv6, and authoritative servers. Our findings offer an early view of the DoQ ecosystem, covering its deployment, certificate ecology, and practical implementations. Overall, the progress of DoQ standardization is satisfactory. Since standardization, DoQ adoption has tripled, and its certificate ecosystem shows a promising trend, with fewer than 10% of certificates being invalid. However, potential security concerns persist. First, the centralization issue in DoQ is more pronounced compared to DoH and DoT. Second, about 30% of DoQ authority servers support recursive parsing, facing the risk of cache poisoning or DDoS attacks. In addition, 2% of DoQ deployments fail to meet RFC requirements, potentially enabling amplification attacks. Therefore, we highlight the need for stricter compliance with standards in future DoQ implementations to enhance security and reliability. Chenglong Li 0006, Wenchong Dong, Cong Dong, Jiahai Yang 0001, Hui Zhang 0052 |
NOMS | 2 |
| 2025 | IPdb: A High-Precision IP Level Industry Categorization of Web ServicesabstractIP addresses with web services are crucial in the Internet ecosystem. Classifying these addresses by industry and organization offers valuable insights into the entities utilizing them, enabling more efficient network management and enhanced security. Previous work in website classification and Internet management struggles to offer an IP-level perspective of the industries of web services due to their limited industry categories or potential industry inconsistencies between IP address owners and AS owners. To this end, we present IPdb, an IP-level industry categorization dataset. To construct the dataset, we developed LLMIC, a Large Language Model-based Industry Categorization framework with a precision of nearly 96%. IPdb serves as a labeled database for future endeavors in developing IP-level industry classifiers, encompassing over 200 million IP addresses. Furthermore, our study indicates that 30% ~ 50% of organizations within critical infrastructure industries deploy web servers across multiple ASes. Our study also validates the problem of mismatched granularity in industry categorization at the AS level with 87.83% ASes in IPv4 and 72.96% ASes in IPv6 containing IP addresses from different industries. Guanglei Song, Jiahai Yang 0001, Songyun Wu, Jinlei Lin, Lin He 0004, Chenglong Li 0006 |
WWW | 8 |
| 2025 | E-DoH: elegantly detecting the depths of open DoH service on the internetabstractAbstract In recent years, DoE methods have been regarded as a novel trend within the realm of the DNS ecosystem. Measuring these DoE services in the wild can promote improvements in DoE methods and facilitate their widespread adoption. A primary requirement for measuring DoE methods is the discovery of these services. The discovery is relatively straightforward for DoT and DoQ, but complex for DoH since it shares port 443 with web services as suggested in RFC 8484. Although previous works primarily analyze the surface of the DoH service, they (1) result in long detection time and large traffic volume by adopting an enumeration strategy to discover the DoH service; (2) lack an in-depth analysis of the status of upper-layer DNS services. In this paper, we propose the E-DoH method for elegant, efficient, and in-depth DoH service measurement. First, we propose a measurement mechanism to enable a single DoH connection to accomplish multiple tasks including service discovery, correctness validation, and dependency construction with minimal backend configuration. Second, we propose a dynamic protocol negotiation strategy to enhance probing efficiency while significantly reducing the required traffic volume. Based on the above optimization methods, we conducted an exploration of the IPv4 space and performed an in-depth analysis of DoH based on the collected information. Through experiments, our approach demonstrates a remarkable 80% improvement in time efficiency and only requires 4–20% traffic volume to complete the detection task. In wild detection, our approach discovered 46k DoH services, which nearly doubles the number discovered by the state-of-the-art. This indicates the growing trend of DoH services. Based on the collected information, we present several intriguing conclusions about the current DoH service ecosystem. Cong Dong, Jiahai Yang 0001, Haoran Jiao, Chenglong Li 0006, Xia Yin 0001 |
Cybersecur. | 5 |
| 2024 | 6Vision: Image-Encoding-Based IPv6 Target Generation in Few-Seed ScenariosabstractEfficient global Internet scanning is crucial for network measurement and security analysis. While existing target generation algorithms verify remarkable performance in largescale detection, their efficiency notably diminishes in few-seed scenarios. This decline is primarily attributed to the intricate configuration rules and sampling bias of seed addresses. Moreover, instances where BGP prefixes have few seed addresses are widespread, constituting$63.65 \%$of occurrences. We introduce 6 Vision to tackle this challenge by introducing a novel approach to encoding IPv6 addresses into images, facilitating comprehensive analysis of intricate configuration rules. Through feature stitching, 6 Vision not only improves the learnable features but also amalgamates addresses associated with configuration patterns for enhanced learning. Moreover, it integrates an environmental feedback mechanism to refine model parameters based on identified active addresses, thereby alleviating the sampling bias inherent in seed addresses. As a result, 6Vision achieves high-accuracy detection even in few-seed scenarios. The HitRate of 6 Vision is improved by$181 \% \sim 2,490 \%$compared to existing algorithms, while the CoverNum is$1.18 \sim 11.20$times that of them. Additionally, 6Vision can function as a preliminary detection module for existing algorithms, yielding a conversion gain (CG) ranging from$242 \% \sim 2,081 \%$. Ultimately, we achieve a conversion rate (CR) of$28.97 \%$for few-seed scenarios. We enrich the IPv6 hitlist, not only enhancing current target generation algorithms for large-scale address detection in few-seed scenarios but also effectively supporting IPv6 network measurement and security analysis. Wenjian Zhang, Guanglei Song, Lin He 0004, Jinlei Lin, Songyun Wu, Chenglong Li 0006, Jiahai Yang 0001 |
ICNP | 7 |
| 2024 | CP-IoT: A Cross-Platform Monitoring System for Smart Home
Chenglong Li 0006, Jiahai Yang 0001, Linna Fan, Chenxin Duan |
NDSS | 2 |
| 2024 | TrafficSiam: More Realistic Few-shot Website Fingerprinting Attack with Contrastive LearningabstractWebsite fingerprinting (WF) attacks pose a serious threat to users’ online privacy, even when using privacy-enhancing tools like Tor. Previous attacks mostly rely on supervised learning and only a few studies have explored the few-shot setting which is more realistic. In this paper, we propose TrafficSiam, a novel WF attack based on few-shot learning with self-supervised learning, which enables our model to be pretrained with unlabeled tor traffic and transferred to new tasks using a small number of labeled samples which is more realistic and yield stronger generalization ability. We conduct a series of experiments, using only a small amount of labeled samples, and find that our model achieves 92.32% accuracy in the closed-world setting, compared to the highest accuracy 88.73%, using previous methods. Furthermore, our model also outperforms previous attacks in the open-world setting. Shangdong Wang, Chenglong Li 0006, Jiahai Yang 0001, Hui Zhang 0052 |
NOMS | 3 |
| 2024 | IoTa: Fine-Grained Traffic Monitoring for IoT Devices via Fully Packet-Level ModelsabstractWith Internet-of-Things (IoT) devices gaining popularity, dedicated monitoring systems which accurately detect intrusion traffic for them are in high demand. Existing methods mainly use statistical spatial-temporal traffic features and machine learning models. Their practicality has been limited due to the lack of detection ability for stealthy and tricky attacks, diagnostic utility and long-term performance. To address these problems and motivated by the simplicity of mini IoT devices, we propose to construct fully packet-level models to profile traffic patterns for IoT devices by constructing automaton for short flow and long flow, where the length and direction of each packet are the representative features. We apply these fine-grained models to design and develop a traffic monitoring system, namelyIoTa, to detect intrusion traffic for IoT devices.IoTamatches the ongoing traffic with patterns extracted from normal traffic traces. With visible and interactive traffic profiles,IoTacan generate interpretable alerts and is available for long-term use under reasonable human efforts. Evaluations on dozens of common IoT devices show thatIoTacan achieve excellent detection accuracy (nearly perfect recalls and always over 0.999 precisions) for various intrusion traffic covering the complete kill chains. Incorrect detection results can be compensated for by error recovery mechanisms and the understandable alert context can be used by the operator to enhance the system. The diagnostic utility and little alert weariness are recognized by the experienced operators. Chenxin Duan, Sainan Li, Guanglei Song, Chenglong Li 0006, Jiahai Yang 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2024 | ProbeGeo: A Comprehensive Landmark Mining Framework Based on Web ContentabstractIP geolocation is essential for various location-aware Internet applications. High-quality IP geolocation landmarks play a decisive role in IP geolocation accuracy. However, the previous research works focusing on mining landmarks from the Internet are hampered by limited quantity, poor coverage, and insufficient landmark quality. In this paper, we present a new framework called ProbeGeo to mine high-quality landmarks automatically. We divide landmarks into common landmarks and probe landmarks, providing systematic mining methods based on online retrieval and web content. ProbeGeo expands traditional common landmarks by taking advantage of the exposure of multiple IoT (Internet of Things) devices on the Internet, mining them based on search engines and webpage contents. Common landmarks, consisting of multi-type devices, significantly improve landmark quantity and coverage. Furthermore, ProbeGeo establishes a methodology for acquiring new probe landmarks from Internet VPs (Vantage Points) webpages, extracting geographical locations from heterogeneous webpages and utilizing active probe functions. Probe landmarks enhance landmark quality and functions, bringing new geolocation frameworks and breaking through the geolocation accuracy bottleneck. We develop the ProbeGeo as a continuously running system and conduct real-world experiments to validate its efficacy. Our results show that ProbeGeo can detect 89,849 high-quality landmarks, including 6,874 probe landmarks and 82,975 common landmarks. ProbeGeo landmarks are about 10x more than existing work, distributed in 181 countries and 7,094 cities. ProbeGeo landmarks cover more than 8 types of devices, and more than 60% of them remain stable over one month. Moreover, the landmark accuracy of more than 58% of ProbeGeo landmarks is above street level, which has not been achieved in previous works. ProbeGeo can provide geolocation services with higher landmark accuracy and broader coverage by correlating a large scale of landmarks. Jinlei Lin, Chenglong Li 0006, Guanglei Song, Linna Fan, Jiahai Yang 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2024 | AddrMiner: A Fast, Efficient, and Comprehensive Global Active IPv6 Address Detection SystemabstractFast Internet-wide scanning is essential for network situational awareness and asset evaluation. However, the vast IPv6 address space makes brute-force scanning infeasible. Despite advancements in state-of-the-art methods, they do not work in seedless regions and suffer low detection efficiency and speed in regions with known active IPv6 addresses (i.e., seed addresses). Moreover, the collected active address list (i.e., IPv6 hitlist) with low coverage cannot truly represent the active IPv6 address landscape of the Internet. This paper introduces AddrMiner, a fast, efficient, and comprehensive global active IPv6 address detection system. We design a systematic active IPv6 address detection strategy that divides the IPv6 space into two detection scenarios based on the presence or absence of seed addresses to discover active IPv6 addresses from scratch and from few to many. In the seedless regions, we present AddrMiner-N, leveraging a multi-level association policy to probe active addresses. It fills the gap of address detection in seedless regions and successfully discovers active addresses in 39,899 BGP prefixes without seed addresses, with a$1.03\times $higher hit rate,$30\sim 911\times $higher speed, and$2.7\times $broader coverage, compared to existing solutions. In the regions with seed addresses, our method AddrMiner-S dynamically generates target addresses using reinforcement learning. Compared to state-of-the-art methods, AddrMiner-S achieves an impressive 56.3% hit rate and a discovery speed of 839.0/s, which is$1.9\sim 2153\times $and$1.5\sim 755\times $of existing works, respectively. Finally, we deploy AddrMiner and discover 2.1B active IPv6 addresses, including 1.7B de-aliased active addresses and 0.4B aliased addresses, through continuous probing for three years. Guanglei Song, Lin He 0004, Feiyu Zhu 0002, Jinlei Lin, Wenjian Zhang, Linna Fan, Chenglong Li 0006, Jiahai Yang 0001 |
IEEE/ACM Trans. Netw. | 7 |
| 2023 | Anomaly Detection in Heterogeneous Time Series Data for Server-Monitoring TasksabstractWhen conducting anomaly detection on server monitoring data, it is important to consider the heterogeneity of the data, which is characterized by the diverse and irregular nature of events. The event values can vary widely, encompassing both continuous and discrete values, and there may be a multitude of randomly occurring events. However, many commonly used anomaly detection methods tend to overlook or discard this heterogeneous data, resulting in a significant loss of valuable information. As such, we propose a novel method, called Heterogeneous Time Series Anomaly Detection (HTSAD), to overcome this difficulty. The approach introduces event gates in the Long Short-Term Memory (LSTM) model while using unsupervised learning to overcome the challenges mentioned above. The results of our experiments on real-world datasets show that HTSAD could achieve an f-score of 0.958, which demonstrates the effectiveness of our approach in detecting anomalies in heterogeneous time series data. Rui Yu 0003, Jiahai Yang 0001, Minghui Jin, Chenglong Li 0006, Enhuan Dong, Shutao Xia |
ISCC | 8 |
| 2023 | Which Doors Are Open: Reinforcement Learning-based Internet-wide Port ScanningabstractInternet-wide scanning is a commonly used research technique in various network surveys, such as measuring service deployment and security vulnerabilities. However, these network surveys are limited to the given port set, not comprehensively obtaining the real network landscape, and even misleading survey conclusions. In this work, we introduce PMap, a port scanning tool that efficiently discovers the majority of open ports from all 65K ports in the whole network. PMap uses the correlation of ports to build an open port correlation graph of each network, using a reinforcement learning framework to update the correlation graph based on feedback results and dynamically adjust the order of port scanning. Compared to current port scanning methods, PMap achieves better performance on hit rate, coverage, and intrusiveness. Our experiments over real-world networks show that PMap can find 90% open ports by only scanning 125 ports (90% @125) to each active address with 136× less than the state-of-the-art port probing methods. PMap reduces the number of scanned ports to decrease the intrusive nature of port scanning. PMap is the first effective practice for scanning open ports using reinforcement learning. It bridges the gap of existing scanning tools and effectively supports subsequent service discovery and security research. Guanglei Song, Lin He 0004, Tianyun Zhao, Yirui Luo, Yichao Wu, Linna Fan, Chenglong Li 0006, Jiahai Yang 0001 |
IWQoS | 7 |
| 2022 | PerfTrace: A New Multi-metric Network Performance Monitoring ToolabstractWe present PerfTrace, an end-to-end tool for efficient, real-time, and multi-metric network performance monitoring. PerfTrace provides a high integration of different existing measurement functions, supporting the measurement of essential metrics such as latency, jitter, packet loss, and available bandwidth. More importantly, innovative schemes and algorithms are proposed to address the weaknesses of existing tools.After conducting comprehensive evaluations, we find that (i) PerfTrace measures one-way and two-way latency, jitter, and packet loss ∼9.4× faster and ∼3.6× more data-efficiently; (ii) PerfTrace measures available bandwidth in our testbed with minimal mean relative error (5.22%), outperforming all the tools compared (ranging from 8.17% to 37.24%). Meanwhile, PerfTrace consumes a more constant percentage of bandwidth resources than other tools when monitoring available bandwidth. PerfTrace’s data overhead is always only about 1/600 of the total bandwidth for a measurement frequency once per minute. Yaozhong Liu, Long Pan, Chenglong Li 0006, Lin He 0004, Yirui Luo, Guanglei Song, Jiahai Yang 0001 |
CNSM | 3 |
| 2022 | Towards a Behavioral and Privacy Analysis of ECS for IPv6 DNS ResolversabstractThe Domain Name System (DNS) is critical to Internet communications. EDNS Client Subnet (ECS), a DNS extension, allows recursive resolvers to include client subnet information in DNS queries to improve CDN end-user mapping, extending the visibility of client information to a broader range. Major content delivery network (CDN) vendors, content providers (CP), and public DNS service providers (PDNS) are accelerating their IPv6 infrastructure development. With the increasing deployment of IPv6-enabled services and DNS being the most foundational system of the Internet, it becomes important to analyze the behavioral and privacy status of IPv6 resolvers. However, there is a lack of research on ECS for IPv6 DNS resolvers.In this paper, we study the ECS deployment and compliance status of IPv6 resolvers. Our measurement shows that 11.12% IPv6 open resolvers implement ECS. We discuss abnormal noncompliant scenarios that exist in both IPv6 and IPv4 that raise privacy and performance issues. Additionally, we measured if the sacrifice of clients’ privacy can enhance IPv6 CDN performance. We find that in some cases ECS helps end-user mapping but with an unnecessary privacy loss. And even worse, the exposure of client address information can sometimes backfire, which deserves attention from both Internet users and PDNSes. Leyao Nie, Lin He 0004, Guanglei Song, Chenglong Li 0006, Jiahai Yang 0001 |
CNSM | 5 |
| 2022 | Both Efficient and Accurate: A Large-scale One-way Delay Measurement SchemeabstractOne-way delay (OWD) is one of the essential network performance metrics. In large-scale resilient overlay networks (RONs), OWD measurements can be used for shortest path selection and troubleshooting. However, OWD measurements remain difficult because of the need for precise time synchronization. Especially in large-scale networks, clock synchronization of all nodes has always been a considerable challenge. Therefore, in many cases, people use half of the round-trip time (RTT/2) as a rough substitute for the OWD. This paper presents an efficient and easy-to-deploy scheme for large-scale OWD measurements with the algorithm ClockConverger at its core. The scheme consists of three steps: Firstly, we perform low-precision time synchronization for all the measured nodes relying on network time protocol daemons (ntpd); Then, we use the open-source tool OWPing to perform OWD measurements; Finally, we correct the errors of the measured OWDs with our proposed ClockConverger. The theory and experiments show that our scheme's accuracy is significantly better than RTT/2. Meanwhile, the complexity of ClockConverger is$O(n^{2})$, which is much lower than the exponential complexity of the existing Maximum-Entropy algorithm. Yaozhong Liu, Jiahai Yang 0001, Long Pan, Lin He 0004, Jinlei Lin, Guanglei Song, Chenglong Li 0006 |
GLOBECOM | 8 |
| 2022 | What Causes Delay Asymmetry: A Large-scale One-way Delay Measurement and Empirical StudyabstractIn global communications, severe one-way delay (OWD) asymmetry often occurs. Due to the difficulties of OWD measurement (need to control both ends and synchronize their clocks), now RTT/2 is commonly used to estimate OWD. However, OWD asymmetry can lead to large errors in the halving RTT method, which in turn affects the end-to-end quality of service (QoS) guarantees. In this paper, we investigate OWD asymmetry through large-scale OWD measurements on a global scale. The measurements show that more than 11% of network paths have OWDs with a relative difference of more than 10% compared to RTT/2. By analyzing the measurement results in depth, we try to explain why the delay asymmetry occurs. We find that 67% is caused by hop inflation or a significant increase in propagation distance, and 33% is caused by variable queuing delays. We also find AS-level paths between node pairs with significant delay asymmetry are much more likely (~ 10 ×) to violate the well-known valley-free rule. Yaozhong Liu, Jiahai Yang 0001, Long Pan, Lin He 0004, Jinlei Lin, Guanglei Song, Chenglong Li 0006 |
GLOBECOM | 8 |
| 2022 | WebIoT: Classifying Internet of Things Devices at Internet Scale through Web CharacteristicsabstractThe number of Internet of Things (IoT) devices connected to the Internet has been growing rapidly. Such a large number of IoT devices bring significant challenges to device man-agement and cyberspace security. The discovery and classification of IoT devices are the prerequisites for monitoring and protecting them. However, existing Internet-scale IoT device classification methods mainly rely on textual analysis of the device response data, whose performance can be affected by the complexity or the multilingualism of the response texts. In this paper, we propose WebIoT, which mainly utilizes the image characteristics of the IoT devices' web interfaces to classify them for the first time. We leverage the observation that many IoT devices have web interfaces for device configuration and device status display, whose visual presentations contain abundant characteristics for device classification. Experiment results show that our method achieves 95.4% precision and 91.5% recall, which significantly outperforms other text analysis-based methods. Yichao Wu, Chenglong Li 0006, Jiahai Yang 0001, Ang Xia, Yong Jiang 0001, Liuli Wu |
ISCC | 2 |
| 2022 | EvoIoT: An evolutionary IoT and non-IoT classification model in open environments
Linna Fan, Lin He 0004, Enhuan Dong, Jiahai Yang 0001, Chenglong Li 0006, Jinlei Lin |
Comput. Networks | 5 |