Jinlei Lin

dblp:295/5881 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0003-1891-8322ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 9 · 1 first-author · 9 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 LMGeo6: A Comprehensive IPv6 Landmark Mining Methodology to Facilitate IPv6 Geolocation
abstract
IP geolocation is essential for various location-aware Internet applications. With the rapid development of the IPv6 protocol, IPv6 geolocation is becoming increasingly important. High-accuracy IPv6 geolocation relies heavily on high-quality landmarks. However, existing landmark mining methods mainly focus on IPv4 landmarks, and cannot mine IPv6 landmarks or are extremely inefficient, making IPv6 geolocation accuracy hard to improve. In this paper, we present a novel IPv6 geolocation framework, LMGeo6, proposing a comprehensive IPv6 landmark mining methodology for high-accuracy IPv6 geolocation. LMGeo6 migrates IPv4 landmark mining methods and designs specific IPv6 landmark mining methods, improving the number and coverage of IPv6 landmarks significantly. Based on these designs, we implement the LMGeo6 system and conduct real-world experiments to validate its efficacy. The experiment results show that LMGeo6 can mine 62,697 IPv6 landmarks, improving 15× over the total of other methods. LMGeo6 landmarks cover 165 countries, 5,805 cities, and 4,030 ASes, improving over 300% and reaching the same scale as IPv4 landmarks.
Jinlei Lin, Chenglong Li 0006, Hui Zhang 0052, Wentong Wang, Jiahai Yang 0001
NOMS1
2025 IPdb: A High-Precision IP Level Industry Categorization of Web Services
abstract
IP addresses with web services are crucial in the Internet ecosystem. Classifying these addresses by industry and organization offers valuable insights into the entities utilizing them, enabling more efficient network management and enhanced security. Previous work in website classification and Internet management struggles to offer an IP-level perspective of the industries of web services due to their limited industry categories or potential industry inconsistencies between IP address owners and AS owners. To this end, we present IPdb, an IP-level industry categorization dataset. To construct the dataset, we developed LLMIC, a Large Language Model-based Industry Categorization framework with a precision of nearly 96%. IPdb serves as a labeled database for future endeavors in developing IP-level industry classifiers, encompassing over 200 million IP addresses. Furthermore, our study indicates that 30% ~ 50% of organizations within critical infrastructure industries deploy web servers across multiple ASes. Our study also validates the problem of mismatched granularity in industry categorization at the AS level with 87.83% ASes in IPv4 and 72.96% ASes in IPv6 containing IP addresses from different industries.
Guanglei Song, Jiahai Yang 0001, Songyun Wu, Jinlei Lin, Lin He 0004, Chenglong Li 0006
WWW6
2024 6Vision: Image-Encoding-Based IPv6 Target Generation in Few-Seed Scenarios
abstract
Efficient global Internet scanning is crucial for network measurement and security analysis. While existing target generation algorithms verify remarkable performance in largescale detection, their efficiency notably diminishes in few-seed scenarios. This decline is primarily attributed to the intricate configuration rules and sampling bias of seed addresses. Moreover, instances where BGP prefixes have few seed addresses are widespread, constituting$63.65 \%$of occurrences. We introduce 6 Vision to tackle this challenge by introducing a novel approach to encoding IPv6 addresses into images, facilitating comprehensive analysis of intricate configuration rules. Through feature stitching, 6 Vision not only improves the learnable features but also amalgamates addresses associated with configuration patterns for enhanced learning. Moreover, it integrates an environmental feedback mechanism to refine model parameters based on identified active addresses, thereby alleviating the sampling bias inherent in seed addresses. As a result, 6Vision achieves high-accuracy detection even in few-seed scenarios. The HitRate of 6 Vision is improved by$181 \% \sim 2,490 \%$compared to existing algorithms, while the CoverNum is$1.18 \sim 11.20$times that of them. Additionally, 6Vision can function as a preliminary detection module for existing algorithms, yielding a conversion gain (CG) ranging from$242 \% \sim 2,081 \%$. Ultimately, we achieve a conversion rate (CR) of$28.97 \%$for few-seed scenarios. We enrich the IPv6 hitlist, not only enhancing current target generation algorithms for large-scale address detection in few-seed scenarios but also effectively supporting IPv6 network measurement and security analysis.
Wenjian Zhang, Guanglei Song, Lin He 0004, Jinlei Lin, Songyun Wu, Chenglong Li 0006, Jiahai Yang 0001
ICNP4
2024 ProbeGeo: A Comprehensive Landmark Mining Framework Based on Web Content
abstract
IP geolocation is essential for various location-aware Internet applications. High-quality IP geolocation landmarks play a decisive role in IP geolocation accuracy. However, the previous research works focusing on mining landmarks from the Internet are hampered by limited quantity, poor coverage, and insufficient landmark quality. In this paper, we present a new framework called ProbeGeo to mine high-quality landmarks automatically. We divide landmarks into common landmarks and probe landmarks, providing systematic mining methods based on online retrieval and web content. ProbeGeo expands traditional common landmarks by taking advantage of the exposure of multiple IoT (Internet of Things) devices on the Internet, mining them based on search engines and webpage contents. Common landmarks, consisting of multi-type devices, significantly improve landmark quantity and coverage. Furthermore, ProbeGeo establishes a methodology for acquiring new probe landmarks from Internet VPs (Vantage Points) webpages, extracting geographical locations from heterogeneous webpages and utilizing active probe functions. Probe landmarks enhance landmark quality and functions, bringing new geolocation frameworks and breaking through the geolocation accuracy bottleneck. We develop the ProbeGeo as a continuously running system and conduct real-world experiments to validate its efficacy. Our results show that ProbeGeo can detect 89,849 high-quality landmarks, including 6,874 probe landmarks and 82,975 common landmarks. ProbeGeo landmarks are about 10x more than existing work, distributed in 181 countries and 7,094 cities. ProbeGeo landmarks cover more than 8 types of devices, and more than 60% of them remain stable over one month. Moreover, the landmark accuracy of more than 58% of ProbeGeo landmarks is above street level, which has not been achieved in previous works. ProbeGeo can provide geolocation services with higher landmark accuracy and broader coverage by correlating a large scale of landmarks.
Jinlei Lin, Chenglong Li 0006, Guanglei Song, Linna Fan, Jiahai Yang 0001
IEEE/ACM Trans. Netw.1
2024 PMap: Reinforcement Learning-Based Internet-Wide Port Scanning
abstract
Internet-wide scanning is a commonly used research technique in various network surveys, such as measuring service deployment and security vulnerabilities. However, these network surveys are limited to the given port set, not comprehensively obtaining the real network landscape, and even misleading survey conclusions. In this work, we introduce PMap, a port scanning tool that efficiently discovers the most open ports from all 65K ports in the whole network. PMap uses the correlation of ports to build an open port correlation graph of each network, using a reinforcement learning framework to update the correlation graph based on feedback results and dynamically adjust the order of port scanning. Compared to current port scanning methods, PMap performs better on hit rate, coverage, and intrusiveness. Our experiments over real networks show that PMap can find 90% open ports by only scanning 125 ports (90%@125) to each address, which is 99.3% less than the state-of-the-art port scanning methods. It reduces the number of scanned ports to decrease the intrusive nature of port scanning. In addition, PMap is highly parallel and lightweight. It scans 500 networks in parallel, achieving a port recommendation rate of up to 18 million per second, consuming only 7GB of memory. PMap is the first effective practice for scanning open ports using reinforcement learning. It bridges the gap of existing scanning tools and effectively supports subsequent service discovery and security research.
Guanglei Song, Lin He 0004, Jinlei Lin, Linna Fan, Jiahai Yang 0001
IEEE/ACM Trans. Netw.4
2024 AddrMiner: A Fast, Efficient, and Comprehensive Global Active IPv6 Address Detection System
abstract
Fast Internet-wide scanning is essential for network situational awareness and asset evaluation. However, the vast IPv6 address space makes brute-force scanning infeasible. Despite advancements in state-of-the-art methods, they do not work in seedless regions and suffer low detection efficiency and speed in regions with known active IPv6 addresses (i.e., seed addresses). Moreover, the collected active address list (i.e., IPv6 hitlist) with low coverage cannot truly represent the active IPv6 address landscape of the Internet. This paper introduces AddrMiner, a fast, efficient, and comprehensive global active IPv6 address detection system. We design a systematic active IPv6 address detection strategy that divides the IPv6 space into two detection scenarios based on the presence or absence of seed addresses to discover active IPv6 addresses from scratch and from few to many. In the seedless regions, we present AddrMiner-N, leveraging a multi-level association policy to probe active addresses. It fills the gap of address detection in seedless regions and successfully discovers active addresses in 39,899 BGP prefixes without seed addresses, with a$1.03\times $higher hit rate,$30\sim 911\times $higher speed, and$2.7\times $broader coverage, compared to existing solutions. In the regions with seed addresses, our method AddrMiner-S dynamically generates target addresses using reinforcement learning. Compared to state-of-the-art methods, AddrMiner-S achieves an impressive 56.3% hit rate and a discovery speed of 839.0/s, which is$1.9\sim 2153\times $and$1.5\sim 755\times $of existing works, respectively. Finally, we deploy AddrMiner and discover 2.1B active IPv6 addresses, including 1.7B de-aliased active addresses and 0.4B aliased addresses, through continuous probing for three years.
Guanglei Song, Lin He 0004, Feiyu Zhu 0002, Jinlei Lin, Wenjian Zhang, Linna Fan, Chenglong Li 0006, Jiahai Yang 0001
IEEE/ACM Trans. Netw.4
2023 GraphIoT: Accurate IoT Identification based on Heterogeneous Graph
abstract
IoT devices deployed on campus and enterprise networks facilitate people's lives and work. However, these devices also bring serious network asset management and security management problems. IoT device identification is the premise to solve these problems. Although current IoT identification methods can identify devices with relatively high accuracy in ideal environments, it is difficult to accurately identify devices in real-world complex environments (e.g., campus networks, enterprise networks). Therefore, we propose to use exact features. To solve the problem of different dimensions of exact features, we creatively model the IoT identification problem as a heterogeneous graph representation learning problem and design a new representation learning algorithm. We are the first to propose an approach to accurately identify IoT devices in real-world complex environments and solve this problem through heterogeneous graphs. The evaluation shows that GraphIoT's macro F1 is on average 13.58% and 12.77% higher than the other methods on two public datasets.
Linna Fan, Lin He 0004, Xiaoqing Sun, Enhuan Dong, Jiahai Yang 0001, Jinlei Lin, Guanglei Song
IWQoS7
2022 Both Efficient and Accurate: A Large-scale One-way Delay Measurement Scheme
abstract
One-way delay (OWD) is one of the essential network performance metrics. In large-scale resilient overlay networks (RONs), OWD measurements can be used for shortest path selection and troubleshooting. However, OWD measurements remain difficult because of the need for precise time synchronization. Especially in large-scale networks, clock synchronization of all nodes has always been a considerable challenge. Therefore, in many cases, people use half of the round-trip time (RTT/2) as a rough substitute for the OWD. This paper presents an efficient and easy-to-deploy scheme for large-scale OWD measurements with the algorithm ClockConverger at its core. The scheme consists of three steps: Firstly, we perform low-precision time synchronization for all the measured nodes relying on network time protocol daemons (ntpd); Then, we use the open-source tool OWPing to perform OWD measurements; Finally, we correct the errors of the measured OWDs with our proposed ClockConverger. The theory and experiments show that our scheme's accuracy is significantly better than RTT/2. Meanwhile, the complexity of ClockConverger is$O(n^{2})$, which is much lower than the exponential complexity of the existing Maximum-Entropy algorithm.
Yaozhong Liu, Jiahai Yang 0001, Long Pan, Lin He 0004, Jinlei Lin, Guanglei Song, Chenglong Li 0006
GLOBECOM6
2022 What Causes Delay Asymmetry: A Large-scale One-way Delay Measurement and Empirical Study
abstract
In global communications, severe one-way delay (OWD) asymmetry often occurs. Due to the difficulties of OWD measurement (need to control both ends and synchronize their clocks), now RTT/2 is commonly used to estimate OWD. However, OWD asymmetry can lead to large errors in the halving RTT method, which in turn affects the end-to-end quality of service (QoS) guarantees. In this paper, we investigate OWD asymmetry through large-scale OWD measurements on a global scale. The measurements show that more than 11% of network paths have OWDs with a relative difference of more than 10% compared to RTT/2. By analyzing the measurement results in depth, we try to explain why the delay asymmetry occurs. We find that 67% is caused by hop inflation or a significant increase in propagation distance, and 33% is caused by variable queuing delays. We also find AS-level paths between node pairs with significant delay asymmetry are much more likely (~ 10 ×) to violate the well-known valley-free rule.
Yaozhong Liu, Jiahai Yang 0001, Long Pan, Lin He 0004, Jinlei Lin, Guanglei Song, Chenglong Li 0006
GLOBECOM6
2022 EvoIoT: An evolutionary IoT and non-IoT classification model in open environments
Linna Fan, Lin He 0004, Enhuan Dong, Jiahai Yang 0001, Chenglong Li 0006, Jinlei Lin
Comput. Networks6
2022 DET: Enabling Efficient Probing of IPv6 Active Addresses
abstract
Fast IPv4 scanning significantly improves network measurement and security research. Nevertheless, it is infeasible to perform brute-force scanning of the IPv6 address space. Alternatively, one can find active IPv6 addresses through scanning the candidate addresses generated by state-of-the-art algorithms. However, the probing efficiency of such algorithms is often very low. In this paper, our objective is to improve the probing efficiency of IPv6 addresses. We first perform a longitudinal active measurement study and build a high-quality dataset, hitlist, including more than 1.95B IPv6 addresses distributed in 58.2K BGP prefixes and collected over 17 months period. Different from the previous works, we probe the announced BGP prefixes using a pattern-based algorithm. This results in a dataset without uneven address distribution and low active rates. Further, we propose an efficient address generation algorithm, DET, which builds a density space tree to learn high-density address regions of the seed addresses with linear time complexity and improves the active addresses’ probing efficiency. We then compare our algorithm DET against state-of-the-art algorithms on the public hitlist and our hitlist by scanning 50M addresses. Our analysis shows that DET increases the de-aliased active address ratio and active address (including aliased addresses) ratio by 10%, and 14%, respectively. Furthermore, we develop a fingerprint-based method to detect aliased prefixes. The proposed method for the first time directly verifies whether the prefix is aliased or not. Our method finds that 10.64% of the public aliased prefixes are false positive.
Guanglei Song, Jiahai Yang 0001, Lin He 0004, Jinlei Lin, Long Pan, Chenxin Duan, Xiaowen Quan
IEEE/ACM Trans. Netw.5
2021 A novel workload scheduling framework for intrusion detection system in NFV scenario
Jia Li 0033, Jiahai Yang 0001, Jinlei Lin
Comput. Secur.4