VLDB 2026 Research / reviewers in the wild / expert
Xionglve Li
dblp:262/6789 · also Xiong-lve Li
· DBLP profile ↗
13ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-6650-7259ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BIFM: an effective similar payload attribution approach for cybercriminal detection using bitmap index table and fuzzy matchingabstractAbstract The payload attribution system has been proposed to analyze network traffic and assist investigators in identifying flows containing specific excerpts to locate criminals and potential victims. However, various attacks or data leakage behaviors can obscure and scatter the crucial portion of flow payloads to evade detection. Although existing payload attribution techniques strive to enhance the data reduction ratio and reduce false positive rates, research on similar payload querying is notably lacking. In this study, we introduce bitmap index table fuzzy matching (BIFM), a method for digesting network traffic to query and trace variants of malicious traffic. Unlike deterministic bitmap-index PAS that require deterministic bit co-occurrence/alignment between the query excerpt and the stored flow bitmap, an assumption violated when payloads are split or jumbled, BIFM overcomes this limitation via progressive relaxation with fuzzy matching and verification. Leveraging the bitmap index table and fuzzy matching, BIFM efficiently identifies flows containing excerpts or their variants (excerpts that change their appearance by splitting or jumbling) by relaxing the matching conditions for candidate malicious flows. To enhance BIFM’s accuracy, we also propose no-shingling and packet caching mechanisms. We extensively evaluate BIFM’s performance using a dataset constructed from real campus network IP-trace data. Our results demonstrate that BIFM outperforms existing state-of-the-art solutions, achieving an accuracy improvement of $\sim $10% without significantly increasing processing time. Changsheng Hou, Ling Hu 0001, Xionglve Li, Bingnan Hou, Zhiping Cai |
Comput. J. | 4 |
| 2026 | Efficient router fingerprinting in IPv6 networksabstractAbstract The pervasive interconnection of heterogeneous routing devices forms the fundamental infrastructure of modern Internet communication, making accurate router vendor identification a critical capability for multiple domains including network topology mapping, intelligent traffic engineering, and proactive cybersecurity defense. While Internet Protocol version 6 (IPv6) has achieved widespread global deployment as the next-generation Internet protocol, the opaque nature of its addressing mechanisms and protocol behaviors has created significant challenges in router attribute detection across IPv6 networks, leaving a crucial gap in network visibility and security analytics. To address this pressing challenge, we present IPv6 Router FingerPrinting (6RFP), an innovative lightweight fingerprinting methodology that establishes a new paradigm for IPv6 router vendor identification by systematically combining two complementary analytical dimensions: (i) comprehensive EUI-64 interface identifier analysis that captures vendor-specific hardware encoding patterns embedded in IPv6 addresses, and (ii) sophisticated IPv6 Identification Field characteristic profiling that reveals distinctive vendor implementations. Through extensive evaluation across diverse network environments, 6RFP demonstrates highly effective detection capabilities, achieving 85.79% accuracy—representing a remarkable 86.01% improvement over current state-of-the-art techniques—while maintaining minimal computational overhead suitable for real-time deployment. Ling Hu 0001, Tao Yang 0041, Xionglve Li, Bingnan Hou, Zhiping Cai |
Comput. J. | 4 |
| 2026 | 6CAI: Efficient large-scale IPv6 cellular address identification
Ling Hu 0001, Xionglve Li, Bingnan Hou, Zhiyuan Jiang, Zhiping Cai |
Comput. Networks | 3 |
| 2026 | HMap: Efficient Internet-Wide IPv6 Scanning With Dynamic SearchabstractInternet-wide scanning is integral to network measurement and security analysis, but the expansive address space of IPv6 limits existing approaches in achieving efficient global-scale scans. This study introduces HMap, an innovative IPv6 scanner that markedly improves scan efficiency and coverage through the implementation of a dynamic search (DS) technique, relying solely on IPv6 routeable BGP prefixes. DS employs a dynamic feedback-driven probing strategy that uses information from previous replies to prioritize more promising address regions in subsequent scans. In Internet-wide scans over IPv6, encompassing both ping-like and traceroute-like scans with DS, HMap has demonstrated its capability to discover 2.29 million non-alias active target addresses, 0.13 million peripheries/middleboxes, and 1.61 million router interfaces, using only million-scale probes. This represents a noteworthy improvement of 1.91 times, 1.63 times, and 12.38 times, respectively, compared to current state-of-the-art alternatives. Additionally, by utilizing an efficient target generation algorithm (TGA) that more effectively leverages seed addresses, HMap expands the non-alias active address count to 44.05 million. This coverage spans 18.97 thousand ASes with a one-hour scan at a limited probing speed of 100 Kpps. The volume of active IPv6 addresses is 4.88 times larger than the currently disclosed largest IPv6 hitlists, providing a more diverse set of IPv6 networks. Unlike prior IPv6 scan studies that preclude their use for Internet-scale security analysis, we also conduct the Internet-wide security scans of IPv6 networks, focusing on the exposed internal IPv6 devices and security-sensitive services in IPv6 routers. Bingnan Hou, Zhenzhong Yang, Xianzheng Meng, Ling Hu 0001, Xionglve Li, Zhiping Cai |
IEEE Trans. Netw. Serv. Manag. | 7 |
| 2026 | Comprehensive Measurement of IPv6 Inbound Source Address Validation Deployment via Global Counter Side-Channel
Ling Hu 0001, Zhihuang Liu, Xionglve Li, Bingnan Hou, Zhiyuan Jiang, Bo Yu 0008, Zhiping Cai |
IEEE Trans. Netw. | 4 |
| 2025 | Accurate and Efficient Fine-Tuning of Quantized Large Language Models Through Optimal Balance in AdaptationabstractAbstract Large Language Models (LLMs) have demonstrated impressive performance across various domains. However, the enormous number of model parameters makes fine-tuning challenging, significantly limiting their application and deployment. Existing solutions combine parameter quantization with Low-Rank Adaptation (LoRA), reducing memory usage but causing performance degradation. Additionally, converting fine-tuned models to low-precision representations further degrades performance. In this paper, we identify an imbalance in fine-tuning quantized LLMs with LoRA: overly complex adapter inputs and outputs versus low effective trainability of the adapter, leading to underfitting during fine-tuning. Thus, we propose Quantized LLMs fine-tuning with Balanced Low-Rank Adaptation (Q-BLoRA), which simplifies the adapter’s inputs and outputs while increasing the adapter’s rank to alleviate underfitting during fine-tuning. For low-precision deployment, we propose Quantization-Aware fine-tuning with Balanced Low-Rank Adaptation (QA-BLoRA), which aligns with the block-wise quantization and facilitates quantization-aware fine-tuning of low-rank adaptation based on the parameter merging of Q-BLoRA. Both Q-BLoRA and QA-BLoRA are easily implemented and offer the following optimizations: (i) Q-BLoRA consistently achieves state-of-the-art accuracy compared to baselines and other variants; (ii) QA-BLoRA enables the direct generation of low-precision inference models, which exhibit significant performance improvements over other low-precision models. We validate the effectiveness of Q-BLoRA and QA-BLoRA across various models and scenarios. Code has been made available at https://github.com/xiaocaigou/qbaraqahira. Zhiquan Lai, Qiang Wang 0006, Xionglve Li, Dongsheng Li 0001 |
Trans. Assoc. Comput. Linguistics | 4 |
| 2025 | Realizing Personalized and Adaptive Inference of AS Paths With a Generative and Measurable ProcessabstractIn the global Internet, understanding paths between autonomous systems (ASes) is valuable for improving the Internet routing system and optimizing various applications. However, due to the business and privacy concerns, only a small portion of paths are disclosed. Moreover, limited by the measurement resources, obtaining paths between any two ASes is impossible. Thus, path inference becomes necessary. Recent work proposes training individual model for each AS to infer paths, but it lacks personalization as it uses a shared approach and data for arbitrary ASes. Moreover, training models from scratch for all the ASes is time-consuming and resource-intensive. This paper introduces Personalized and Adaptive Generative Measurable Path Inference (PA-GMPI), a prefix-grained path inference process. PA-GMPI is capable of achieving superior performance and faster model training by fully leveraging the exclusive information of each AS. These improvements come from a personalized path generator, a 3-layer graph kernel based adaptive training warm-starter, and a real-world walks based AS representation learner. In evaluation, PA-GMPI significantly outperforms the state-of-the-art method, achieving a maximal accuracy improvement of 28.72% and ESR (exact same ratio) improvement of 49.95%. Furthermore, PA-GMPI achieves an average reduction of 20.21% in training resource consumption across over two thousand training sessions, using vantage ASes from five snapshots, which included 439 distinct ASes. Xionglve Li, Chengyu Wang 0008, Tao Yang 0041, Zhenyu Qiu, Bingnan Hou, Zhiping Cai |
IEEE Trans. Netw. | 1 |
| 2024 | DRL-Tomo: a deep reinforcement learning-based approach to augmented data generation for network tomographyabstractAbstract Accurate and current comprehension of network status is crucial for efficient network management. Nevertheless, direct network measurement strategies entail substantial traffic overhead and demand intricate coordination among network entities, making them impractical. Network tomography, an indirect measurement approach, utilizes insights garnered from measured parts to deduce characteristics of the entire network. Past studies frequently depend on acquiring challenging-to-access information, such as the complete network topology or support from specialized protocols. Unfortunately, these constraints pose challenges in non-cooperative scenarios where obtaining such information is difficult. Recent endeavors pursue emancipating tomography from dependence on copious information, striving to predict unmeasured path performance using limited data. Nevertheless, the disparity between the measured data and actual performance has hindered the accuracy. In response, we introduce an innovative tomography framework named DRL-Tomo, designed to alleviate potential biases. DRL-Tomo initiates by generating augmented data through deep reinforcement learning, gradually approximating the genuine performance of unmeasured paths. Subsequently, a neural network model is trained using this augmented data, enabling precise inferences. Our experiments, encompassing both real-world and synthetic datasets, vividly demonstrate DRL-Tomo’s remarkable enhancement. Specifically, it achieves a substantial 10%–67% improvement in path delay prediction and an impressive 30%–98% enhancement in path loss rate prediction. Changsheng Hou, Bingnan Hou, Xionglve Li, Tongqing Zhou, Yingwen Chen 0001, Zhiping Cai |
Comput. J. | 3 |
| 2024 | DGA domain embedding with deep metric learningabstractAbstract Botnets currently use domain-generation algorithms to produce fast-flux domains that enable them to evade detection. Accurately categorizing these botnet domains is crucial to develop cybersecurity solutions against botnet threats. However, existing methods, requiring labeled data, are ineffective against new botnets. To address this issue, we propose Domain2Vec, a metric learning-based approach that can explore new botnets. Domain2Vec integrates a framework of metric learning, which uses individual domains from known botnets for categorization of unknown botnet domains. The training involves an attention-based encoder, and it includes a constraint to ensure that samples with the same labels are closer in the embedding space. The categorization uses the encoder to project domain names into appropriate representations (numerical vectors), even for domains from new botnets. Finally, Domain2Vec uses numerical vectors to explore botnets. Experiments showed that Domain2Vec performs well on domain retrieval and clustering tasks without labeled data, outperforming the state of the art by 13% and 100%, respectively. Real-world tests demonstrate that Domain2Vec can effectively identify unreported malicious domains and monitor botnet activities. Xionglve Li, Tao Yang 0041, Bingnan Hou, Lingbin Zeng, Zhiping Cai, Wenyuan Kuang |
Comput. J. | 2 |
| 2024 | Generative adversarial minority enlargement - A local linear over-sampling synthetic method
Ke Wang 0044, Tongqing Zhou, Menghua Luo, Xionglve Li, Zhiping Cai |
Expert Syst. Appl. | 4 |
| 2023 | Realizing Fine-Grained Inference of AS Path With a Generative Measurable ProcessabstractIn the global Internet, the paths between two autonomous systems (ASes), which are used for the exchange of traffic, are essential for understanding the behavior of the Internet routing system and they can help improve the performance of many applications of the Internet. Popular approaches to obtain the AS path between an AS pair (AP) are measurement based (e.g., Traceroute), but considering the size of the modern Internet and the limitations of measurement resources, only paths between a very small portion of APs can be measured. In recent years, a large body of path inference approaches has been proposed to bridge the gap in measurement resources. However, as we show with experiments, they perform poorly in accuracy and coverage. We propose a generative measurable path inference (GMPI) framework for AS-level path measurement, which performs well in accuracy and coverage. GMPI addresses two limitations of previous approaches: 1) Information incompleteness due to unrevealed real-world AS-level routing policies and insufficient measuring resources. 2) Knowledge isolation caused by distributed AS knowledge with different sources and inconsistent forms. To overcome these challenges, the data-driven GMPI framework invents heuristic path generation to address incompleteness and a dual-attention network to integrate the isolated knowledge. GMPI does not perform any measurement or impose any burden on the network. Our performance evaluation shows that our framework GMPI outperforms state-of-the-art approaches in terms of accuracy and coverage. In particular, compared to the state-of-the-art stitching-based baseline, GMPI provides a 42.45% improvement in coverage and a 39.97% improvement in accuracy. The experimental results demonstrate that GMPI can accurately infer paths for nearly arbitrary APs. Xionglve Li, Tongqing Zhou, Zhiping Cai, Jinshu Su |
IEEE/ACM Trans. Netw. | 1 |
| 2020 | ProbInfer: Probability-based AS path inference from multigraph perspective
Xionglve Li, Zhiping Cai, Bingnan Hou, Ning Liu 0015, Fang Liu 0002, Jieren Cheng |
Comput. Networks | 1 |
| 2020 | Large-scale graph processing systems: a surveyabstractGraph is a significant data structure that describes the relationship between entries. Many application domains in the real world are heavily dependent on graph data. However, graph applications are vastly different from traditional applications. It is inefficient to use general-purpose platforms for graph applications, thus contributing to the research of specific graph processing platforms. In this survey, we systematically categorize the graph workloads and applications, and provide a detailed review of existing graph processing platforms by dividing them into general-purpose and specialized systems. We thoroughly analyze the implementation technologies including programming models, partitioning strategies, communication models, execution models, and fault tolerance strategies. Finally, we analyze recent advances and present four open problems for future research. Ning Liu 0015, Dongsheng Li 0001, Yiming Zhang 0003, Xionglve Li |
Frontiers Inf. Technol. Electron. Eng. | 4 |