EDBT 2026 Demo / reviewers in the wild / expert
Chengxi Xu
dblp:273/1561
· DBLP profile ↗
18ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0001-7405-7710ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 9 · 9 since 2021Computer networks · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Device Type Identification with Deep Metric Learning
Xinyu Yin, Fan Shi 0003, Chengxi Xu, Jinfeng Peng, Jiatang Zhao |
ICIC (4) | 3 |
| 2026 | Beyond Up and Down: Analyzing Temporal Phenomena in IPv6 Probing Responses
Yudong Lian, Jinfeng Peng, Chengxi Xu, Fan Shi 0003, Min Zhang 0054, Jiatang Zhao, Quming Peng |
IWQoS | 4 |
| 2026 | CoT-DPG: A Co-Training based Dynamic Password Guessing Method
Fan Shi 0003, Min Zhang 0054, Chengxi Xu, Shasha Guo 0001, Jinghua Zheng |
NDSS | 4 |
| 2026 | TNet: Efficient IPv6 active network discovery
Jiatang Zhao, Fan Shi 0003, Chengxi Xu, Jinfeng Peng, Mingyi Ge, Min Zhang 0054 |
Comput. Networks | 3 |
| 2026 | Darkness at dawn: understanding illicit websites in newly registered domain namesabstractAbstract Illicit website represents a significant challenge on the Internet. Miscreants exploit the inherent flexibility and invisibility of the Internet to promote illicit activities, particularly online gambling and pornography, intending to generate substantial profits. Previous studies have primarily focused on illicit website detection techniques and analyzed illicit activities using passive datasets. However, constrained by the limitations of passive dataset perspectives, the security community lacks a global understanding of illicit website deployment and operational behavior patterns, particularly during the early stages of website activation. In this paper, we conduct an in-depth analysis of the activities of illicit websites through the advantageous lens of newly registered domains (NRDs). The NRD dataset’s key strength is its broad coverage of emerging illicit activities during observation, complementing previous studies. Specifically, we designed and implemented a framework, NRDMiner, for tracking and analyzing illicit activities associated with large-scale NRDs. This framework supports long-term monitoring of vast quantities of domains and enables accurate identification of illicit websites. Over a 133-day period (July 1–Nov 10, 2024), we collected 27,623,326 NRDs across 481 top-level domains (e.g., and ), and identified 910,794 abusive domains. Our analysis highlights several important patterns. First, illicit activity shows a consistent and steady pattern, with an average of 3.3% of NRDs flagged for illicit website. Moreover, 98% of these domains are first-time registrations. Second, 60% of abusive domains are activated on the same day they are registered, indicating mature automated domain abuse techniques. Third, from a global NRD perspective, we observed regional tendencies in illicit activities, like Asia identified as the primary concentration area, with over 70% of illicit website pages being in Asian languages. Furthermore, we analyzed the deployment and operation of illicit websites. Our work provides a large-scale empirical study of the early-stage activities of illicit websites from the perspective of NRDs, offering valuable evidence that contribute to the timely mitigation of illicit activities. Bingyang Guo, Fan Shi 0003, Min Zhang 0054, Chengxi Xu, Yi Shen 0012 |
Cybersecur. | 6 |
| 2026 | Trilink: discovering embedded siblings using a novel approachabstractAbstract Due to the bucket effect, dual-stack hosts face more severe security risks than single-stack hosts, making the discovery and identification of dual-stack hosts particularly important. Traditional studies employ methods such as domain name association and service fingerprinting for dual-stack identification; however, these methods suffer from incomplete identification and limited dual-stack scale. To solve this issue, we introduce the Trilink algorithm, which performs dual-stack host discovery and identification, as well as conducts security analysis, by verifying whether IPv6 addresses conform to the potential dual-stack address format standards, comparing the consistency of port fingerprints between IPv4 and IPv6, and utilizing IP geolocation and IP address ASN matching. The results show that we have discovered a total of 204,825 dual-stack devices across 118 countries and 269 autonomous systems. Meanwhile, our research reveals that dual-stack devices have 27% higher asset exposure across common service types than single-stack devices. Fan Shi 0003, Mingyi Ge, Chengxi Xu, Jiatang Zhao, Xinyu Yin |
Cybersecur. | 3 |
| 2026 | CyMapNER: a named entity recognition model for cyberspace surveying and mapping domainabstractAbstract Cyberspace Surveying and Mapping (CSM) involves the identification and analysis of digital assets to support network management and security, yet its domain-specific named entity recognition (NER) remains underexplored. A key challenge is the semantic gap between general-domain corpora and CSM domain texts, the suboptimal performance of existing named entity recognition (NER) models in accurately identifying entities within CSM data. To tackle obstacles, we proposed a NER model CyMapNER for the CSM domain. A clear definition of named entity categories pertinent to the CSM domain was established initially, followed by the creation of a dedicated NER dataset tailored to this domain. Subsequently, we present a domain adaptation training framework that integrates large language models. It combines with data augments, pseudo-labeling and domain-adaptive pretraining to enhance the adaptability of the NER model. The comparative experimental results demonstrate that CyMapNER models outperforms traditional NER models in CSM datasets. The results reveal that by domain adaptation training framework, the recognition accuracy of CyMapNER model reaches 97%, which achieves an improvement from 5.6% to 18.3% over the state-of-the-art NER models, and it performs well in recognizing complex and sparse entities, highlighting its effectiveness in handling the intricacies of CSM data. Fan Shi 0003, Chengxi Xu, Xinyu Yin, Mingyi Ge |
Cybersecur. | 3 |
| 2026 | PGMaP: Password generation based on mask predictionabstractNumerous studies have focused on data-driven password guessing methods in recent years, aiming to reduce the use of weak passwords by users and improve password security. Existing password generation models learn the distribution of password datasets and generate candidate guesses by fitting sequential conditional probabilities. These methods are based on a key assumption: users construct passwords in one direction from left to right. However, with the more complex password policy requirements of authentication systems and the increasing security awareness of people, users construct passwords by modifying existing or popular passwords. At this point, users consider global and bi-directional information of passwords. This breaks the key assumption of uni-directional construction and leads to omissions when generating passwords by existing methods. Motivated by this, we propose a password generation method based on mask prediction, named PGMaP, which captures this large number of omitted passwords. First, we design a password construction template extraction algorithm to cluster the templates used by users for constructing and modifying passwords. Then we construct a transformer-based masked language model to learn password bi-directional features. The extracted templates are fed into the model to generate password guesses by means of mask prediction. Different from existing auto-regressive model based methods that generate in one direction, PGMaP uses the auto-encoding model to generate passwords based on the bidirectional information. Finally, through password guessing experiments all eight real-world datasets, we demonstrate that PGMaP can effectively generate a large number of omitted passwords, and its password guessing performance outperforms existing methods. Fan Shi 0003, Shasha Guo 0001, Min Zhang 0054, Yi Shen 0012, Chengxi Xu |
Expert Syst. Appl. | 6 |
| 2025 | Email Cloaking: Deceiving Users and Spam Email Detectors with Invisible HTML Settings
Bingyang Guo, Mingxuan Liu 0006, Yihui Ma, Ruixuan Li 0008, Fan Shi 0003, Min Zhang 0054, Baojun Liu 0002, Chengxi Xu, Hai-Xin Duan, Geng Hong, Min Yang 0002, Qingfeng Pan |
ESORICS (4) | 8 |
| 2025 | Misty Registry: An Empirical Study of Flawed Domain Registry Operation
Mingming Zhang 0010, Baojun Liu 0002, Hai-Xin Duan, Min Zhang 0054, Fan Shi 0003, Chengxi Xu |
USENIX Security Symposium | 7 |
| 2025 | Understanding and Characterizing the Adoption of Internationalized Domain Names in PracticeabstractInternationalized Domain Names (IDNs) allow users to access the internet using domain names in their native languages. This technology provides significant convenience for non-English speaking users. However, despite the widespread acceptance and use of IDNs, the risks associated with using IDNs remain unclear in practice, such as the IDN homograph problem. To address this issue, we conduct a systematic analysis of the IDN homograph problem and explore the adoption characteristics of IDNs in practice. Specifically, we design and implement an effective IDN analysis framework, named as IDNMon. We perform a large-scale measurement study covering 863 top-level domain zone files and historical top lists based on IDNMon. Our findings indicate that the IDN registration and usage in Europe exceeds that in East Asia. Our results confirm that the IDN homograph problem is universal (12.32% of 2,623,161 IDNs face this problem), which raises serious challenges when designing protection strategies for browsers. Our work provides new insights into the adoption of IDNs in practice, contributes to a better understanding, and promotes the development of IDNs. Chengxi Xu, Fan Shi 0003, Min Zhang 0054, Yuwei Li 0002, Zhijie Xie |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2024 | CAKGC: A Clustering Method of Cybercrime Assets Knowledge Graph Based on Feature Fusion
Fan Shi 0003, Chengxi Xu, Jiankun Sun |
ICIC (9) | 3 |
| 2024 | Poster: A Fistful of Queries: Accurate and Lightweight Anycast Enumeration of Public DNS
Chengxi Xu, Fan Shi 0003 |
IMC | 1 |
| 2024 | Poster: Capture the List: Ranking Manipulation Leveraging Open Forwarders
Chengxi Xu, Fan Shi 0003 |
IMC | 1 |
| 2024 | Rethinking the Security Threats of Stale DNS Glue Records
Baojun Liu 0002, Hai-Xin Duan, Min Zhang 0054, Xiang Li 0108, Fan Shi 0003, Chengxi Xu, Eihal Alowaisheq |
USENIX Security Symposium | 7 |
| 2024 | Cross the Zone: Toward a Covert Domain Hijacking via Shared DNS Infrastructure
Mingming Zhang 0010, Baojun Liu 0002, Jia Zhang 0004, Hai-Xin Duan, Min Zhang 0054, Fan Shi 0003, Chengxi Xu |
USENIX Security Symposium | 9 |
| 2024 | An Intelligent Penetration Testing Method Using Human FeedbackabstractPenetration testing is widely acknowledged as the foremost method for evaluating network security. However, three challenges impede the generation of strategies that align with human expectations. In this article, we present, for the first time, a method based on human feedback to enhance strategy generation. Our approach comprises two components: agent training and decision-making. During agent training, we establish a hierarchical framework to decompose tasks and a knowledge base to offer advice for improving data efficiency. We then impose constraints on the action space to mitigate ineffective exploration. Finally, we train a reward model based on human feedback and fine tune the model guided by this reward model. In decision-making, we process the model output to enhance decision accuracy. We crafted scenarios based on real-world networks, and the results demonstrate the effectiveness of our method in generating penetration testing strategies that align more closely with human intentions. Qianyu Li 0001, Min Zhang 0054, Fan Shi 0003, Yi Shen 0012, Bingyang Guo, Chengxi Xu |
IEEE Trans. Ind. Informatics | 8 |
| 2021 | Mining Centralization of Internet Service Infrastructure in the WildabstractThe last decade has witnessed the rapid evolution of the Internet structure, one of which is centralization, that is, Internet core infrastructure has been constantly transferred into the hands of a few popular market participants. Researchers are trying to measure centrality and analyze its security impact from the perspective of traffic analysis. But the underlying distribution of service providers is still enveloped in mysterious veils. In order to address this problem and assess the security risk associated with such centralization. Firstly, we performed linear regression on the data of each kind of service provider in the Alexa Top 1M domains to study the current underlying distribution of various services for the first time. The results show that Zipf’s law is universal in various service providers’ market share, which proves that Internet service infrastructures are centralized. Secondly, we explored the security impacts of centralized infrastructures on the Internet. we conducted attack simulations on providers. Results show that intentional attacks on core providers can greatly downgrade the performance of the Internet. To make matters worse, the quantitative analysis of the provider’s infrastructures found that a considerable number of provider’s infrastructures have low diversity. In addition, we proposed an algorithm to calculate the dependencies between different types of service providers and carried out an evaluation of our datasets, and found the tendency for different services to depend on each other. Our results indicate that the Internet is facing huge security challenges, because the centralized infrastructure will impair service redundancy, and at the same time, it will also cause dependence between infrastructures, which in turn strengthens its centralization. Bingyang Guo, Fan Shi 0003, Chengxi Xu, Min Zhang 0054, Yang Li 0215 |
MSN | 3 |