EDBT 2026 Demo / reviewers in the wild / expert
Xiaodong Lee
dblp:86/4349 · also Xiao-Dong Lee
· DBLP profile ↗
24ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0002-1757-8365ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 5 since 2021Artificial intelligence and machine learning · 8 · 2 since 2021Computer networks · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Security and privacy · 3 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Structured policy modeling and context-aware generation for multi-jurisdictional compliance in global software systems
Zhixian Zhuang, Xiaodong Lee, Aiyao Zhang, Jiuqi Wei, Yufan Fu, Botao Peng |
Inf. Softw. Technol. | 2 |
| 2026 | PDET-LSH: Scalable In-Memory Indexing for High-Dimensional Approximate Nearest Neighbor Search With Quality GuaranteesabstractLocality-sensitive hashing (LSH) is a well-known solution for approximate nearest neighbor (ANN) search with theoretical guarantees. Traditional LSH-based methods mainly focus on improving the efficiency and accuracy of query phase by designing different query strategies, but pay little attention to improving the efficiency of the indexing phase. They typically fine tune existing data-oriented partitioning trees to index data points and support their query strategies. However, their strategy to directly partition the multidimensional space is time-consuming, and performance degrades as the space dimensionality increases. In this paper, we design an encoding-based tree called Dynamic Encoding Tree (DE-Tree) to improve the indexing efficiency and support efficient range queries. Based on DE-Tree, we propose a novel LSH scheme called DET-LSH. DET-LSH adopts a novel query strategy, which performs range queries in multiple independent index DE-Trees to reduce the probability of missing exact NN points. Extensive experiments demonstrate that while achieving best query accuracy, DET-LSH achieves up to 6x speedup in indexing time and 2x speedup in query time over the state-of-the-art LSH-based methods. In addition, to further improve the performance of DET-LSH, we propose PDET-LSH, an in-memory method adopting the parallelization opportunities provided by multicore CPUs. PDET-LSH exhibits considerable advantages in indexing and query efficiency, especially on large scale datasets. Extensive experiments show that, while achieving the same query accuracy as DET-LSH, PDET-LSH offers up to 40x speedup in indexing time and 62x speedup in query answering time over the state-of-the-art LSH-based methods. Our theoretical analysis demonstrates that DET-LSH and PDET-LSH offer probabilistic guarantees on query answering accuracy. Jiuqi Wei, Xiaodong Lee, Botao Peng, Quanqing Xu, Chuanhui Yang, Themis Palpanas |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | DeFiAD: A Unified Method for Early-Stage Domain Abuse Detection Through Automated Deep Feature InteractionabstractAs a key infrastructure of the Internet, the Domain Name System (DNS) is frequently abused because of its open-ness, which makes early detection of domain abuse a critical task in cybersecurity. However, most existing methods rely on external data sources such as webpage content or long-term DNS resolution logs, which are unavailable at the registration stage, delaying detection and limiting adaptability across attack types. To overcome these limitations, we propose DeFiAD, a unified framework for early-stage domain abuse detection that relies solely on multi-source static information available at registration, enabling detection at the beginning of the domain lifecycle. We first introduce UDIRP, a unified domain information representation paradigm, which maps multi-source heterogeneous DNS native data into three subspaces for model understanding, forming a unified multi-channel representation. Building on this representation, we design DCINet, a Deep & Cross Network–based feature interaction model that effectively learns and integrates multi-source heterogeneous features, improving robustness and adaptability across diverse domain abuse types. Experiments on phishing, DGA, and mixed-abuse datasets show that DeFiAD consistently achieves the best overall performance, improving detection accuracy by 4%–10% over existing methods while maintaining high inference efficiency, highlighting its practicality for timely deployment. Zhaojun Dai, Xiaodong Lee, Yufan Fu, Botao Peng |
TrustCom | 2 |
| 2025 | POLARIS: Cross-Domain Access Control via Verifiable Identity and Policy-Based AuthorizationabstractAccess control is a security mechanism designed to ensure that only authorized users can access specific resources. Cross-domain access control involves access to resources across different organizations, institutions, or applications. Traditional access control, however, which handles authentication and authorization separately in centralized environments, faces challenges in identity dispersion, privacy leakage, and diversified permission requirements, failing to adapt to cross-domain scenarios. Thus, there is an urgent need for a new access control mechanism that empowers autonomous control over user identity and resources, addressing the demands for privacy-preserving authentication and flexible authorization in cross-domain scenarios.To address cross-domain access control challenges, we propose POLARIS, a unified and extensible architecture that enables policy-based, verifiable and privacy-preserving access control across different domains. POLARIS features a structured commitment mechanism for reliable, fine-grained, policy-based identity disclosure. It further introduces VPPL, a lightweight policy language that supports issuer-bound evaluation of selectively revealed attributes. A dedicated session-level security mechanism ensures binding between authentication and access, enhancing confidentiality and resilience to replay attacks.We implement a working prototype and conduct comprehensive experiments, demonstrating that POLARIS effectively provides scalable, privacy-preserving, and interoperable access control across heterogeneous domains. Our results highlight the practical viability of POLARIS for enabling secure and privacy-preserving access control in decentralized, cross-domain environments. Aiyao Zhang, Xiaodong Lee, Zhixian Zhuang, Jiuqi Wei, Yufan Fu, Botao Peng |
TrustCom | 2 |
| 2025 | Subspace Collision: An Efficient and Accurate Framework for High-dimensional Approximate Nearest Neighbor SearchabstractApproximate Nearest Neighbor (ANN) search in high-dimensional Euclidean spaces is a fundamental problem with a wide range of applications. However, there is currently no ANN method that performs well in both indexing and query answering performance, while providing rigorous theoretical guarantees for the quality of the answers. In this paper, we first design SC-score, a metric that we show follows the Pareto principle and can act as a proxy for the Euclidean distance between data points. Inspired by this, we propose a novel ANN search framework called Subspace Collision (SC), which can provide theoretical guarantees on the quality of its results. We further propose SuCo, which achieves efficient and accurate ANN search by designing a clustering-based lightweight index and query strategies for our proposed subspace collision framework. Extensive experiments on real-world datasets demonstrate that both the indexing and query answering performance of SuCo outperform state-of-the-art ANN methods that can provide theoretical guarantees, performing 1-2 orders of magnitude faster query answering with only up to one-tenth of the index memory footprint. Moreover, SuCo achieves top performance (best for hard datasets) even when compared to methods that do not provide theoretical guarantees. Jiuqi Wei, Xiaodong Lee, Zhenyu Liao 0001, Themis Palpanas, Botao Peng |
Proc. ACM Manag. Data | 2 |
| 2025 | Dominate data by yourself: a decentralized scheme for data interoperation when data is decoupled from applications
Jiuqi Wei, Xiaodong Lee, Yufan Fu, Ying Li 0051, Botao Peng |
World Wide Web (WWW) | 2 |
| 2024 | CBCMS: A Compliance Management System for Cross-Border Data TransferabstractCross-border data transfer is vital for the digital economy by enabling data flow across different countries or regions. However, ensuring compliance with diverse data protection regulations during the transfer introduces significant complexities. Existing solutions either focus on a single legal framework or neglect real-time and concurrent processing demands, resulting in incomplete and inconsistent compliance management. To address this issue, we propose Cross-Border Compliance Management System (CBCMS), which not only enables the unified management of data processing policies across multiple jurisdictions to ensure compliance with various legal frameworks involved in cross-border data transfer, but also supports real-time and high-concurrency processing capabilities. We design Policy Definition Language (PDL) that supports the unified management of data processing policies, bridging the gap between natural language policies and machine-processable expressions, thereby allowing various legal frameworks to be seamlessly integrated into CBCMS. We present Compliance Policy Generation Model (CPGM), the core component of CBCMS, which generates compliant data processing policies with high accuracy, achieving up to 25.16% improvement in F1 score (reaching 97.32%) compared to rule-based baseline. CPGM achieves inference time in the order of milliseconds (6 to 13 ms), and keeps low latency even under high-load scenarios, demonstrating high real-time and concurrent performance. To our knowledge, CBCMS is the first system to support unified compliance management across jurisdictions while ensuring real-time and concurrent processing capabilities. Zhixian Zhuang, Xiaodong Lee, Jiuqi Wei, Yufan Fu, Aiyao Zhang |
IEEE Big Data | 2 |
| 2024 | Securing the internet's backbone: A blockchain-based and incentive-driven architecture for DNS cache poisoning defenseabstractDomain Name System (DNS) is the backbone of the Internet infrastructure, converting human-friendly domain names into machine-processable IP addresses. However, DNS remains vulnerable to various security threats, such as cache poisoning attacks , where malicious attackers inject false information into DNS resolvers’ caches. Although efforts have been made to enhance DNS against such vulnerabilities, existing countermeasures often fall short in one or more areas: they may offer limited resistance to the collusion attack, introduce significant overhead, or require complex implementation that hinders widespread adoption. To address these challenges, this paper introduces TI-DNS+, a trusted and incentivized blockchain-based DNS resolution architecture for cache poisoning defense. TI-DNS+ introduces a Verification Cache exploiting blockchain ledger’s immutable nature to detect and correct forged DNS responses . The architecture also incorporates a multi-resolver Query Vote mechanism, enhancing the ledger’s credibility by validating each record modification through a stake-weighted algorithm. This algorithm selects resolvers as validators based on their stake proportion. To promote well-behaved participation, TI-DNS+ also implements a novel stake-based incentive mechanism that optimizes the generation and distribution of stake rewards. This ensures that incentives align with participants’ contributions, achieving incentive compatibility, fairness, and efficiency. Moreover, TI-DNS+ possesses high practicability as it requires only resolver-side modifications to current DNS. Finally, through comprehensive prototyping and experimental evaluations, the results demonstrate that our solution effectively mitigates DNS cache poisoning. Compared to competitors, our solution improves attack resistance by 1-3 orders of magnitude, while also reducing resolution latency by 5% to 68%. Yufan Fu, Xiaodong Lee, Jiuqi Wei, Ying Li 0051, Botao Peng |
Comput. Networks | 2 |
| 2024 | DiSAuth: A DNS-based secure authorization framework for protecting data decoupled from applications
Ying Li 0051, Jiuqi Wei, Ziyu Fei, Yufan Fu, Xiaodong Lee |
Comput. Networks | 5 |
| 2024 | Streaming Data Collection With a Private Sketch-Based ProtocolabstractData stream collection is critical to analyze service conditions and detect anomalies in time, especially in Internet of Things. However, it may undermine the individual privacy. Local differential privacy (LDP) has recently become a popular privacy-preserving technique protecting users’ privacy. However, most of them are still limited to the assumption of one-item collection, resulting in poor utility when extended to the multi-item collection from a very large domain. This article proposes a private streaming data collection framework, private sketch-based framework (PSF), which takes advantage of sketches. Combining the proposed background information and a decode-first collection-side workflow, the framework improves the utility by reducing the errors introduced by the sketching algorithm and the privacy budget utilization when collecting multiple items. We analytically prove the superior accuracy and privacy characteristics of PSF. In order to support specific computing tasks, we build two private protocols based on PSF, PrivSketch and PrivSketch+, aiming at frequency estimation and mean estimation, respectively. We demonstrate the utility of PrivSketch and PrivSketch+ theoretically, and also evaluate them experimentally. Our evaluation, with several diverse synthetic and real data sets, demonstrates that PrivSketch is 1–3 orders of magnitude better than the competitors in terms of utility in both frequency estimation and frequent item estimation, while being up to ~100x faster. PrivSketch+ performs ~4 orders of magnitude better than advanced solutions, such as piecewise mechanism (PM) and hybrid mechanism (HM), under a limited privacy budget. Ying Li 0051, Xiaodong Lee, Botao Peng, Themis Palpanas, Jing'an Xue |
IEEE Internet Things J. | 2 |
| 2024 | DET-LSH: A Locality-Sensitive Hashing Scheme with Dynamic Encoding Tree for Approximate Nearest Neighbor SearchabstractLocality-sensitive hashing (LSH) is a well-known solution for approximate nearest neighbor (ANN) search in high-dimensional spaces due to its robust theoretical guarantee on query accuracy. Traditional LSH-based methods mainly focus on improving the efficiency and accuracy of the query phase by designing different query strategies, but pay little attention to improving the efficiency of the indexing phase. They typically fine-tune existing data-oriented partitioning trees to index data points and support their query strategies. However, their strategy to directly partition the multi-dimensional space is time-consuming, and performance degrades as the space dimensionality increases. In this paper, we design an encoding-based tree called Dynamic Encoding Tree (DE-Tree) to improve the indexing efficiency and support efficient range queries based on Euclidean distance. Based on DE-Tree, we propose a novel LSH scheme called DET-LSH. DET-LSH adopts a novel query strategy, which performs range queries in multiple independent index DE-Trees to reduce the probability of missing exact NN points, thereby improving the query accuracy. Our theoretical studies show that DET-LSH enjoys probabilistic guarantees on query accuracy. Extensive experiments on real-world datasets demonstrate the superiority of DET-LSH over the state-of-the-art LSH-based methods on both efficiency and accuracy. While achieving better query accuracy than competitors, DET-LSH achieves up to 6x speedup in indexing time and 2x speedup in query time over the state-of-the-art LSH-based methods. Jiuqi Wei, Botao Peng, Xiaodong Lee, Themis Palpanas |
Proc. VLDB Endow. | 3 |
| 2023 | PrivSketch: A Private Sketch-Based Frequency Estimation Protocol for Data Streams
Ying Li 0051, Xiaodong Lee, Botao Peng, Themis Palpanas, Jing'an Xue |
DEXA (1) | 2 |
| 2017 | A robust internet abuse detection methodabstractJavaScript can modify HyperText Markup Language(HTML) code.tag of HTML can load pages from other websites. These technologies are widely used nowadays for constructing flexible and robust web services. However, illegal websites also use these technologies to hide illegal contents. For traditional methods using web text to detect these illegal webpages (just getting initial source codes that aren't parsed by a browser), they can't give a precise judgment. In this paper, we solved the challenge and proposed a robust Internet abuse detection method. We get HTML codes that had been parsed by simulating the progress how browsers work in order to gather the necessary materials for text detection methods. Besides, we not only extracted the text features of HTML code, but also extracted structure features of HTML codes. With the experiments under detecting Internet abuse scenarios, we demonstrated that the proposal is efficient. Zhou Fa, Guanggang Geng, Zhiwei Yan, Xiaodong Lee |
IEEE BigData | 4 |
| 2017 | Towards tackling privacy disclosure issues in Domain Name ServiceabstractServing as the global Internet's phonebook, the Domain Name Service (DNS) helps to translate human-friendly domain names into machine-readable IP addresses, which makes DNS of great importance to the operation of the Internet and virtually relied on by today's almost all kinds of Internet-based activities. As such, people whoever want to go anywhere over the Internet will need to refer to the DNS first. Therefore, it has become an ideal way to conduct online privacy exploitations through the DNS due to people's pervasive usage of the Internet. However, the current DNS doesn't provide any countermeasure against this kind of exploitation, and thus risks severe privacy disclosure problems. In this paper, we give a comprehensive empirical analysis of DNS privacy disclosure problems by exploring potential privacy leaking paths in the DNS. Then we further identify and describe multiple criterions of validity systematically that are obligated when considering DNS privacy preservation. Finally, we propose a simple DNS privacy preserving solution with significant deployment potential in the current DNS, which can only lead to a moderate level of extra query latency perceived by end users. Xuebiao Yuchi, Guanggang Geng, Zhiwei Yan, Xiaodong Lee |
IM | 4 |
| 2016 | Phishing detection based on newly registered domainsabstractPhishing is a security attack that involves the creation of websites that mimic legitimate websites, and these fraud websites bring Internet users a lot of loss. Traditional anti-phishing methods usually worked in a passive way by receiving report data of user. Due to the growing shorter survival time of phishing, this kind of methods is not efficient enough to find and take down new phishing attacks. In this paper, we propose an Intelligent Phishing Detection (IPD) system to address phishing detection problem actively. Specially, IPD first generates the detection dataset from the global massive domain name registration data automatically; then it applies the Naïve Bayes algorithm which is optimized by position-based features to achieve the high precision detection; finally, in order to find more phishing websites, IPD expands detection dataset by generating Uniform Resource Locator (URL) templates based on the detection results. The experimental results of IPD demonstrate the effectiveness and timeliness in detection phishing websites. Xueni Li, Guanggang Geng, Zhiwei Yan, Xiaodong Lee |
IEEE BigData | 5 |
| 2016 | Dealing with temporary domain name issues in the DNSabstractRecently, a new type of domain names, namely temporary domain names, has become heavily used by many kinds of Internet services, such as cloud storage and social networks. Generally, these services would generate a large volume of temporary domain names within their private domain zones, and use them to convey “one-time-signals” with their customers. By analyzing the real world DNS traffic, we find that over 40% of observed domain names from the Internet are temporary. While this creative usage of temporary domain names could benefit many Internet services, it may also cause some unanticipated (even negative) consequences to the DNS infrastructure. In this paper, we first describe the critical features of temporary domain names and present our observation results of their pervasiveness based on real world DNS traces collected from some major ISP. Then we verify and analyze quantitatively the negative impact that temporary domain names would cause on the DNS, especially on the DNS caching functionalities. To solve this problem, we finally introduce segmented caching strategies into the DNS cache, and further validate its capability for ensuring the effectiveness of DNS caching when facing the temporary domain name problems. Xuebiao Yuchi, Xiaodong Lee, Lanlan Pan |
ISCC | 2 |
| 2016 | A Novel DMM Architecture Based on NDN
Zhiwei Yan, Jong-Hyouk Lee, Guanggang Geng, Xiaodong Lee |
QSHINE | 4 |
| 2015 | Combating phishing attacks via brand identity and authorization featuresabstractAbstract Phishing, also called brand spoofing, has become the most troubling scam on the Internet, which seriously threatens the Web security. The essence of phish is that “robbers” use false sites, which look like a trustworthy brand site, where favicon, logo and copyright notice are important brand identities. We analyzed 78‐day phishing data of PhishTank and Anti‐Phishing Working Group (APWG). The statistics show that more than 98.93% phishing sites contain at least one brand entity—favicon, logo or copyright notice. Indeed, only a few lowest‐quality phishing campaigns do not use such brand elements. Obviously, brand entities are powerful weapons of phishers to trick users. By analyzing the characteristics of brand entities in phishing sites, several brand identity features are extracted. However, only brand entities do not consider whether the Web page with brand entities belongs to the corresponding brand or has an authorization to use the brand entities. To solve this problem, redirection, incoming links and Domain Name System (DNS) information‐based brand authorization features are further extracted to discriminate the sites with branding rights from phishing sites. Based on extracted features, statistical anti‐phishing classification models are trained. We collected a diverse spectrum of corpora containing 3863 phishing cases from PhishTank and APWG, and 17 571 legitimate samples from DMOZ, Google and DNS resolution log. Experimental evaluations show that the model achieves 98.8% true positive rate and 0.09% false positive rate, which demonstrates the competitive performances of extracted features for statistical anti‐phishing in practice. Copyright © 2014 John Wiley & Sons, Ltd. Guanggang Geng, Xiaodong Lee, Yan-Ming Zhang 0001 |
Secur. Commun. Networks | 2 |
| 2014 | Enhanced HMIPv6 with cascaded tunnelabstractHierarchical Mobile IPv6 (HMIPv6) is designed to reduce the handover latency and signaling load of Mobile IPv6 (MIPv6). However, HMIPv6 only improves local mobility performance while the handover latency in the inter-Mobility Anchor Point (MAP) handover is still long and packet transmission cost is still heavy due to the overlapped tunnels. In this paper, we propose a cascaded tunnel scheme in HMIPv6 to reduce its handover latency and signaling cost. In the proposed scheme, the home registration of Mobile Node (MN) is conducted by the MAP in a network-based manner and the global tunnel is established between MAP and Home Agent (HA), which allows the MN to originate only the local registration signaling even during its inter-MAP handover. The improvements contributed by the proposed scheme over the basic HMIPv6 are shown to be significant from the analyzing results. Zhiwei Yan, Xiaodong Lee |
PIMRC | 2 |
| 2010 | Investigating Sequential Patterns of DNS Usage and Its Applications
Xiaodong Lee, Baoping Yan |
ADMA (1) | 3 |
| 2010 | Modeling DNS Activities Based on Probabilistic Latent Semantic Analysis
Xuebiao Yuchi, Xiaodong Lee, Baoping Yan |
ADMA (2) | 2 |
| 2010 | A New Statistical Approach to DNS Traffic Anomaly Detection
Xuebiao Yuchi, Xiaodong Lee, Baoping Yan |
ADMA (2) | 3 |
| 2010 | Detecting DDoS Attack towards DNS Server Using a Neural Network Classifier
Xiaodong Lee, Baoping Yan |
ICANN (3) | 3 |
| 2009 | Efficient Filtering Algorithm for Repeated Queries in DNS LogabstractThe domain name system (DNS) is a fundamental component of the modern Internet, and repeated queries make up of a large amount of DNS traffic. To filter out the repeated queries in BIND DNS log file, an efficient algorithm is proposed. The algorithm is characterized by maintaining the time sequence of original queries during the processing. The experimental results using CN TLD root server log are analyzed. Xiaodong Lee, Baoping Yan |
NCA | 2 |