EDBT 2026 Demo / reviewers in the wild / expert
Jia Zhang 0004
dblp:80/2266-4
· DBLP profile ↗
42ranked-venue papers
6as first author
27since 2021 · last 2026
0000-0001-7896-3382ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 17 · 14 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 8 since 2021Computer networks · 8 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorTheory of computation · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CtPhishCapture: Uncovering Credential-Theft-Based Phishing Scams Targeting Cryptocurrency Wallets
Zhenrui Zhang, Xiang Li 0108, Anpeng Zhou, Chenghui Wu, Man Hou, Jia Zhang 0004, Zongpeng Li |
NDSS | 8 |
| 2026 | Toward Robust Detection of Malicious Encrypted Traffic Using Only Low-Quality Training DataabstractMachine learning (ML) is promising in accurately detecting malicious flows in encrypted network traffic; however, it is challenging to collect a training dataset that contains a sufficient amount of encrypted malicious data with correct labels. When ML models are trained with low-quality training data, they suffer degraded performance. In this paper, we aim to address a real-world low-quality training dataset problem, namely, detecting encrypted malicious traffic generated by continuously evolving malware. We develop RAPIER+ that fully utilizes different distributions of normal and malicious traffic data in the feature space, where normal data is tightly distributed in a certain area, and the malicious data is scattered over the entire feature space to augment training data for model training. RAPIER+ includes two pre-processing modules to convert traffic into feature vectors and correct label noises. We evaluate our system on two public datasets and one combined dataset. With 1000 samples and 45% noise from each dataset, our system achieves the F1 scores of 0.78, 0.84, and 0.87, respectively, achieving average improvements of 358.5%, 314.0%, and 221.1% over the existing methods, respectively. Furthermore, we evaluate RAPIER+ with a real-world dataset obtained from a security enterprise. RAPIER+ effectively achieves encrypted malicious traffic detection with the best F1 score of 0.81 and improves the F1 score of existing methods by an average of 288.7%. Yuqi Qing, Qilei Yin, Xinhao Deng 0001, Zhuotao Liu, Kun Sun 0001, Ke Xu 0002, Jia Zhang 0004, Qi Li 0002 |
IEEE Trans. Netw. | 8 |
| 2025 | RebirthDay Attack: Reviving DNS Cache Poisoning with the Birthday ParadoxabstractDNS cache poisoning is a persistent game of attack and defense, posing an enduring challenge for the DNS community. Significant efforts have been made to uncover, detect, and mitigate vulnerabilities that increase the risk of cache poisoning. However, no work has systematically revisited whether the original cache poisoning attack based on the Birthday Paradox remains effective. In this work, we introduce RebirthDay, a novel DNS cache poisoning attack targeting recursive resolvers and forwarders, reviving the classic DNS Birthday attack that no longer works since 2002. RebirthDay exploits newly uncovered, protocol-compliant vulnerabilities in DNS extension implementations to bypass the query aggregation mechanism intended to prevent DNS Birthday attacks that has not been well understood. We uncovered that 18 out of 22 mainstream DNS software are vulnerable due to weaknesses in the processing of a DNS extension (i.e., ECS option), specifically lacking or incorrectly implemented ECS coherence checks when handling DNS queries and responses, demonstrating the widespread susceptibility to RebirthDay. These flaws could be exploited to circumvent the query aggregation mechanism and launch RebirthDay attacks. Through comprehensive evaluation, we showed that RebirthDay attacks are highly practical and can have significant real-world impact, affecting 16 router vendors, 14 public DNS services, and 365K (15%) open DNS resolvers. We have reported the identified vulnerabilities to affected vendors and discussed mitigation solutions with them. To date, we have received acknowledgments from 8 vendors, including BIND, Unbound, PowerDNS, and Quad9, and have been assigned 50 CVE-ids. Our study emphasizes the need for greater attention to the importance of ECS verification and DNS extension implementations, revealing new security risks introduced by them. Xiang Li 0108, Mingming Zhang 0010, Zuyao Xu, Fasheng Miao, Yuqi Qiu, Baojun Liu 0002, Jia Zhang 0004, Hai-Xin Duan, Zheli Liu, Yunhai Zhang, Dunqiu Fan |
CCS | 7 |
| 2025 | Decoding DNS Centralization: Measuring and Identifying NS Domains Across Hosting ProvidersabstractThe Domain Name System (DNS) is designed to be distributed, which aims to provide services with low latency and great reliability. However, after decades of development and changes in Internet business models, various aspects of the DNS ecosystem have begun to show signs of centralization. To investigate the centralization from the viewpoint of hosting service providers, we develop an automated method based on similarity among NS domains and co-hosting relationship to identify the hosting providers for authoritative name servers, so that we can identify hosting providers in DNS zone file to count the number of domains which a hosting provider host. This tool demonstrates greater accuracy than previous methods and our testing demonstrates the ability to identify hosting service providers for most domains in real-world measurement tasks. Using this tool, we conducted measurements on the dataset combined with .com, .net and .org TLD zones. We find that the top 10 providers collectively host over 54.19% of domains while top 100 providers host over 82.99% domains, which shows a significant level of centralization in hosting service providers within the DNS. Through an analysis of NSone’s NS domains and a statistical examination of top providers’ NS domains, we find that directly identifying the base domain as the provider is inappropriate. Furthermore, we discover relationships among hosting providers and between hosting providers and infrastructure that are more complex than previously anticipated. Finally, based on our research findings, we offer corresponding suggestions to mitigate the continued development of centralization. Qihang Peng, Mingming Zhang 0010, Deliang Chang, Jia Zhang 0004, Baojun Liu 0002, Hai-Xin Duan |
DSN | 4 |
| 2025 | Enhancing the Scalability and Applicability of Kohn-Sham Hamiltonians for Molecular SystemsabstractDensity Functional Theory (DFT) is a pivotal method within quantum chemistry and materials science, with its core involving the construction and solution of the Kohn-Sham Hamiltonian. Despite its importance, the application of DFT is frequently limited by the substantial computational resources required to construct the Kohn-Sham Hamiltonian. In response to these limitations, current research has employed deep-learning models to efficiently predict molecular and solid Hamiltonians, with roto-translational symmetries encoded in their neural networks. However, the scalability of prior models may be problematic when applied to large molecules, resulting in non-physical predictions of ground-state properties. In this study, we generate a substantially larger training set (PubChemQH) than used previously and use it to create a scalable model for DFT calculations with physical accuracy. For our model, we introduce a loss function derived from physical principles, which we call Wavefunction Alignment Loss (WALoss). WALoss involves performing a basis change on the predicted Hamiltonian to align it with the observed one; thus, the resulting differences can serve as a surrogate for orbital energy differences, allowing models to make better predictions for molecular orbitals and total energies than previously possible. WALoss also substantially accelerates self-consistent-field (SCF) DFT calculations. Here, we show it achieves a reduction in total energy prediction error by a factor of 1347 and an SCF calculation speed-up by a factor of 18\%. These substantial improvements set new benchmarks for achieving accurate and applicable predictions in larger molecular systems. Yunyang Li, Zaishuo Xia, Xinran Wei, Sam Harshe, Erpai Luo, Zun Wang 0006, Jia Zhang 0004, Chang Liu 0030, Bin Shao 0002, Mark Gerstein |
ICLR | 9 |
| 2025 | Efficient and Scalable Density Functional Theory Hamiltonian Prediction through Adaptive SparsityabstractHamiltonian matrix prediction is pivotal in computational chemistry, serving as the foundation for determining a wide range of molecular properties. While SE(3) equivariant graph neural networks have achieved remarkable success in this domain, their substantial computational cost—driven by high-order tensor product (TP) operations—restricts their scalability to large molecular systems with extensive basis sets. To address this challenge, we introduce SPHNet, an efficient and scalable equivariant network, that incorporates adaptive SParsity into Hamiltonian prediction. SPHNet employs two innovative sparse gates to selectively constrain non-critical interaction combinations, significantly reducing tensor product computations while maintaining accuracy. To optimize the sparse representation, we develop a Three-phase Sparsity Scheduler, ensuring stable convergence and achieving high performance at sparsity rates of up to 70%. Extensive evaluations on QH9 and PubchemQH datasets demonstrate that SPHNet achieves state-of-the-art accuracy while providing up to a 7x speedup over existing models. Beyond Hamiltonian prediction, the proposed sparsification techniques also hold significant potential for improving the efficiency and scalability of other SE(3) equivariant networks, further broadening their applicability and impact. Erpai Luo, Xinran Wei, Yunyang Li, Zaishuo Xia, Zun Wang 0006, Chang Liu 0030, Bin Shao 0002, Jia Zhang 0004 |
ICML | 10 |
| 2025 | Dive into the Cloud: Unveiling the (Ab)Usage of Serverless Cloud Function in the WildabstractServerless cloud functions transfer server management responsibilities to service providers, offering scalability and cost-efficiency. This convenience not only facilitates normal activities but also raises abuse concerns. So far, public understanding of real-world cloud functions remains limited. To fill this gap, we conducted an in-depth measurement study to uncover their practical usage and abuse. Through empirical analysis of nine leading providers (e.g., AWS, Tencent), we identified 531,089 function domains from a passive DNS dataset spanning April 2022 to March 2024. We first investigated the usage status of serverless cloud functions, showing the different practices between providers. Additionally, based on active requests to these functions, we pointed out privacy risks of unauthorized access and identified four abuse types, including covert C2 communication, hosting malicious websites, promoting illicit services, and abusing egress nodes as IP proxies. Alarmingly, 4.89% of cloud functions are being abused, with over 614k invocations recorded. Only four abused functions were flagged by existing threat intelligence systems, indicating critical gaps in security monitoring for serverless environments. Our work offers insights into the serverless cloud ecosystem and provides recommendations for better management. With responsible disclosure, we hope to raise awareness and improve protective measures against abuses among cloud function providers. Yijing Liu 0007, Mingxuan Liu 0006, Yiming Zhang 0009, Baojun Liu 0002, Jia Zhang 0004, Geng Hong, Hai-Xin Duan, Min Yang 0002 |
IMC | 5 |
| 2025 | Poster: RMap: Uncovering Risky DNS Resolution Chains and MisconfigurationsabstractIn recent years, large-scale network outages caused by DNS misconfigurations have become increasingly common. The intricate inter-domain dependencies, along with emerging mechanisms (Such as DNSSEC, EDNS, and 0x20), have made DNS resolution increasingly complex and fault localization more challenging. We present RMap, a tool that rapidly probes all potential resolution chains of a domain, reveals its resolution dependency topology, and detects security risks. We experimentally demonstrate the effectiveness of RMap and its broad applicability. Our findings reveal that domain configurations in real-world environments remain concerning, with potential issues observed even in several well-known top-level domains. RMap is avaliable in https://github.com/ahlien/rmap. Fasheng Miao, Shuying Zhuang, Xiang Li 0108, Changqing An, Deliang Chang, Baojun Liu 0002, Jia Zhang 0004, Jilong Wang 0001 |
IMC | 7 |
| 2025 | E2Former: An Efficient and Equivariant Transformer with Linear-Scaling Tensor ProductsabstractEquivariant Graph Neural Networks (EGNNs) have demonstrated significant success in modeling microscale systems, including those in chemistry, biology and materials science. However, EGNNs face substantial computational challenges due to the high cost of constructing edge features via spherical tensor products, making them almost impractical for large-scale systems.
To address this limitation, we introduce E2Former, an equivariant and efficient transformer architecture that incorporates a Wigner $6j$ convolution (Wigner $6j$ Conv). By shifting the computational burden from edges to nodes, Wigner $6j$ Conv reduces the complexity from $O(| \mathcal{E}|)$ to $O(| \mathcal{V}|)$ while preserving both the model's expressive power and rotational equivariance.
We show that this approach achieves a 7x–30x speedup compared to conventional $\mathrm{SO}(3)$ convolutions. Furthermore, our empirical results demonstrate that the derived E2Former mitigates the computational challenges of existing approaches without compromising the ability to capture detailed geometric information. This development could suggest a promising direction for scalable molecular modeling. Yunyang Li, Zhihao Ding, Xinran Wei, Zun Wang 0006, Chang Liu 0030, Peiran Jin, Tao Qin 0001, Mark Gerstein, Jia Zhang 0004 |
NeurIPS | 13 |
| 2025 | Detection and Mitigation of Unknown Threats in IPv6 Networks via Layered Data AdaptationabstractWith the rapid proliferation of IPv6 deployment, traditional threat detection approaches are increasingly challenged by the scale, diversity, and concealment of emerging attack behaviors. This paper tackles three key limitations in IPv6 threat detection: poor adaptability across diverse network environments, insufficient result validation mechanisms, and the absence of globally applicable detection toolsets. To address these challenges, we propose a data-driven, layer-adaptive detection framework that forms a closed-loop pipeline of detection, validation, and mitigation. Our framework constructs a hierarchical dataset covering multiple detection scenarios, including open sensitive ports, entropy-based traffic anomalies, and TensorFlow-enhanced machine learning inputs. We introduce a novel validation technique, the Daily Active Mutual Access Index Matrix, which captures inter-prefix interaction patterns to identify coordinated malicious behavior. Additionally, we deploy a global-scale threat intelligence resolution and measurement system to validate detection outcomes and uncover cross-border threats missed by conventional models. Extensive analysis of abuse reports and complaint email interactions reveals that 65.53% of observed abuse involves sensitive service ports, and 38.64% of detected addresses exhibit verifiable abuse behavior. Notably, detection methods based on information entropy and AI models demonstrate higher abuse report delivery rates compared to commercial threat intelligence sources. Experimental results confirm that our multi-modal, adaptive approach significantly enhances the accuracy of unknown threat detection and the effectiveness of coordinated response, offering scalable and robust technical support for IPv6 security operations. Youjun Huang, Xiang Li 0108, Jia Zhang 0004, Hai-Xin Duan |
TrustCom | 3 |
| 2025 | NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines
Mingxuan Liu 0006, Lijie Wu, Baojun Liu 0002, Geng Hong, Yiming Zhang 0009, Jia Zhang 0004, Hai-Xin Duan, Min Zhang 0054, Fan Shi 0003, Min Yang 0002 |
USENIX Security Symposium | 8 |
| 2024 | Long-Short-Range Message-Passing: A Physics-Informed Framework to Capture Non-Local Interaction for Scalable Molecular Dynamics SimulationabstractComputational simulation of chemical and biological systems using *ab initio* molecular dynamics has been a challenge over decades. Researchers have attempted to address the problem with machine learning and fragmentation-based methods. However, the two approaches fail to give a satisfactory description of long-range and many-body interactions, respectively. Inspired by fragmentation-based methods, we propose the Long-Short-Range Message-Passing (LSR-MP) framework as a generalization of the existing equivariant graph neural networks (EGNNs) with the intent to incorporate long-range interactions efficiently and effectively. We apply the LSR-MP framework to the recently proposed ViSNet and demonstrate the state-of-the-art results with up to 40% MAE reduction for molecules in MD22 and Chignolin datasets. Consistent improvements to various EGNNs will also be discussed to illustrate the general applicability and robustness of our LSR-MP framework. The code for our experiments and trained model weights could be found at https://github.com/liyy2/LSR-MP. Yunyang Li, Xinran Wei, Jia Zhang 0004, Tong Wang 0014, Zun Wang 0006, Bin Shao 0002, Tie-Yan Liu |
ICLR | 6 |
| 2024 | Bounce in the Wild: A Deep Dive into Email Delivery Failures from a Large Email Service ProviderabstractAbnormal email bounces seriously disrupt user lives and company transactions. Proliferating security protocols and protection strategies have made email delivery increasingly complex. A natural question is how and why email delivery fails in the wild. Filling this knowledge gap requires a representative global email delivery dataset, which is rarely disclosed by email service providers (ESPs). Ruixuan Li 0008, Shaodong Xiao, Baojun Liu 0002, Yanzhong Lin, Hai-Xin Duan, Qingfeng Pan, Jianjun Chen 0005, Jia Zhang 0004, Ximeng Liu, Xiuqi Lu, Jun Shao 0001 |
IMC | 8 |
| 2024 | Low-Quality Training Data Only? A Robust Framework for Detecting Encrypted Malicious Network Traffic
Yuqi Qing, Qilei Yin, Xinhao Deng 0001, Zhuotao Liu, Kun Sun 0001, Ke Xu 0002, Jia Zhang 0004, Qi Li 0002 |
NDSS | 8 |
| 2024 | TuDoor Attack: Systematically Exploring and Exploiting Logic Vulnerabilities in DNS Response Pre-processing with Malformed PacketsabstractDNS can be compared to a game of chess in that its rules are simple, yet the possibilities it presents are endless. While the fundamental rules of DNS are straightforward, DNS implementations can be extremely complex. In this study, we intend to explore the complexities and vulnerabilities in DNS response pre-processing by systematically analyzing DNS RFCs and DNS software implementations. We present the discovery of three new types of logic vulnerabilities, leading to the proposal of three novel attacks, namely the TuDoor attack. These attacks involve the use of malformed DNS response packets to carry out DNS cache poisoning, denial- of-service, and resource consuming attacks. By performing comprehensive experiments, we demonstrate the attack’s feasibility and significant real-world impacts of TUDOOR. In total, 24 mainstream DNS software, including BIND, PowerDNS, and Microsoft DNS, are affected by TuDoor. Attackers can instigate cache poisoning and denial-of-service attacks against vulnerable resolvers using a handful of crafted packets within 1 second or circumvent the query limit to deplete resolution resources (e.g., CPU). Besides, to determine the vulnerable resolver population in the wild, we collect and evaluate 16 popular Wi-Fi routers, 6 prevalent router OSes, 42 public DNS services, and around 1.8M open DNS resolvers. Our measurement results indicate that TUDOOR could exploit 7 routers (OSes), 18 public DNS services, and 424,652 (23.1%) open DNS resolvers. Following the best practice of responsible disclosure, we have reported these vulnerabilities to all affected vendors, and 18 of them, including BIND, Chrome, Cloudflare, and Microsoft, have acknowledged our findings and discussed mitigation solutions with us. Furthermore, 33 CVE IDs are assigned to our discovered vulnerabilities, and we provide an online detection tool as one of the mitigation measures. Our research highlights the urgent need for standardization of DNS response pre-processing logic to enhance the security of DNS. Xiang Li 0108, Wei Xu 0064, Baojun Liu 0002, Mingming Zhang 0010, Zhou Li 0001, Jia Zhang 0004, Deliang Chang, Chuhan Wang 0001, Jianjun Chen 0005, Hai-Xin Duan, Qi Li 0002 |
SP | 6 |
| 2024 | Cross the Zone: Toward a Covert Domain Hijacking via Shared DNS Infrastructure
Mingming Zhang 0010, Baojun Liu 0002, Jia Zhang 0004, Hai-Xin Duan, Min Zhang 0054, Fan Shi 0003, Chengxi Xu |
USENIX Security Symposium | 5 |
| 2024 | LordNet: An efficient neural network for learning to solve parametric partial differential equations without simulated data
Xinquan Huang, Wenlei Shi, Xiaotian Gao, Xinran Wei, Jia Zhang 0004, Jiang Bian 0002, Mao Yang 0004, Tie-Yan Liu |
Neural Networks | 5 |
| 2023 | Learning Physics-Informed Neural Networks without Stacked Back-propagationabstractPhysics-Informed Neural Network (PINN) has become a commonly used machine learning approach to solve partial differential equations (PDE). But, facing high-dimensional secondorder PDE problems, PINN will suffer from severe scalability issues since its loss includes second-order derivatives, the computational cost of which will grow along with the dimension during stacked back-propagation. In this work, we develop a novel approach that can significantly accelerate the training of Physics-Informed Neural Networks. In particular, we parameterize the PDE solution by the Gaussian smoothed model and show that, derived from Stein’s Identity, the second-order derivatives can be efficiently calculated without back-propagation. We further discuss the model capacity and provide variance reduction methods to address key limitations in the derivative estimation. Experimental results show that our proposed method can achieve competitive error compared to standard PINN training but is significantly faster. Di He 0001, Shanda Li, Wenlei Shi, Xiaotian Gao, Jia Zhang 0004, Jiang Bian 0002, Liwei Wang 0001, Tie-Yan Liu |
AISTATS | 5 |
| 2023 | Stolen Risks of Models with Security PropertiesabstractVerifiable robust machine learning, as a new trend of ML security defense, enforces security properties (e.g., Lipschitzness, Monotonicity) on machine learning models and achieves satisfying accuracy-security trade-off. Such security properties identify a series of evasion strategies of ML security attackers and specify logical constraints on their effects on a classifier (e.g., the classifier is monotonically increasing along some feature dimensions). However, little has been done so far to understand the side effect of those security properties on the model privacy. Zhuoqun Fu, Chuyun Deng, Xiaojing Liao, Jia Zhang 0004, Hai-Xin Duan |
CCS | 5 |
| 2023 | TsuKing: Coordinating DNS Resolvers and Queries into Potent DoS AmplifiersabstractIn this paper, we present a new DNS amplification attack, named TsuKing. Instead of exploiting individual DNS resolvers independently to achieve an amplification effect, TsuKing deftly coordinates numerous vulnerable DNS resolvers and crafted queries together to form potent DoS amplifiers. We demconstrate that with TsuKing, an initial small amplification factor can inrease exponentially through the internal layers of coordinated amplifiers, resulting in an extremely powerful amplification attack. TsuKing has three variants, including DNSRetry, DNSChain, and DNSLoop, all of which exploit a suite of inconsistent DNS implementations to achieve enormous amplification effect. With comprehensive measurements, we found that about 14.5% of 1.3M open DNS resolvers are potentially vulnerable to TsuKing. Real-world controlled evaluations indicated that attackers can achieve a packet amplification factor of at least 3,700X (DNSChain). We have reported vulnerabilities to affected vendors and provided them with mitigation recommendations. We have received positive responses from 6 vendors, including Unbound, MikroTik, and AliDNS, and 3 CVEs were assigned. Some of them are implementing our recommendations. Wei Xu 0064, Xiang Li 0108, Chaoyi Lu, Baojun Liu 0002, Hai-Xin Duan, Jia Zhang 0004, Jianjun Chen 0005, Tao Wan 0004 |
CCS | 6 |
| 2023 | Under the Dark: A Systematical Study of Stealthy Mining Pools (Ab)use in the WildabstractCryptocurrency mining is a crucial operation in blockchains, and miners often join mining pools to increase their chances of earning rewards. However, the energy-intensive nature of PoW cryptocurrency mining has led to its ban in New York State of the United States, China, and India. As a result, mining pools, serving as a central hub for mining activities, have become prime targets for regulatory enforcement. Furthermore, cryptojacking malware refers to self-owned stealthy mining pools to evade detection techniques and conceal profit wallet addresses. However, no systematic research has been conducted to analyze it, largely due to a lack of full understanding of the protocol implementation, usage, and port distribution of the stealth mining pool. Zhenrui Zhang, Geng Hong, Xiang Li 0108, Zhuoqun Fu, Jia Zhang 0004, Mingxuan Liu 0006, Chuhan Wang 0001, Jianjun Chen 0005, Baojun Liu 0002, Hai-Xin Duan, Chao Zhang 0008, Min Yang 0002 |
CCS | 5 |
| 2023 | NeuralStagger: Accelerating Physics-constrained Neural PDE Solver with Spatial-temporal DecompositionabstractNeural networks have shown great potential in accelerating the solution of partial differential equations (PDEs). Recently, there has been a growing interest in introducing physics constraints into training neural PDE solvers to reduce the use of costly data and improve the generalization ability. However, these physics constraints, based on certain finite dimensional approximations over the function space, must resolve the smallest scaled physics to ensure the accuracy and stability of the simulation, resulting in high computational costs from large input, output, and neural networks. This paper proposes a general acceleration methodology called NeuralStagger by spatially and temporally decomposing the original learning tasks into several coarser-resolution subtasks. We define a coarse-resolution neural solver for each subtask, which requires fewer computational resources, and jointly train them with the vanilla physics-constrained loss by simply arranging their outputs to reconstruct the original solution. Due to the perfect parallelism between them, the solution is achieved as fast as a coarse-resolution neural solver. In addition, the trained solvers bring the flexibility of simulating with multiple levels of resolution. We demonstrate the successful application of NeuralStagger on 2D and 3D fluid dynamics simulations, which leads to an additional $10\sim100\times$ speed-up. Moreover, the experiment also shows that the learned model could be well used for optimal control. Xinquan Huang, Wenlei Shi, Yue Wang 0017, Xiaotian Gao, Jia Zhang 0004, Tie-Yan Liu |
ICML | 6 |
| 2022 | HDiff: A Semi-automatic Framework for Discovering Semantic Gap Attack in HTTP ImplementationsabstractThe Internet has become a complex distributed network with numerous middle-boxes, where an end-to-end HTTP request is often processed by multiple intermediate servers before it reaches its destination. However, a general problem in this distributed network is the semantic gap attack, which is defined as inconsistent semantic interpretations in the processing chain. While some studies have found individual semantic gap attacks, most of them are based on ad-hoc manual analysis, which is inadequate for fundamentally enhancing the security assurance of a system as complex as the HTTP network.In this work, we propose HDiff, a novel semi-automatic detecting framework, systematically exploring semantic gap attacks in HTTP implementations. We designed a documentation analyzer that employs natural language processing techniques to extract rules from specifications, and utilized differential testing to discover semantic gap attacks. We implemented and evaluated it to find three kinds of semantic gap attacks in 10 popular HTTP implementations. In total, HDiff found 14 vulnerabilities and 29 affected server pairs covering all three types of attacks. In particular, HDiff also discovered three new types of attack vectors. We have already duly reported all identified vulnerabilities to the involved HTTP software vendors and obtained 7 new CVEs from well-known HTTP software, including Apache, Tomcat, Weblogic, and Microsoft IIS Server. Kaiwen Shen, Jianyu Lu, Jianjun Chen 0005, Mingming Zhang 0010, Hai-Xin Duan, Jia Zhang 0004 |
DSN | 7 |
| 2022 | ValCAT: Variable-Length Contextualized Adversarial Transformations Using Encoder-Decoder Language ModelabstractChuyun Deng, Mingxuan Liu, Yue Qin, Jia Zhang, Hai-Xin Duan, Donghong Sun. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Chuyun Deng, Mingxuan Liu 0006, Jia Zhang 0004, Hai-Xin Duan, Donghong Sun |
NAACL-HLT | 4 |
| 2022 | Encrypted Malware Traffic Detection via Graph-based Network AnalysisabstractMalicious activities on the Internet continue to grow in volume and damage, posing a serious risk to society. Malware with remote control capabilities is considered one of the most threatening malicious activities, as it can enable arbitrary types of cyber-attacks. As a countermeasure, many malware detection methods are proposed to identify malicious behaviours based on traffic characteristics. However, the emerging encryption and evasion techniques pose substantial barriers to the full exploitation of network information. This significantly impairs the effectiveness of existing malware detection methods relying on a singular type of characteristics. In this paper, we propose ST-Graph to resolve this issue. In addition to traditional stream attributes, ST-Graph explores spatial and temporal characteristics of network behaviours based on a graph representation learning algorithm and integrates all available information to boost the detection decision. To illustrate the effectiveness of ST-Graph, we evaluate it on two datasets. Experimental results demonstrate that ST-Graph outperforms state-of-the-art malware detection systems and also shows good performance in efficiency, generalizability, and robustness. Specifically, it achieves over 99% precision and recall, and its False Positive Rate is even two orders of magnitude lower than (nearly 0.02 times) that of baseline models. Meanwhile, the deployment of ST-Graph in two real network scenarios for around one year shows an outstanding efficiency with only 160 seconds time cost for 5-minute traffic in 1.7 Gbps bandwidth. Zhuoqun Fu, Mingxuan Liu 0006, Jia Zhang 0004, Yuan Zou, Qilei Yin, Qi Li 0002, Hai-Xin Duan |
RAID | 4 |
| 2021 | Mingling of Clear and Muddy Water: Understanding and Detecting Semantic Confusion in Blackhat SEO
Kun Du, Yubao Zhang, Shuai Hao 0001, Haining Wang 0001, Jia Zhang 0004, Hai-Xin Duan |
ESORICS (1) | 6 |
| 2021 | On Evaluating Delegated Digital Signing of Broadcasting Messages in 5GabstractIn 5G networks, base stations, namely gNBs (5G NodeB, as per 3GPP nomenclature) periodically broadcast the system information messages including network identifiers to facilitate User Equipment (UE) to connect to the network. As in prior generations, the system information messages in 5G are transmitted in clear text without any security protection. Therefore, an adversary could spoof a legitimate gNB to become a man-on-the-side (MOTS) or man-in-the-middle (MITM) attacker. This vulnerability is being studied by 3GPP and a number of solutions have been proposed in the Technical Report (TR 33.809), including a promising solution namely Digital Signing Network Function (DSnF). In this paper, we provided an evaluation of DSnF, including the practicality of its assumption, feasibility of its certificate trans-mission within the system information message, and quantitative analysis of its performance. Our evaluation results show that DSnF is practical in general. Initial results from this paper have been provided to 3GPP and incorporated into TR 33.809. Yiming Zhang 0009, Tao Wan 0004, Jia Zhang 0004, Hai-Xin Duan |
GLOBECOM | 4 |
| 2020 | Light Multi-Segment Activation for Model CompressionabstractModel compression has become necessary when applying neural networks (NN) into many real application tasks that can accept slightly-reduced model accuracy but with strict tolerance to model complexity. Recently, Knowledge Distillation, which distills the knowledge from well-trained and highly complex teacher model into a compact student model, has been widely used for model compression. However, under the strict requirement on the resource cost, it is quite challenging to make student model achieve comparable performance with the teacher one, essentially due to the drastically-reduced expressiveness ability of the compact student model. Inspired by the nature of the expressiveness ability in NN, we propose to use multi-segment activation, which can significantly improve the expressiveness ability with very little cost, in the compact student model. Specifically, we propose a highly efficient multi-segment activation, called Light Multi-segment Activation (LMA), which can rapidly produce multiple linear regions with very few parameters by leveraging the statistical information. With using LMA, the compact student model is capable of achieving much better performance effectively and efficiently, than the ReLU-equipped one with same model complexity. Furthermore, the proposed method is compatible with other model compression techniques, such as quantization, which means they can be used jointly for better compression performance. Experiments on state-of-the-art NN architectures over the real-world tasks demonstrate the effectiveness and extensibility of the LMA. Zhenhui Xu, Guolin Ke, Jia Zhang 0004, Jiang Bian 0002, Tie-Yan Liu |
AAAI | 3 |
| 2020 | CDN Backfired: Amplification Attacks Based on HTTP Range RequestsabstractContent Delivery Networks (CDNs) aim to improve network performance and protect against web attack traffic for their hosting websites. And the HTTP range request mechanism is majorly designed to reduce unnecessary network transmission. However, we find the specifications failed to consider the security risks introduced when CDNs meet range requests. In this study, we present a novel class of HTTP amplification attack, Range-based Amplification (RangeAmp) Attacks. It allows attackers to massively exhaust not only the outgoing bandwidth of the origin servers deployed behind CDNs but also the bandwidth of CDN surrogate nodes. We examined the RangeAmp attacks on 13 popular CDNs to evaluate the feasibility and real-world impacts. Our experiment results show that all these CDNs are affected by the RangeAmp attacks. We also disclosed all security issues to affected CDN vendors and already received positive feedback from 12 vendors. Kaiwen Shen, Run Guo, Baojun Liu 0002, Jia Zhang 0004, Hai-Xin Duan, Shuang Hao 0001, Xiarun Chen |
DSN | 5 |
| 2020 | CDN Judo: Breaking the CDN DoS Protection with Itself
Run Guo, Baojun Liu 0002, Shuang Hao 0001, Jia Zhang 0004, Hai-Xin Duan, Kaiwen Shen, Jianjun Chen 0005, Ying Liu 0024 |
NDSS | 5 |
| 2019 | DeepGBM: A Deep Learning Framework Distilled by GBDT for Online Prediction TasksabstractOnline prediction has become one of the most essential tasks in many real-world applications. Two main characteristics of typical online prediction tasks include tabular input space and online data generation. Specifically, tabular input space indicates the existence of both sparse categorical features and dense numerical ones, while online data generation implies continuous task-generated data with potentially dynamic distribution. Consequently, effective learning with tabular input space as well as fast adaption to online data generation become two vital challenges for obtaining the online prediction model. Although Gradient Boosting Decision Tree (GBDT) and Neural Network (NN) have been widely used in practice, either of them yields their own weaknesses. Particularly, GBDT can hardly be adapted to dynamic online data generation, and it tends to be ineffective when facing sparse categorical features; NN, on the other hand, is quite difficult to achieve satisfactory performance when facing dense numerical features. In this paper, we propose a new learning framework, DeepGBM, which integrates the advantages of the both NN and GBDT by using two corresponding NN components: (1) CatNN, focusing on handling sparse categorical features. (2) GBDT2NN, focusing on dense numerical features with distilled knowledge from GBDT. Powered by these two components, DeepGBM can leverage both categorical and numerical features while retaining the ability of efficient online update. Comprehensive experiments on a variety of publicly available datasets have demonstrated that DeepGBM can outperform other well-recognized baselines in various online prediction tasks. Guolin Ke, Zhenhui Xu, Jia Zhang 0004, Jiang Bian 0002, Tie-Yan Liu |
KDD | 3 |
| 2019 | Finding the best answer: measuring the optimization of public and authoritative DNS
Jia Zhang 0004, Hai-Xin Duan, Jian Jiang 0002, Jinjin Liang |
Sci. China Inf. Sci. | 1 |
| 2018 | Analysis and Measurement of Zone Dependency in the Domain Name SystemabstractThe Domain Name System (DNS) is a hierarchical distributed system organized through top-down zone delegation. Consequently resolution of a zone depends on its ancestors. However, since the delegation in DNS is designed by name rather than address, the dependency could further extend to other zones. If not configured well, the dependency of a zone could be large and complicated, potentially harmful to its availability and integrity. In this paper, we propose a graph-based model to comprehensively analyze zone dependency in DNS. Our approach classifies zone dependency into four different relations: general dependency, explicit dependency, critical dependency and essential dependency. We also propose an empirical method to quantitatively measure the zone dependencies of given zones. Our survey with over 1 million DNS zones shows that more than 99% of the zones depend on some 3-rd party zone; about 41% of the zones critically rely on more than 2 zones except their ancestors; some TLDs such as .org, .info and .cn tend to have more dependencies than others. Jian Jiang 0002, Jia Zhang 0004, Hai-Xin Duan, Kang Li 0001 |
ICC | 2 |
| 2018 | Abusing CDNs for Fun and Profit: Security Issues in CDNs' Origin ValidationabstractContent Delivery Networks (CDNs) are critical Internet infrastructure. Besides high availability and high performance, CDNs also provide security services such as anti-DoS and Web Application Firewalls to CDN-powered websites. However, the massive resources of CDNs may also be leveraged by attackers exploiting their architectural, implementation, or operational weaknesses. In this paper, we show that today's CDN operation is overly loose in customer-controlled forwarding policy and the lack of origin validation leads to a wide range of abuse cases such as DoS attack and stealthy port scan. We systematically study these abuse cases and demonstrate their feasibility in popular CDNs. Further, we evaluate the impact of these abuses by discovering that there are millions of CDN edge servers, and a substantial fraction of them can be abused. Lastly, we propose mitigation solutions against such abuses and discuss their feasibility. Run Guo, Jianjun Chen 0005, Baojun Liu 0002, Jia Zhang 0004, Chao Zhang 0008, Hai-Xin Duan, Tao Wan 0004, Jian Jiang 0002, Shuang Hao 0001, Yaoqi Jia |
SRDS | 4 |
| 2017 | Randomized Mechanisms for Selling Reserved Instances in Cloud ComputingabstractSelling reserved instances (or virtual machines) is a basic service in cloud computing. In this paper, we consider a more flexible pricing model for instance reservation, in which a customer can propose the time length and number of resources of her request, while in today's industry, customers can only choose from several predefined reservation packages. Under this model, we design randomized mechanisms for customers coming online to optimize social welfare and providers' revenue. We first consider a simple case, where the requests from the customers do not vary too much in terms of both length and value density. We design a randomized mechanism that achieves a competitive ratio 1/42 for both social welfare and revenue, which is a improvement as there is usually no revenue guarantee in previous works such as (Azar et al. 2015; Wang et al. 2015. This ratio can be improved up to 1/11 when we impose a realistic constraint on the maximum number of resources used by each request. On the hardness side, we show an upper bound 1/3 on competitive ratio for any randomized mechanism.We then extend our mechanism to the general case and achieve a competitive ratio 1/42⌈log k⌉ log T for both social welfare and revenue, where T is the ratio of the maximum request length to the minimum request length and k is the ratio of the maximum request value density to the minimum request value density. This result outperforms the previous upper bound 1/CkT for deterministic mechanisms (Wang et al. 2015). We also prove an upper bound 2/log 8kT for any randomized mechanism. All the mechanisms we provide are in a greedy style. They are truthful and easy to be integrated into practical cloud systems. Jia Zhang 0004, Weidong Ma, Tao Qin 0001, Xiaoming Sun 0001, Tie-Yan Liu |
AAAI | 1 |
| 2017 | Efficient Delivery Policy to Minimize User Traffic Consumption in Guaranteed AdvertisingabstractIn this work, we study the guaranteed delivery model which is widely used in online advertising. In the guaranteed delivery scenario, ad exposures (which are also called impressions in some works) to users are guaranteed by contracts signed in advance between advertisers and publishers. A crucial problem for the advertising platform is how to fully utilize the valuable user traffic to generate as much as possible revenue. Different from previous works which usually minimize the penalty of unsatisfied contracts and some other cost (e.g. representativeness), we propose the novel consumption minimization model, in which the primary objective is to minimize the user traffic consumed to satisfy all contracts. Under this model, we develop a near optimal method to deliver ads for users. The main advantage of our method lies in that it consumes nearly as least as possible user traffic to satisfy all contracts, therefore more contracts can be accepted to produce more revenue. It also enables the publishers to estimate how much user traffic is redundant or short so that they can sell or buy this part of traffic in bulk in the exchange market. Furthermore, it is robust with regard to priori knowledge of user type distribution. Finally, the simulation shows that our method outperforms the traditional state-of-the-art methods. Jia Zhang 0004, Qian Li 0012, Jialin Zhang 0001, Yanyan Lan, Qiang Li 0043, Xiaoming Sun 0001 |
AAAI | 1 |
| 2017 | How to Notify a Vulnerability to the Right Person? Case Study: In an ISP ScopeabstractHow to inform the right person is an important step in the network security incident response. In previous studies, researchers focused on email notification mode in the whole Internet, and main objective is to find an effective mode to notify the ISP or the related institutions. In this paper, we extend previous research in an ISP scope. We plan to analyze factors which can affect the effectiveness of notification in an ISP scope and try to find some reasonable vulnerability notification modes for ISP to use. We use a Chinese ISP CERNET as a research case, and identify three different types of vulnerabilities in the customers of CERNET and notify them through three different notification methods(Customer Service Phone, Email and Instant Messenger). Then we analyze all the feedbacks of customers, and study the effectiveness of each notification method. Through the study we find that the customer pays more attention to the high-risk vulnerability, while other potential risks have not been given adequate attention. We also find that the current vulnerability notification mode for ISP is not perfect. IM (Instant Messenger) is the most effective way to notify vulnerability, but it is not commonly used. For different types of vulnerabilities, the remediation ratio may be related to the role who we should notify: the vulnerabilities with more complexity need to notify the person with higher level technical capability, while the vulnerabilities which are related to application system need to notify the person who has the authority to fix it. If we do not find the right contact, repeated notification is useless. At last, for the effectiveness of notification, we propose to establish an IM group in an ISP scope with the participation of network operation directors, security operation experts and system administrators etc. of each customer. Jia Zhang 0004, Hai-Xin Duan, Xingkun Yao |
GLOBECOM | 1 |
| 2016 | Computing the least-core and nucleolus for threshold cardinality matching games
Qizhi Fang, Bo Li 0037, Xiaoming Sun 0001, Jia Zhang 0004, Jialin Zhang 0001 |
Theor. Comput. Sci. | 4 |
| 2014 | Solving Multi-choice Secretary Problem in Parallel: An Optimal Observation-Selection Protocol
Xiaoming Sun 0001, Jia Zhang 0004, Jialin Zhang 0001 |
ISAAC | 2 |
| 2014 | Computing the Least-Core and Nucleolus for Threshold Cardinality Matching Games
Qizhi Fang, Bo Li 0037, Xiaoming Sun 0001, Jia Zhang 0004, Jialin Zhang 0001 |
WINE | 4 |
| 2011 | Anonymity analysis of P2P anonymous communication systems
Jia Zhang 0004, Hai-Xin Duan |
Comput. Commun. | 1 |
| 2008 | AMCAS: An Automatic Malicious Code Analysis SystemabstractWith the development of malicious code technology, the number of malicious code has continued to increase. So it is imperative to optimize the traditional manual analysis method by automatic malicious code analysis system. This paper presents AMCAS - an automatic malicious code analysis system. It includes malicious code static analyzer, dynamic analyzer and network behavior analyzer. Compared with some existing automatic analysis systems, this system integrates the advantages of static and dynamic analysis, and imports network behavior analysis. Static analyzer can get the unpacked binary code and CallGraph; dynamic analyzer can get the host behavior of malicious code and network behavior analyzer can get the malicious network behavior profile. Experiment shows that this system can get malicious code information efficiently. Jia Zhang 0004, Yuntao Guan, Xiaoxin Jiang, Hai-Xin Duan |
WAIM | 1 |