Yujia Zhu

dblp:80/4063 · DBLP profile ↗
← Back
44ranked-venue papers
8as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 13 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 5 since 2021Security and privacy · 9 · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 STAR: Semantic-Traffic Alignment and Retrieval for Zero-Shot HTTPS Website Fingerprinting
Yujia Zhu, Baiyang Li, Xinhao Deng 0001, Yitong Cai, Yaochen Ren, Qingyun Liu 0001
INFOCOM2
2025 Unraveling DoH Traces: Padding-Resilient Website Fingerprinting via HTTP/2 Key Frame Sequences
Baiyang Li, Yujia Zhu, Qingyun Liu 0001, Li Guo 0001
ESORICS (3)2
2025 Knocking on IP: Unveiling Websites through Cache-Aware Fingerprinting
abstract
As user privacy becomes increasingly critical in the digital landscape, traditional methods of website fingerprinting (WF) face significant challenges, particularly in caching scenarios. Existing WF studies are limited by the assumption of disabled caching. Recently, only a few have explored how to address the potential problem of model performance degradation in caching scenarios, but what they propose have limitations in terms of transferability and practicality in existing conditions. However, we propose a deep learning method based solely on IP fingerprints, which is more resilient to network fluctuations than traditional features and is able to extract effective local features from the fluctuating sequences. Furthermore, when tested on simulated cached datasets, our method achieves a remarkable stability with an accuracy of 96.10%. These findings underscore our model’s suitability for identifying websites amidst mixed-length sequences in caching contexts. The solution we provide not only improves the accuracy of identification, but also reveals the impact of caching mechanisms on privacy. Meanwhile, our comprehensive dataset and innovative methods pave the way for future research in user privacy and WF.
Yujia Zhu, Xiaoou Zhang
ICASSP3
2025 Node-Centric Meta Structure Search in Heterogeneous Graphs
abstract
Heterogeneous graphs are increasingly used to represent complex real-world scenarios with diverse entities and interactions by meta structures. Recently, the search of meta structures is combined with graph neural architecture search to automatically extract the semantic knowledge for various tasks in heterogeneous graphs. However, prior research primarily focuses on identifying meta structures that are universally applicable across all nodes in a graph, neglecting the variations in meta structure selection that arise from the unique features and topology of individual nodes. To address this challenge, we introduce a Node-Centric approach to search Meta Structures in heterogeneous graphs (NC-MS for short). NC-MS implements a node level method that discover meaningful meta structures tailored to each node, capturing subtle differences in meta structure choices between nodes and providing nuanced identification. Additionally, NC-MS utilizes an efficient and differentiable network to enhance operational efficiency. Empirical studies across three real-world datasets validate the superiority of NC-MS, demonstrating its ability to outperform existing models in heterogeneous graph neural networks.
Xiaoou Zhang, Yang Gao 0024, Yang Aron Liu, Yujia Zhu, Chuan Zhou 0001, Peng Zhang 0001, Qingyun Liu 0001, Hongyang Chen 0001
ICASSP4
2025 Heterogeneous Graph Anomaly Detection with Graph Wavelet Transformer
abstract
Graph Anomaly Detection (GAD) identifies deviant patterns including anomalous nodes, edges, and subgraphs in graph data, with significant applications in social networks, cybersecurity, and financial risk control. While spectral methods have proven effective for homogeneous graph anomaly detection, their application to heterogeneous graphs remains challenging due to structural complexity and semantic richness. Existing heterogeneous graph anomaly detection methods either rely on manually designed meta-paths or decompose the graph into homogeneous subgraphs, leading to limited flexibility or loss of structural integrity. To address these limitations, we propose the Graph Wavelet Transformer (GWT), a novel spectral-based approach that integrates global graph properties and spectral analysis without requiring meta-path information. GWT employs a three-stage process: heterogeneous-to-homogeneous graph conversion, global dependency modeling via graph transformers, and spectral-aware feature enhancement focusing on frequency band components. Extensive experiments on multiple benchmarks demonstrate that GWT significantly outperforms ten baseline methods, providing a new paradigm for heterogeneous graph anomaly detection that preserves structural completeness while achieving computational efficiency.
Xiaoou Zhang, Chuan Zhou 0001, Yang Aron Liu, Shuai Zhang 0007, Peng Zhang 0001, Yujia Zhu, Qingyun Liu 0001
ICDM6
2025 Lightweight Relational Proposal Network with Dual-Branch Distillation for Video Moment Retrieval
Yujia Zhu, Hao Yang 0067, Yibo Zhao 0001, Chunjie Ma, Weili Guan, Zan Gao 0001
ACM Multimedia1
2025 Broken Chains: An Empirical Analysis of DNS Resolution in IPv6-only Environments
abstract
The global transition to IPv6 is impeded by failures within the Domain Name System (DNS), where the mere presence of an AAAA record does not guarantee a domain’s resolvability in IPv6-only environments. In this paper, we presents a comprehensive measurement study analyzing the complete DNS dependency chain—including parental, delegation, and alias dependencies—to reveal the true state of IPv6 resolvability. We introduce 6ChainChecker, a lightweight tool developed for this analysis. Our analysis of Tranco top domains reveals a critical discrepancy: 7.68% of domains with published AAAA records are nevertheless unresolvable from a strict IPv6-only stack due to structural failures in their dependency chains. To understand the broader landscape of failures, we find that while the absence of an AAAA record is the most common reason for unresolvability (58.0%), a substantial portion of failures stem from broken dependency paths, including delegation (17.1%) and alias chain (24.8%) issues. Crucially, we also identify significant dependency concentration, where the non-compliance of a few critical infrastructure zones creates cascading failures for hundreds of their dependent domains. These findings demonstrate that upstream infrastructure, not just endpoint configuration, is a significant impediment to the IPv6 transition. Our tool and dataset are publicly available to foster further research.
Yujia Zhu, Baiyang Li, Qingyun Liu 0001
TrustCom2
2025 HOLMES & WATSON: A Robust and Lightweight HTTPS Website Fingerprinting through HTTP Version Parallelism
abstract
Website Fingerprinting (WF) is a traffic analysis technique that aims to identify websites visited by users through the analysis of encrypted traffic patterns.Existing approaches often exhibit limited robustness against network variability and concept drift, resulting in significant performance degradation under real-world HTTPS conditions.Moreover, these methods typically require large-scale training datasets and substantial computational resources, which further increases the complexity of deployment.In this paper, we propose HOLMES, a novel approach that exploits HTTP version parallelism to extract enhanced application-layer features.These features, including the number of web resources transmitting in various HTTP versions, expose up to 4.28 bits of information-surpassing 98% of previously reported features and demonstrate increased stability across varying network conditions.Complementary to this, we introduce WATSON, a lightweight classification method based on lazy learning, which substantially reduces the dependency on large training datasets.To further enhance the identification accuracy, we incorporate two fingerprint-specific distance metrics that ensure high intra-class similarity.Our experimental evaluation demonstrates that HOLMES & WATSON significantly enhance both robustness and efficiency, achieving an average accuracy of 87.7% with only a single sample per website, marking an improvement of over 15% compared to state-of-the-art methods.
Yujia Zhu, Baiyang Li, Peishuai Sun, Xinhao Deng 0001, Qingyun Liu 0001
WWW2
2025 DMC-Watermark: A backdoor richer watermark for dual identity verification by dynamic mask covering
Yujia Zhu, Daoxun Xia
Appl. Intell.1
2025 Cascade Ownership Verification Framework Based on Invisible Watermark for Model Copyright Protection
abstract
ABSTRACT Successfully training a model requires substantial computational power, excellent model design, and high training costs, which implies that a well‐trained model holds significant commercial value. Protecting a trained Deep Neural Network (DNN) model from Intellectual Property (IP) infringement has become a matter of intense concern recently. Particularly, embedding and verifying watermarks in black‐box models without accessing internal model parameters, while ensuring the robustness and invisibility of the watermark, remains a challenging issue. Unlike many existing methods, we propose a cascade ownership verification framework based on invisible watermarks, with a focus on how to effectively protect the copyright of black‐box watermark models and detect unauthorized users' infringement behaviors. This framework consists of two parts: watermark generation and copyright verification. In the watermark generation phase, watermarked samples are generated from key samples and label images. The difference between watermarked samples and key samples is imperceptible, while a specific identifier has been injected into the watermarked samples, leaving a backdoor as an entry point for copyright verification. The copyright verification phase employs hypothesis testing to enhance the confidence level of verification. In image classification tasks based on MNIST, CIFAR‐10, and CIFAR‐100 datasets, experiments were conducted on several popular deep learning models. The experimental results show that this framework offers high security and effectiveness in protecting model copyrights and demonstrates strong robustness against pruning and fine‐tuning attacks.
Yujia Zhu, Daoxun Xia
Concurr. Comput. Pract. Exp.2
2025 Multiuser Hierarchical Authorization Using Sparsity Polarization Pruning for Model Active Protection
abstract
ABSTRACT Currently, artificial intelligence technology is rapidly penetrating into various fields of socioeconomic development with increasing depth and breadth, becoming an important force driving innovation and development, empowering thousands of industries, while also bringing challenges such as security governance. The application of deep neural network models must implement hierarchical access based on user permissions to prevent unauthorized users from accessing and abusing the model, and to prevent malicious attackers from tampering or damaging the model, thereby reducing its vulnerabilities and security risks. To address this issue, the model provider must implement a hierarchical authorization policy for the model, which can grant users access to the model based on their specific needs, while ensuring that unauthorized users cannot use the model. Common methods for implementing hierarchical authorization of models include pruning and encryption, but existing technologies require high computational complexity and have unclear hierarchical effects. In this article, we propose a sparsity polarization pruning approach for layered authorization, which combines sparsity regularization to filter insignificant channels and a polarization technique to cluster critical channels into distinct intervals. By pruning channels based on polarized scaling factors from the batch normalization (BN) layer, our method dynamically adjusts model precision to match user authorization levels. Initially, we extract the scaling factor of the BN layer to assess the importance of each channel. A sparsity regularizer is then applied to filter out irrelevant scaling factors. To enhance the clarity and rationality of pruning intervals, we use a polarization technique to induce clustering of scaling factors. So we proposed multiuser hierarchical authorization using sparsity polarization pruning for model active protection. Based on the grading requirements, we prune channels corresponding to varying numbers of significant scaling factors. Access is granted at different levels depending on the precision key provided by the user, thereby ensuring a secure and efficient means of accessing the model's resources. Experimental results demonstrate that our approach achieves superior grading performance across three datasets and two different neural networks, showcasing its broad applicability. Moreover, our method achieves effective grading just by pruning a small portion of the channels, offering a high level of efficiency.
Yujia Zhu, Xiaojie Du, Daoxun Xia
Concurr. Comput. Pract. Exp.1
2024 6GAI: Active IPv6 Address Generation via Adversarial Training with Leaked Information
abstract
Global IPv6 scanning has always been a challenge for researchers because of the limited network speed and computational power. In this paper, we introduce 6GAI to implement more efficient target address generation. 6GAI is built with Generative Adversarial Net (GAN) integrated with Convolutional Bottleneck Attention Module (CBAM). 6GAI allows the discriminative net to leak generated address’s high-level features extracted by CBAM to the generative net, while the generative net incorporates such informative signals into all generation steps through an additional Manger module, which takes the extracted features of current generated address nybbles and outputs a latent vector to guide the Worker module for active IPv6 address generation. This work outperformed the state-of-the-art target generation algorithms on two datasets including one public dataset and one independently collected dataset.
Liang Jiao, Yujia Zhu, Wen-Xiu Zhang, Qingyun Liu 0001
CSCWD2
2024 From Fingerprint to Footprint: Characterizing the Dependencies in Encrypted DNS Infrastructures
Baiyang Li, Yujia Zhu, Qingyun Liu 0001, Li Guo 0001
ESORICS (2)2
2024 Meta Structure Search for Link Weight Prediction in Heterogeneous Graphs
abstract
Recently link weight prediction has attracted an increasing research interest due to its merits in quantifying the strength between nodes within a graph. Nonetheless, current link weight prediction methods focus solely on graph topology, disregarding node feature information embedded in graphs. In real-world applications, we often collect heterogeneous graph data where multiple types of nodes linked by multiple types of edges are available for analysis, and it is essential and challenging to quantify the proximity of different types of nodes. To solve this challenge, we present a new model for Heterogeneous Graph Link Weight Prediction (HLWP for short). In HLWP, message passing in heterogeneous graph neural networks is described as a meta structure, which can be effectively designed by Differentiable Neural Architecture Search (DARTS) algorithms. Thus, HLWP can enhance the message passing in heterogeneous graphs by DARTS. In addition, HLWP employs a perturbation-based algorithm to enhance stability and precision. Through empirical experiments conducted on three real-world datasets, we demonstrate that HLWP achieves accurate predictions of link weights. Our results highlight the superiority of HLWP over existing methods for link weight prediction and baseline GNN models in terms of accurately predicting link weights within heterogeneous graphs.
Xiaoou Zhang, Yang Gao 0024, Yang Aron Liu, Yujia Zhu, Peng Zhang 0001, Chuan Zhou 0001, Qingyun Liu 0001, Hongyang Chen 0001
ICASSP4
2024 Failed Yet Stored: A First Look at DNS Negative Caching
abstract
Caching is a critical method for enhancing the efficiency and the security of the Domain Name System (DNS). Initially, only successful domain name resolution results were cached. To mitigate failures in DNS transactions (e.g. NXDomain), the IETF proposed standards, further developed into RFC 9520 as of December 2023. In addition to the basic implementation, RFC 9520 standardizes more sophisticated forms of negative caching. This new standard aims to reduce redundant query retries in DNS traffic and protect resolvers from Denial of Service (DoS) attacks.In this study, we present a comprehensive examination of the specific implementations of Negative Caching in resolvers. We designed and validated a method for measuring negative caching and conducted experiments on 44 public resolvers, including their Do53, DoH, and DoT interfaces. Our findings indicate that while public resolvers generally implement various types of negative caching, some exhibit unexpected cache handling behaviors when encountering specific negative responses. Additionally, we discovered that most public resolvers modify the TTL value of negative responses before returning them to clients. Despite the lack of explicit TTL values for newly specified negative responses, we devised a method to approximate the default TTL values used by public resolvers.
Meng Zeng, Yujia Zhu, Baiyang Li, Qingyun Liu 0001, Binxing Fang
IPCCC2
2024 Measuring Encrypted DNS Service with TLS1.3 Support over IPv6
abstract
The Encrypted Domain Name System (DNS) and Encrypted Server Name Indication (ESNI) are recently proposed to enhance network security and privacy protection; we refer to these schemes collectively as domain name encryption technologies. Previous research has shown that the destination IP address accessed by the user cannot be associated with common web services such as websites because a large number of websites are hosted through cloud or CDN over IPv6. However, encrypted DNS, as an internet infrastructure service, is typically deployed independently by the service provider rather than hosted through cloud or CDN. In this paper, we propose a method to discover the unique service provider of encrypted DNS resolvers on a large-scale encrypted traffic with TLS1.3 support over IPv6. The model utilizes a Siamese Network to determine whether two IPv6 resolver addresses belong to the same service provider of encrypted DNS, even if the DNS query is protected by ESNI. Through a comprehensive analysis of two real-world datasets, which include encrypted DNS data and common web data, we find that the implementation of TLS1.3, especially ESNI, does not impact the association of encrypted DNS server addresses. Our model achieves an accuracy rate of 95.29%.
Liang Jiao, Wen-Xiu Zhang, Tianyu Cui, Yujia Zhu, Qingyun Liu 0001
ISCC6
2024 Backdoor Richer Watermarks Using Dynamic Mask Covering for Dual Identity Verification
Yujia Zhu, Daoxun Xia
PRCV (4)1
2024 A Comprehensive Evaluation of the Impact on Tor Network Anonymity Caused by ShadowBridge
Baiwei Duan, Yujia Zhu, Can Zhao 0005, Jinqiao Shi
SecureComm (4)4
2024 LayyerX: Unveiling the Hidden Layers of DoH Server via Differential Fingerprinting
abstract
As a rapidly developing DNS security enhancement technology, DoH(DNS over HTTPS) is gaining popularity among people. It allows users to quickly set up a DoH server by combining several components which create a layered structure. However, the multi-layer setup, which involves both HTTPS and DNS protocols, makes internal structural details more difficult to be detected. To address this issue, we propose a method that utilizes cross-protocol fingerprinting and analysis techniques, which is capable of identifying various components of multi-layer DoH servers with a focus on the underlying differences within protocol. Using this approach, we developed LayyerX, a system for detecting multi-layer DoH servers. Finally, through experiments and large-scale measurements in the wild, we showcased LayyerX’s outstanding capabilities and presented a meaningful structural overview of multi-layer DoH server.
Yunyang Qin, Yujia Zhu, Linkang Zhang, Baiyang Li, Qingyun Liu 0001
TrustCom2
2023 DualDNSMiner: A Dual-Stack Resolver Discovery Method Based on Alias Resolution
Dingkang Han, Yujia Zhu, Liang Jiao, Dikai Mo, Qingyun Liu 0001
CollaborateCom (3)2
2023 SpoofingGuard: A Content-agnostic Framework for Email Spoofing Detection via Delivery Graph
abstract
Email spoofing is an effective attack vector for infiltrating companies and organizations. Traditional detectors are primarily based on the content of emails, but they ignore the frequent contextual changes. The blacklist-based solutions commonly used in the industry suffers from latency issues. Additionally, there are protocol-based solutions, such as SPF, DKIM, etc., but their adoption rates are unsatisfactory. To address these issues, this work presents a new framework named SpoofingGuard that detects email spoofing based on graph representation learning. As SpoofingGuard extracts important delivery path information related to the email service infrastructure from email headers, it is completely content-agnostic, and is expected to be more robust in the face of complex content variations. Finally, the evaluation results on two public datasets show that SpoofingGuard can achieve 99.51% precision and less than 0.5% false positive rate, demonstrating its effectiveness and advancement.
Yujia Zhu, Xiaoou Zhang, Zhen Jie, Qingyun Liu 0001
CSCWD2
2023 Unveiling Flawed Cache Structures in DNS Infrastructure via Record Watermarking
abstract
The Domain Name System (DNS) is an essential component of the internet, providing name resolution services to navigate clients to various resources on the network. Caches are critical to the efficient operation of DNS, both in terms of service quality and security. Major service providers maintain complex DNS infrastructure with multiple cache layers to handle client queries. Unfortunately, access to these caches is not available, making it difficult to understand the cache structure. In this study, we propose methodologies for identifying hidden cache structures in DNS infrastructure using watermark records. We further applied our methods to conduct global measurements, utilizing open resolvers as vantage points. Our measurement results indicated that flawed DNS cache structures exist, leaving the DNS infrastructure vulnerable and inefficient. Thousands of client networks suffer from severe cache fragmentation. Furthermore, a large number of recursive resolvers rely on fragile or poor cache structures.
Dikai Mo, Yujia Zhu, Zhen Jie, Qingyun Liu 0001, Binxing Fang
GLOBECOM2
2023 A Time Series Clustering Method for Network Big Data
abstract
In the era of big data, network data increase rapidly in a distributed manner, giving birth to the network big data. Network big data with the extra features such as distributed and decentralized data collection and storage, distributed and parallel data processing, more complex and evolving relationships among data, and heterogeneous data representation, pose opportunities together with challenges to the traditional network analysis algorithms. A hopeful solution is combining machine learning techniques with network big data analysis. In this paper, we proposed a novel time series clustering method which effectively combines machine learning techniques with network big data analysis for fault diagnosis task based on network logs. Verification experimental results on classic HDFS dataset demonstrate the outstanding performance of the proposed method.
Yujia Zhu, Geyong Min, Yulei Wu, Haozhe Wang 0001
ICPADS1
2023 DRDoSHunter: A Novel Approach Based on FDA and Inter-flow Features for DRDoS Detection
abstract
In recent years, Distributed Reflective Denial of Service (DRDoS) attacks have emerged as a major threat to network security, utilizing IP spoofing and amplification mechanisms to drain network bandwidth. Existing approaches for DRDoS detection lack sophistication in feature selection and focus primarily on detection rather than fine-grained classification and targeted mitigation. In this paper, we propose DRDoSHunter, a novel approach that addresses these limitations. DRDoSHunter employs Frequency Domain Analysis (FDA) and inter-flow features to extract effective and robust features from continuous time series data. By utilizing a deep residual network model, our approach achieves accurate and efficient classification of DRDoS attacks at a fine-grained level. Experimental results on the CIC-DDoS2019 public dataset demonstrate that DRDoSHunter outperforms popular detection models, achieving an Fl-Score of over 98.44% for DRDoS attack detection and classification.
Yujia Zhu, Jiang Xie 0004, Yitong Cai
ISCC2
2023 CCSv6: A Detection Model for DNS-over-HTTPS Tunnel Using Attention Mechanism over IPv6
abstract
In this paper, we first show DNS-over-HTTPS (DoH) tunneling detection methods verified to be effective over IPv4 can be applied to IPv6, and then propose a new model called CCSv6, using attention-based convolution neural network to build classifiers with flow-based features to detect DoH tunneling over IPv6, achieve 99.99% accuracy on the IPv6 dataset. In addition, we discuss the influence of various factors such as locations or DoH resolvers on the detection results in detail over IPv6. All the more important, our model shows better transfer learning ability, which can achieve the F1-score of 96% when trained on the IPv6 dataset and tested on the IPv4 dataset.
Liang Jiao, Yujia Zhu, Fenglin Qin, Qingyun Liu 0001
ISCC2
2023 Before Toasters Rise Up: A View into the Emerging DoH Resolver's Deployment Risk
abstract
As an encryption protocol for DNS queries, DNS-over-HTTPS (DoH) is becoming increasingly popular, and it mainly addresses the last-mile privacy protection problem. However, the security of DoH is in urgent need of measurement and analysis due to its reliance on certificates and upstream servers. In this paper, we focus on the DoH ecosystem and conduct a one-month measurement to analyze the current deployment of DoH resolvers. Our findings indicate that some of these resolvers use invalid certificates, which can compromise the security and privacy advantages of the protocol. Furthermore, we found that many providers are at risk of certificate outages, which could cause significant disruptions to the DoH ecosystem. Additionally, we observed that the centralization of DoH resolvers and upstream DNS servers is a potential issue that needs addressing to ensure the stability of the ecosystem.
Yuqi Qiu, Baiyang Li, Zhiqian Li, Liang Jiao, Yujia Zhu, Qingyun Liu 0001
ISCC5
2023 Measuring DNS-over-Encryption Performance Over IPv6
abstract
In recent years, encrypted DNS such as DNS-over-HTTPS (DoH) and DNS-over-TLS (DoT) has gained significant traction as privacy-preserving alternative to conventional DNS. While several studies have measured the performance of encrypted DNS relative to conventional DNS, they are only performed over IPv4, little has been done to understand their status over IPv6. Besides, previous studies can not obtain the absolute query latency due to lack of control over vantage points.This paper performs by far the fist end-to-end performance measurements on encrypted DNS over IPv6. By analyzing measurement results, we have gained several insights. In general, the quality of service for encrypted DNS is satisfying. Over IPv6, encrypted DNS performance varies across resolvers, and is affected by the location issuing DNS queries, the type of encrypted DNS protocol used and the latency to resolvers. Compared with IPv4, the performance of encrypted DNS of different resolvers over IPv6 is improved to some extent. In addition, we also find other problems such as the quality of service of resolver Ahadns is significantly low both over IPv6 and IPv4, as well as the performance of encrypted DNS for resolver Alidns significantly deteriorates when switching from IPv4 to IPv6. Based on our observations, we provide recommendations and discuss situations in which switching to IPv6 may be beneficial. We hope that our tools developed for performing measurements can help people in different regions to choose to the right recursive resolver and network environment, and that our findings can contribute to improve IPv6 Internet infrastructure and inform continuing encrypted DNS deployment over IPv6.
Liang Jiao, Yujia Zhu, Baiyang Li, Qingyun Liu 0001
TrustCom2
2023 A framework for deep neural network multiuser authorization based on channel pruning
abstract
Summary Various deep neural network (DNN) model watermarks have been proposed by researchers to verify copyrights for deep neural networks DNN. However, most DNN watermarking methods cannot prevent attackers from stealing and using the model. Unlike many existing approaches, this paper uses a channel pruning algorithm to protect DNN models, which verifies DNN models copyrights but also prevents the illegal use of DNN models. In this work, the pruning threshold or pruning rate is used as the secret key of a DNN model. After the secret key is distributed to multiple users, they prune the DNN model with the secret key, and the pruned and fine‐tuned model is provided to the users. The users can verify ownership of the model according to the pruning accuracy and fine‐tuning accuracy. If the secret key is incorrect, the accuracy of the model after fine‐tuning will be very low, and users will be unable to use the reasoning function of the fine‐tuned model. Based on the CIFAR‐10 and CIFAR‐100 datasets, we conducted experiments on five popular DNN models. The experimental results show that we can authorize multiple users by pruning very few channels in the convolution layers of the DNN model.
Linna Wang, Yunfei Song, Yujia Zhu, Daoxun Xia, Guoquan Han
Concurr. Comput. Pract. Exp.3
2022 Evading Encrypted Traffic Classifiers by Transferable Adversarial Traffic
Hanwu Sun, Chengwei Peng, Yafei Sang, Yongzheng Zhang 0002, Yujia Zhu
CollaborateCom (2)6
2022 Detection of DoH Tunnels with Dual-Tier Classifier
abstract
DNS over HTTPS (DoH) has been deployed to provide confidentiality in the DNS resolution process. However, encryption is a double-edged sword in providing security while increasing the risk of data tunneling attacks. Current approaches for plaintext DNS tunnel detection are disabled. Due to the diversity of tunneling tool variations and the low proportion of tunneled traffic in real situations, detecting malicious behaviors is becoming more and more challenging. In this paper, we propose a novel behavior-based model with Dual-Tier Tunnel Classifier (DTC) for tool-level DoH tunneling detection. The major advantage of DTC is that it can not only capture existing tunneling tools but also explore unknown ones in the wild. In particular, DTC considers data imbalance, which improves robustness of the model in the open environment. Our method has been proven successful in both closed and open scenarios, achieving 99.99 % accuracy in detecting known malicious DoH traffic, 96.93% accuracy in unknown and 95.31 % accuracy in identifying malicious DoH tunnel tools.
Yuqi Qiu, Baiyang Li, Liang Jiao, Yujia Zhu, Qingyun Liu 0001
MSN4
2021 Peek Inside the Encrypted World: Autoencoder-Based Detection of DoH Resolvers
abstract
DNS-over-HTTPS (DoH), as a rising star to improve DNS security and privacy, has developed rapidly in recent years. It mixes with HTTP features, shares ports with other web services and provides API with URI templates. The unique characteristics of DoH, as well as its fast growth, bring both promising prospects and new risks, e.g. botnet communication, name abuse and data exfiltration. It is essential for network operators to learn about adoption and usage of DoH resolvers. Active scanning may be a possible way. However, it is considered to incur significantly additional overhead, which can be inefficient and aggressive. In this paper, we present DOHUNTER, a system for automati-cally discovering DoH resolvers. DOHUNTER: (i) picks DoH flow from miscellaneous HTTPS traffic, (ii)confirms DoH resolvers based on the detected DoH flow, (iii)mines other related DoH resolvers from the known ones. Our real-world experiments demonstrate the effectiveness of DOHUNTER in detecting DoH flow and finding DoH resolvers. Utilizing DOHUNTER, we witness an alarming increase in DoH adoption. Additionally, we also reveal oblivious growing trends of DoH, which may provide advice for both users and network operators.
Jiating Wu, Yujia Zhu, Baiyang Li, Qingyun Liu 0001, Binxing Fang
TrustCom2
2019 Hunting for Invisible SmartCam: Characterizing and Detecting Smart Camera Based on Netflow Analysis
abstract
Nowadays, the rapid growth of cloud computing and IoT enabled services among multiple organizations brings both promising prospects and security & privacy challenges. IP cameras have become a top target for hackers because of their relatively high computing power and throughput. To understand the risks of these threats requires learning about IP cameras-where are they, how many are there? Active scanning is considered to be an effective way, like SHODAN. However, deployment of smart cameras in the network address translation (NAT) environments with dynamic locations is usually desired. To find these Invisible Cameras, CamHunter: (i) introduces three statements of smart cameras when they are online, (ii) concludes the most popular smart cameras in China have very similar communication patterns, (iii) proposes a model to detect smart cameras in a passive way constructed by nineteen feature sets, and (iv) raises alarms for IoT manufacturers. Our real-world experiments demonstrate the effectiveness of CamHunter in finding smart cameras even if they are behind NATs and using encrypted connections like SSL/TLS or private protocols. We argue that CamHunter represents an important view of IoT security and privacy, and it can guide the effort of designing and protecting smart cameras.
Baiyang Li, Yujia Zhu, Qingyun Liu 0001, Zhou Zhou 0007, Li Guo 0001
ICC2
2019 IDNS: A High-Performance Model for Identification of DNS Infrastructures on Large-scale Traffic
abstract
Domain Name System (DNS) is indispensable in a large number of network applications. Identifying DNS infrastructures into different roles hierarchically is highly desired for a variety of purposes such as network management and threat evaluation. However, traditional measurements almost all depend on active scanning without considering dynamic packet-level features of different DNS infrastructures.In this paper, we propose a high-performance model IDNS (Identifying DNS) based on passive measurement. IDNS: (i) extracts single-packet field features (SFF) and multi-packet statistical features (MSF) from DNS traffic, (ii) utilizes an estimation algorithm to calculate MSF for satisfying online processing speed, and (iii) applies several classifiers in Ensemble Learning and Incremental Learning. We perform an extensive evaluation based on a large volume of DNS queries and responses collected from one ISP. The evaluation results demonstrate that the best classifier in Ensemble Learning can reach 90% accuracy rate while the classifier in Incremental Learning can reach 80% with the highest scalability.
Caiyun Huang, Yujia Zhu, Qingyun Liu 0001, Binxing Fang
ISCC2
2018 SASD: A Self-Adaptive Stateful Decompression Architecture
abstract
Due to the increasing threats in the current network environment, many researchers have shifted their interests to network content audit, which combines deep packet inspection and natural language processing. However, the performance of network content audit systems is becoming the bottle-neck because of the demand on processing fast growing compressed traffic. While compressed traffic is often split into multiple out-of-order packets for transmission, stateful decompression ensures that the compressed data are processed in a timely manner without waiting for all the compressed traffic to arrive before decompressing. In the meanwhile, hardware innovations lead to new type of devices being invented, which shows promise to fully handle the offloaded traffic for complex calculations at higher throughput than software-based solutions. We consider both software-based and hardware-based solutions for decompressing traffic from network content audit systems and study the workload. We notice that the performance is data-dependant: hardware-based decompression solutions perform better for longer compressed data than software method. On the contrary, software-based decompressing methods are more preferred for the short content in terms of the processing speed. So there is no one-size-fits-all solution. In this paper, we combine the advantages of hardware and software and propose a novel self-adaptive stateful decompression architecture to support fast decompression in accordance with the traffic status and system state. Experiments on real-world traffic show that our proposed architecture can achieve about three times of the data decompression efficiency, compared to the best pure software and hardware algorithm, which can significantly improve the detection efficiency of many network content audit systems.
Zhou Zhou 0007, Qingyun Liu 0001, Yujia Zhu, Da Li 0002, Li Guo 0001
GLOBECOM4
2018 WDMTI: Wireless Device Manufacturer and Type Identification Using Hierarchical Dirichlet Process
abstract
Wireless devices have been widely adopted across all domains. With the convenience brought by wireless communication technology, increasing number of conventional (wired) devices are evolving to become wireless. However, significant security issues arise with the popularity of wireless devices. To start an attack, the adversary usually performs a network reconnaissance to discover exposed devices, identify device manufacturers and types, and then scan for vulnerabilities. From the defense side, network administrators are expected to identify the potential vulnerabilities/risks and enforce Network Access Control (or Network Admission Control, NAC) on all the connecting devices. To do this, it is essential to accurately identify the make/model/type of each device that attempts to connect to the network, e.g., MacBooks, Samsung smart phones (Android), Amazon kindles, DLink surveillance cameras, TP-Link smart plugs, etc. In this paper, we present a novel approach, namely WDMTI, for the identification of wireless device manufacturer and type. We tackle the challenge from two aspects: the features and the classification model. First, we claim that it is critical to discover the device manufacturer and type as soon as the device requests to join the WLAN, and it is unrealistic to make other assumptions on the status of the device, e.g., assuming that the device is booting up or initializing a new connection to corresponding servers/clouds. We primarily depend on the features extracted from the network connection phase, while features from device booting are considered "bonus". In particular, we propose to utilize features from the raw HDCP packets, which is shown to be sufficient for device manufacturer and type recognition with high accuracy. Meanwhile, in the WDMTI system, we employ the Hierarchical Dirichlet Process (HDP), which is a nonparametric Bayesian model for grouped data. HDP allows new groups to be introduced with new data being added, i.e. previously unknown devices connect to the network and the extracted features receive new labels. The WDMTI mechanism is dynamically retrained on-line, instead of requiring a time-consuming off-line retraining process. Our experiments show that WDMTI identifies known types of devices with average accuracy of 0.89, and new types of devices with average accuracy of 0.96, both of which is higher than the state-of-art approaches. In summary, we present a wireless device manufacturer and type identification (WDMTI) system that is both scalable and accurate, and capable of adapting to unknown types of devices on-the-fly.
Lingjing Yu, Zhaoyu Zhou, Yujia Zhu, Qingyun Liu 0001, Jianlong Tan
MASS4
2018 A Privacy Protection Model of Data Publication Based on Game Theory
abstract
With the rapid development of sensor acquisition technology, more and more data are collected, analyzed, and encapsulated into application services. However, most of applications are developed by untrusted third parties. Therefore, it has become an urgent problem to protect users’ privacy in data publication. Since the attacker may identify the user based on the combination of user’s quasi-identifiers and the fewer quasi-identifier fields result in a lower probability of privacy leaks, therefore, in this paper, we aim to investigate an optimal number of quasi-identifier fields under the constraint of trade-offs between service quality and privacy protection. We first propose modelling the service development process as a cooperative game between the data owner and consumers and employing the Stackelberg game model to determine the number of quasi-identifiers that are published to the data development organization. We then propose a way to identify when the new data should be learned, as well, a way to update the parameters involved in the model, so that the new strategy on quasi-identifier fields can be delivered. The experiment first analyses the validity of our proposed model and then compares it with the traditional privacy protection approach, and the experiment shows that the data loss of our model is less than that of the traditional k-anonymity especially when strong privacy protection is applied.
Li Kuang, Yujia Zhu, Xuejin Yan, Shuiguang Deng
Secur. Commun. Networks2
2017 Semantics and locality preserving correlation projections
abstract
Multi-view correlation learning has attracted great attention with the proliferation of heterogeneous data. Typical methods, such as Canonical Correlation Analysis (CCA) and its variants, usually maximize one-to-one corresponding correlation of inter-view data, while most of them neglect discriminative multi-label information and local structure of each view data. In this paper, we propose multi-label Semantics and Locality Preserving Correlation Projections method (SLPCP), which seeks for a semantic common subspace by jointly learning view-specific linear projections from intra-view and interview perspectives simultaneously. SLPCP can be easily optimized with generalized eigen value decomposition via concatenating the projections of multi-views. Applied to retrieval tasks of image and text data in experiments, SLPCP out performs state-of-the-art methods on a widely used dataset NUS-WIDE. The extensive experiments also validate that it is effective to preserve the multi-label semantics and locality of multi-view data.
Yan Hua, Jianke Du, Yujia Zhu
ICME3
2017 WiFi fingerprint releasing for indoor localization based on differential privacy
abstract
WiFi fingerprint-based localization is regarded as one of the most promising techniques for indoor localization. However, this raises serious privacy concerns. Current approaches to mitigate the privacy concerns rely on the encryption with large calculation consumption. In this paper, we propose a data obfuscation mechanism based on the generalized version of differential privacy. We extend the standard definition to the indoor WiFi fingerprint data for spatial counting where the inputs belong to multiple dimensions of numerical data in a limited range. With a given privacy budget, the proposed method generalizes the original dataset, and then specializes it using differential privacy. As the designed novel scheme expand the range for specialization, the data set released by the proposed algorithm can yield better mining results. Furthermore, experimental results give out comparisons between nonuniform and uniform ε selection scheme, and find uniform ε selection scheme can fully use the privacy budget in our situation.
Yujia Zhu, Qingyun Liu 0001, Yang Aron Liu, Peng Zhang 0001
PIMRC1
2017 Collaborative sparse representation leaning model for RGBD action recognition
Su-hua Li, Yujia Zhu, Hua Zhang 0003
J. Vis. Commun. Image Represent.3
2017 Evaluation of regularized multi-task leaning algorithms for single/multi-view human action recognition
Su-hua Li, Guotai Zhang, Yujia Zhu, Hua Zhang 0003
Multim. Tools Appl.4
2014 Location Privacy in Buildings: A 3-Dimensional K-Anonymity Model
abstract
Privacy protection has recently received considerable attention in location-based services. In this paper, we show that most of the existing k-anonymity location cloaking algorithms are concerned only and cannot effectively prevent location-dependent attacks when users' locations have height information. Therefore, adopting the three dimensional location information, we propose a new clique-based cloaking algorithm, called 3d Clique Cloak, to defend against location leaks in indoor environment. The main idea is to expand the MBV (minimum bounding volume) to a three-dimensional space, thus for a user who initiated location services can find k-anonymity cloaking set in the three-dimensional space. The efficiency and effectiveness of the proposed 3d Clique Cloak algorithm are validated by series of carefully designed experiments.
Yujia Zhu, Lidong Zhai
MSN1
2012 AP Selection for Indoor Localization Based on Neighborhood Rough Sets
abstract
In this paper, a new AP (access points) selection method is proposed based on neighborhood rough sets (NRS) for indoor localization. Due to the existence of virtual APs in large scale buildings, received signal strength (RSS) may produce similar measurements, leading to biased estimates and redundant computations. Our method utilize neighborhood relations to transform the radio map into an extended rough set model such that offers a more discriminative way to select a small subset of available APs. Experimental results show that our method can be used to further improve the performance of the original AP selection methods.
Yujia Zhu
VTC Fall1
2011 Robust head pose estimation via semi-supervised manifold learning with ℓ1-graph regularization
abstract
In this paper; a new ℓ1-graph regularized semi- supervised manifold learning (LRSML) method is proposed for robust human head pose estimation problem. The manifold is constructed under Biased Manifold Embedding (BME) framework which computes a biased neighborhood of each point in the feature space with ℓ1-graph regularization. The construction process of ℓ1-graph is assumed to be unsupervised without harnessing any data label information and uncovers the underlying ℓ1-norm driven sparse reconstruction relationship of each sample. The LRSML is more robust to noises and has the potential to convey more discriminative information compared to conventional manifold learning methods. Furthermore, utilizing both labeled and unlabeled information improve the pose estimation accuracy and generalization capability. Numerous experiments show the superiority of our method over several current state of the art methods on publicly available dataset.
Yujia Zhu
IJCB3
2011 Incremental nonparametric discriminant analysis for robust object tracking
abstract
In this paper, a new adaptive subspace learning model based on incremental nonparametric discriminant analysis (INDA) is proposed for visual tracking. Traditional subspace trackers focus on updating eigenvectors in handling with appearance variation of the target object, ignoring the non-target background region during tracking. The INDA features take both of them into consideration, thereby promoting the tracking process in the ever-changing environment. Meanwhile, INDA relaxes the Gaussian assumption in Fisher discriminant analysis (FDA), so it can handle more general class distributions problem. The scatter matrices are also reformulated to update the subspace incrementally based on previous results. In conjunction with efficient feature extraction method, the system is real time capable. Numerous experiments show the superiority of our tracker over current states of art methods on several publicly available datasets.
Yujia Zhu
ICME3