Liyao Xiang

dblp:115/6308 · DBLP profile ↗
← Back
49ranked-venue papers
9as first author
41since 2021 · last 2026
0000-0003-0165-4930ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 16 · 7 first-author · 11 since 2021Security and privacy · 14 · 14 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FedTopo: Topology-Informed Representation Alignment in Federated Learning Under Non-I.I.D. Conditions
abstract
Current federated-learning models deteriorate under heterogeneous (non-I.I.D.) client data, as their feature representations diverge and pixel- or patch-level objectives fail to capture the global topology which is essential for high-dimensional visual tasks. We propose FedTopo, a framework that integrates Topological-Guided Block Screening (TGBS) and Topological Embedding (TE) to leverage topological information, yielding coherently aligned cross-client representations by Topological Alignment Loss (TAL). First, Topology-Guided Block Screening (TGBS) automatically selects the most topology-informative block, i.e., the one with maximal topological separability, whose persistence-based signatures best distinguish within- versus between-class pairs, ensuring that subsequent analysis focuses on topology-rich features. Next, this block yields a compact Topological Embedding, which quantifies the topological information for each client. Finally, a Topological Alignment Loss (TAL) guides clients to maintain topological consistency with the global model during optimization, reducing representation drift across rounds. Experiments on Fashion-MNIST, CIFAR-10, and CIFAR-100 under four non-I.I.D. partitions show that FedTopo accelerates convergence and improves accuracy over strong baselines.
Liyao Xiang, Peng Tang 0002, Weidong Qiu
AAAI2
2026 Faster Than Ever: A New Lightweight Private Set Intersection and Its Variants
Guowei Ling, Peng Tang 0002, Jinyong Shan, Liyao Xiang, Weidong Qiu
NDSS4
2026 PrivSniffer: Graph-based Contextual Privacy Leakage Detection for User-Generated Texts
abstract
A vast amount of user-generated content is uploaded on social media everyday, potentially being collected as training data for large language models, and thus poses severe threats to individual privacy. Existing approaches for protecting user-generated content often fail to detect implicit privacy leaks which are not directly given but could be inferred from the contextual text (clues). The precise detection of leaks and clues is difficult but critical in text sanitization. We propose a graph-based model for the user-generated content and build a privacy leakage detector PrivSniffer which integrates the inference capability of language models with the precise graph search, to capture the contextual privacy leakage (CPL). To evaluate the effectiveness of the framework, we create SynthLeak, a dataset of dialogues containing human-labeled implicit leaks, while featuring diversity and naturalness. Our experimental results on SynthLeak and several other benchmarks reveal PrivSniffer's superior capability in detecting implicit leaks and CPLs. Particularly in CPL detection, the state-of-the-art approach achieves only an F1 Score of 0.36, whereas PrivSniffer attains 0.59 on SynthLeak, showing great promise in real-world text sanitization.
Hangyu Ye, Liyao Xiang, Naixuan Huang, Dongyue Yu
WWW2
2026 DeepFP: Deep-Unfolded Fractional Programming for MIMO Beamforming
abstract
This work proposes a mixed learning-based and optimization-based approach to the weighted-sum-rates beamforming problem in a multiple-input multiple-output (MIMO) wireless network. The conventional methods, i.e., the fractional programming (FP) method and the weighted minimum mean square error (WMMSE) algorithm, can be computationally demanding for two reasons: (i) they require inverting a sequence of matrices whose sizes are proportional to the number of antennas; (ii) they require tuning a set of Lagrange multipliers to account for the power constraints. The recently proposed method called the reduced WMMSE addresses the above two issues for a single cell. In contrast, for the multicell case, another recent method called the FastFP eliminates the large matrix inversion and the Lagrange multipliers by using an improved FP technique, but the update stepsize in the FastFP can be difficult to decide. As such, we propose integrating the deep unfolding network into the FastFP for the stepsize optimization. Numerical experiments show that the proposed method is much more efficient than the learning method based on the WMMSE algorithm.
Jianhang Zhu, Tsung-Hui Chang, Liyao Xiang, Kaiming Shen
IEEE Trans. Commun.3
2025 Interpretable Rotation-Equivariant Multiary-Valued Network for Attribute Obfuscation
abstract
This paper focuses on the problem of preventing information leakage in neural networks, i.e., assuming that attackers have obtained intermediate-layer features of a neural network, and preventing attackers from inverting these features to the input with private information. We propose a generic method to slightly revise each arbitrary traditional neural network into a multiary-valued rotation-equivariant neural network (RENN) for preventing information leakage. Specifically, we convert real-valued features in the network into multi-ary features, and each element in the feature vector is a multi-ary number. We hide the input information into a certain phase of the multi-ary feature, and rotate the multi-ary feature for attribute obfuscation in the encryption process. The rotation axis and angle can be considered as the private key. In this way, even when attackers have obtained network parameters and intermediate-layer features, they still cannot extract input information without knowing the rotation information. More crucially, the encryption operation does not damage the spatial correlations between features, so that the encrypted features can be easily processed by convolution operations in the neural network without difficulties. In order to implement successful encryption and decryption, the RENN is designed to satisfy the rotation equivariance property. To this end, we propose a set of rules to revise classic operations in the neural network to ensure the rotation equivariance property. Besides, we prove that the $d$d-ary RENN is downward compatible with the $d^{\prime }$d'-ary RENN when $d^{\prime }< d$d'
Quanshi Zhang, Hao Zhang 0063, Yiting Chen 0003, Qihan Ren, Jie Ren 0018, Xu Cheng 0005, Liyao Xiang
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 CodeMark: Contextual and Natural Watermarking for Tracing Code Snippet Provenance
abstract
Determining the origins of code snippets has gained increasing attention due to the popularity of large language models and the concern about their misuse in generating unlicensed or malicious code. Watermarking is considered a working solution for tracing code snippet provenance. However, source code watermarking requires more stringent and intricate rules than natural language or software watermarking, since one needs to ensure both readability and functionality of the watermarked code snippets. To this end, we propose a novel watermarking systemCodeMark, featured by variable renaming as the key. Surrounding variable renaming, several challenges emerge such as determining renaming candidates, defining the variable context, and providing diverse variable substitutes, etc. The design ofCodeMarkconquers these challenges by an end-to-end learning system, which interprets the code context through Graph Neural Networks (GNNs) and generates natural substitutes fitting the context by distilling from CodeBERT. Experiments illustrate thatCodeMarksurpasses the state-of-the-art watermarking systems in terms of watermarking requirements.
Wei Li 0254, Borui Yang, Yujie Sun 0001, Suyu Chen, Yuting Chen 0001, Liyao Xiang
IEEE Trans. Dependable Secur. Comput.6
2025 Privacy-Preserving Authorized Set Matching via Dishonest Majority Multiparty Computation
abstract
Private Set Intersection (PSI) enables each party with a private set to compute the intersection without disclosing other information. However, even in maliciously secure PSI, it does not guarantee input authenticity and output integrity, which becomes problematic in certain scenarios. For instance, in Web 3.0, one of the essential requirements is to find common certifiers among the parties. However, if certifier identities are meant to be protected, some parties may attempt to forge certifier identities or intentionally exclude a particular certifier during protocol execution. Recently, the Private Certifier Intersection (PCI), a variant of PSI, has been proposed to address this problem. Nevertheless, it incurs significantly high computational and communication overhead. This work proposes thePrivate Identity Intersection(PII), which takes private identifiers and corresponding anonymous signatures from mutually distrusting parties as input, verifies them, and delivers the intersection of the successfully verified identifiers to all parties while ensuring the integrity of the output. Furthermore, PII can naturally extend from two to multiple-party settings while resisting the collusion attack. To achieve the ideal functionality of PII, we implement a user-friendly MPC framework called$\mathsf {Oryx}$without third-party libraries. Based on$\mathsf {Oryx}$, we instantiate PII with two digital signature schemes, one proposed in this paper. Compared to existing work, our PII protocols reduce the computation overhead by up to$163\times$and the communication overhead by up to$190\times$, representing an improvement of two orders of magnitude. To demonstrate the practicality of our work, we evaluate its performance in WAN environments with bandwidths of 100 Mbps and 500 Mbps, under a fixed latency of 20 ms.
Guowei Ling, Peng Tang 0002, Fei Tang 0001, Shifeng Sun 0001, Jinyong Shan, Liyao Xiang, Weidong Qiu
IEEE Trans. Dependable Secur. Comput.6
2025 Shuffling for Semantic Secrecy
abstract
Deep learning draws heavily on the latest progress in semantic communications. The present paper aims to examine the security aspect of this cutting-edge technique from a novel shuffling perspective. Our goal is to improve upon the conventional secure coding scheme to strike a desirable tradeoff between transmission rate and leakage rate. To be more specific, for a wiretap channel, we seek to maximize the transmission rate while minimizing the semantic error probability under the given leakage rate constraint. Toward this end, we devise a novel semantic security communication system wherein the random shuffling pattern plays the role of the shared secret key. Intuitively, the permutation of feature sequences via shuffling would distort the semantic essence of the target data to a sufficient extent so that eavesdroppers cannot access it anymore. The proposed random shuffling method also exhibits its flexibility in working for the existing semantic communication system as a plugin. Simulations demonstrate the significant advantage of the proposed method over the benchmark in boosting secure transmission, especially when channels are prone to strong noise and unpredictable fading.
Fupei Chen, Liyao Xiang, Haoxiang Sun, Hei Victor Cheng, Kaiming Shen
IEEE Trans. Inf. Forensics Secur.2
2024 Curator Attack: When Blackbox Differential Privacy Auditing Loses Its Power
abstract
A surge in data-driven applications enhances everyday life but also raises serious concerns about private information leakage. Hence many privacy auditing tools are emerging for checking if the data sanitization performed meets the privacy standard of the data owner. Blackbox auditing for differential privacy is particularly gaining popularity for its effectiveness and applicability to a wide range of scenarios. Yet, we identified that blackbox auditing is essentially flawed with its setting --- small probabilities/densities are ignored due to inaccurate observation. Our argument is based on a solid false positive analysis from a hypothesis testing perspective, which is missed out by prior blackbox auditing tools. This oversight greatly reduces the reliability of these tools, as it allows malicious or incapable data curators to pass the auditing with an overstated privacy guarantee, posing significant risks to data owners. We demonstrate the practical existence of such threats in classical differential privacy mechanisms against four representative blackbox auditors with experimental validations. Our findings aim to reveal the limitations of blackbox auditing tools, empower the data owner with the awareness of risks in using these tools, and encourage the development of more reliable differential privacy auditing methods.
Liyao Xiang, Bowei Cheng, Tianran Sun, Xinbing Wang
CCS2
2024 Permutation Equivariance of Transformers and its Applications
abstract
Revolutionizing the field of deep learning, Transformer-based models have achieved remarkable performance in many tasks. Recent research has recognized these models are robust to shuffling but are limited to inter-token permutation in the forward propagation. In this work, we propose our definition of permutation equivariance, a broader concept covering both inter- and intra- token per-mutation in the forward and backward propagation of neural networks. We rigorously proved that such permutation equivariance property can be satisfied on most vanilla Transformer-based models with almost no adaptation. We examine the property over a range of state-of-the-art models including ViT, Bert, GPT, and others, with experimental validations. Further, as a proof-of-concept, we explore how real-world applications including privacy-enhancing split learning, and model authorization, could exploit the permutation equivariance property, which implicates wider, intriguing application scenarios. The code is available at https://github.com/Doby-Xu/ST
Hengyuan Xu, Liyao Xiang, Hangyu Ye, Dixi Yao, Pengzhi Chu, Baochun Li
CVPR2
2024 CROSSWORD: A Semantic Approach To Text Compression Via Masking
abstract
Conventional data compression methods typically model the information source as an i.i.d. stochastic process, thereby establishing the fundamental limit as entropy for lossless compression and as mutual information for lossy compression. However, the source in the real world (e.g., text, music, and speech) is often statistically ill-defined because of its close connection to human perception. This work aims to exploit the semantic aspect of text as inspired by the puzzle crossword. The main idea is to only compress those semantically important words while masking the rest; the proposed decompressor can recover all the missing words automatically according to context. Experiments show that the proposed semantic approach can achieve much higher compression efficiency than the state-of-the-art semantic compression method.
Liyao Xiang, Kaiming Shen, Shuguang Cui
ICASSP3
2024 Feature Norm Regularized Federated Learning: Utilizing Data Disparities for Model Performance Gains
Liyao Xiang, Peng Tang 0002, Weidong Qiu
IJCAI2
2024 A Principled Approach to Natural Language Watermarking
abstract
Recently, there has been a surge in machine-generated natural language content being misused by unauthorized parties. Watermarking is a well-recognized technique to address the issue by tracing the provenance of the text. However, we found that most existing watermarking systems for texts are subject to ad hoc design and thus suffer from fundamental vulnerabilities.
Qiansiqi Hu, Yicheng Zheng, Liyao Xiang, Xinbing Wang
ACM Multimedia4
2024 Crafter: Facial Feature Crafting against Inversion-based Identity Theft on Deep Models
Liyao Xiang, Hao Zhang 0063, Xinbing Wang, Chenghu Zhou, Bo Li 0115
NDSS3
2024 Lambda: Learning Matchable Prior For Entity Alignment with Unlabeled Dangling Cases
abstract
We investigate the entity alignment (EA) problem with unlabeled dangling cases, meaning that partial entities have no counterparts in the other knowledge graph (KG), yet these entities are unlabeled. The problem arises when the source and target graphs are of different scales, and it is much cheaper to label the matchable pairs than the dangling entities. To address this challenge, we propose the framework \textit{Lambda} for dangling detection and entity alignment. Lambda features a GNN-based encoder called KEESA with a spectral contrastive learning loss for EA and a positive-unlabeled learning algorithm called iPULE for dangling detection. Our dangling detection module offers theoretical guarantees of unbiasedness, uniform deviation bounds, and convergence. Experimental results demonstrate that each component contributes to overall performances that are superior to baselines, even when baselines additionally exploit 30\% of dangling entities labeled for training.
Hang Yin 0009, Liyao Xiang, Yuheng He, Pengzhi Chu, Xinbing Wang, Chenghu Zhou
NeurIPS2
2024 SrcMarker: Dual-Channel Source Code Watermarking via Scalable Code Transformations
abstract
The expansion of the open source community and the rise of large language models have raised ethical and security concerns on the distribution of source code, such as misconduct on copyrighted code, distributions without proper licenses, or misuse of the code for malicious purposes. Hence it is important to track the ownership of source code, in which watermarking is a major technique. Yet, drastically different from natural languages, source code watermarking requires far stricter and more complicated rules to ensure the readability as well as the functionality of the source code. Hence we introduce SrcMarker, a watermarking system to unobtrusively encode ID bitstrings into source code, without affecting the usage and semantics of the code. To this end, SrcMarker performs transformations on an AST-based intermediate representation that enables unified transformations across different programming languages. The core of the system utilizes learning-based embedding and extraction modules to select rule-based transformations for watermarking. In addition, a novel feature-approximation technique is designed to tackle the inherent non-differentiability of rule selection, thus seamlessly integrating the rule-based transformations and learning-based networks into an interconnected system to enable end-to-end training. Extensive experiments demonstrate the superiority of SrcMarker over existing methods in various watermarking requirements.
Borui Yang, Wei Li 0254, Liyao Xiang, Bo Li 0001
SP3
2024 A Verifiable and Privacy-Preserving Federated Learning Training Framework
abstract
Federated learning allows multiple clients to collaboratively train a global model without revealing their private data. Despite its success in many applications, it remains a challenge to prevent malicious clients to corrupt the global model through uploading incorrect model updates. Hence, one critical issue arises in how to validate the training is truly conducted on legitimate neural networks. To address the issue, we proposeVPNNT, a zero-knowledge proof scheme for neural network backpropagation.VPNNTenables each client to prove to others that the model updates (gradients) are indeed calculated on the global model of the previous round, without leaking any information about the client's private training data. Our proof scheme is generally applicable to any type of neural network. Different from conventional verification schemes constructing neural network operations by gate-level circuits, we improve verification efficiency by formulating the training process using custom gates — matrix operations, and apply an optimized linear time zero knowledge protocol for verification. Thanks to the recursive structure of neural network backward propagation, common custom gates are combined in verification thereby reducing prover and verifier costs over conventional zero knowledge proofs. Experimental results show thatVPNNTis a lightweighted verification scheme for neural network backpropagation with an improved prove time, verification time and proof size.
Haohua Duan, Zedong Peng, Liyao Xiang, Yuncong Hu, Bo Li 0001
IEEE Trans. Dependable Secur. Comput.3
2024 MiniTracker: Large-Scale Sensitive Information Tracking in Mini Apps
abstract
Running on host mobile applications, mini apps have gained increasing popularity these days for its convenience in installation and usage. However, being easy to use allows mini apps to freely access a large amount of user information, mostly without close inspection of privacy violations. Hence it becomes a crucial issue to automatically track sensitive flows in mini apps. Although flow analysis has been widely studied, unique challenges emerge: the analysis tool should not only handle mini app-specific features such as flows that interweave between rendering and logic, and asynchronous executions, but also deal with problems raised by Javascript development: the performance tradeoff between precision and efficiency, and function aliases. To this end, we proposeMiniTracker, an automatic sensitive flow tracking tool which well handles mini app features, constructs assignment flow graphs as common representation across different host apps, searches function aliases, and analyzes the graph by property chains. We show our design choices achieve a sweet spot in the tradeoff between precision and efficiency, with superior performance compared to the state-of-the-art. We also perform a large-scale study on 150 k mini apps, which reveals the common leakage patterns and offers insights into the privacy threats of mini apps.
Wei Li 0254, Borui Yang, Hangyu Ye, Liyao Xiang, Qingxiao Tao, Xinbing Wang, Chenghu Zhou
IEEE Trans. Dependable Secur. Comput.4
2024 Certified Distributional Robustness on Smoothed Classifiers
abstract
The robustness of deep neural networks (DNNs) against adversarial example attacks has raised wide attention. For smoothed classifiers, we propose the worst-case adversarial loss over input distributions as a robustness certificate. Compared with previous certificates, our certificate better describes the empirical performance of the smoothed classifiers. By exploiting duality and the smoothness property, we provide an easy-to-compute upper bound as a surrogate for the certificate. We adopt a noisy adversarial learning procedure to minimize the surrogate loss to improve model robustness. We show that our training method provides a theoretically tighter bound over the distributional robust base classifiers. Experiments on a variety of datasets further demonstrate superior robustness performance of our method over the state-of-the-art certified or heuristic methods.
Jungang Yang 0002, Liyao Xiang, Pengzhi Chu, Xinbing Wang, Chenghu Zhou
IEEE Trans. Dependable Secur. Comput.2
2024 Distributional Learning for Network Alignment with Global Constraints
abstract
Network alignment, pairing corresponding nodes across the source and target networks, plays an important role in many data mining tasks. Extensive studies focus on learning node embeddings across different networks in a unified space. However, these methods have not taken the large structural discrepancy between aligned nodes into account and, thus, are largely confined by the deterministic representations of nodes. In this work, we propose a novel network alignment framework highlighted by distributional learning and globally optimal alignment. By modeling the uncertainty of each node by Gaussian distribution, our framework builds similarity matrices on the Wasserstein distance between distributions and applies Sinkhorn operation, which learns the globally optimal mapping in an end-to-end fashion. We show that each integrated part of the framework contributes to the overall performance. Under a variety of experimental settings, our alignment framework shows superior accuracy and efficiency to the state-of-the-art.
Hui Xu 0011, Liyao Xiang, Xiaoying Gan, Luoyi Fu, Xinbing Wang, Chenghu Zhou
ACM Trans. Knowl. Discov. Data2
2024 Open-World Graph Active Learning for Node Classification
abstract
The great power of Graph Neural Networks (GNNs) relies on a large number of labeled training data, but obtaining the labels can be costly in many cases. Graph Active Learning (GAL) is proposed to reduce such annotation costs, but the existing methods mainly focus on improving labeling efficiency with fixed classes, and are limited to handle the emergence of novel classes. We term the problem as Open-World Graph Active Learning (OWGAL) and propose a framework of the same name. The key is to recognize novel-class as well as informative nodes in a unified framework. Instead of a fully connected neural network classifier, OWGAL employs prototype learning and label propagation to assign high uncertainty scores to the targeted nodes in the representation and topology space, respectively. Weighted sampling further suppresses the impact of unimportant classes by weighing both the node and class importance. Experimental results on four large-scale datasets demonstrate that our framework achieves a substantial improvement of 5.97% to 16.57% on Macro-F1 over state-of-the-art methods.
Hui Xu 0011, Liyao Xiang, Junjie Ou, Yuting Weng, Xinbing Wang, Chenghu Zhou
ACM Trans. Knowl. Discov. Data2
2024 Learning to Prevent Input Leakages in the Mobile Cloud Inference
abstract
Powered by machine learning services in the cloud, numerous learning-driven mobile applications are gaining popularity in the market. As deep learning tasks are mostly computation-intensive, it has become a trend to process raw data on devices and send the deep neural network (DNN) features to the cloud, where the features are further processed to return final results. However, there is always an unexpected leakage with the release of features, by which an adversary could infer much information on the original data. We propose a privacy-preserving framework on top of the mobile cloud infrastructure from the perspective of DNN structures. Our framework aims to learn a policy to modify the base DNNs to prevent information leakage while maintaining high inference accuracy. The policy can also be readily transferred to large-size DNNs and large-scale datasets to speed up learning. Extensive evaluations on a variety of DNNs have shown that our framework successfully finds privacy-preserving DNN structures to defend privacy attacks.
Liyao Xiang, Shuang Zhang 0007, Quanshi Zhang
IEEE Trans. Mob. Comput.1
2023 Privacy-Preserving Split Learning via Pareto Optimal Search
Liyao Xiang, Chengnian Long
ESORICS (4)2
2023 Deep Learning Enabled Semantic-Secure Communication with Shuffling
abstract
Deep learning and natural language processing draw heavily on the recent progress in semantic communications; this paper examines the security aspect of this cutting-edge technique. Our goal is to improve upon the conventional secure coding methods to strike a superior tradeoff between transmission rate and leakage rate. Toward this end, we devise a novel semantic security communication system wherein the random shuffling pattern serves as the secret key shared. Intuitively, the permutation of words in the same text via shuffling would result in the meaning distortion of the target text to such a great extent that an eavesdropper can no longer recover the semantic truth. The proposed method can be rephrased as maximizing the transmission rate while minimizing the semantic error probability under the given leakage rate constraint. Simulations demonstrate the significant advantage of the proposed method over the benchmark in boosting secure transmission, especially when channels are prone to strong noise and unpredictable fading, can achieve up to 60% performance gain.
Fupei Chen, Liyao Xiang, Hei Victor Cheng, Kaiming Shen
GLOBECOM2
2023 Mixup Training for Generative Models to Defend Membership Inference Attacks
abstract
With the popularity of machine learning, it has been a growing concern on the trained model revealing the private information of the training data. Membership inference attack (MIA) poses one of the threats by inferring whether a given sample participates in the training of the target model. Although MIA has been widely studied for discriminative models, for generative models, neither it nor its defense is extensively investigated. In this work, we propose a mixup training method for generative adversarial networks (GANs) as a defense against MIAs. Specifically, the original training data is replaced with their interpolations so that GANs would never overfit the original data. The intriguing part is an analysis from the hypothesis test perspective to theoretically prove our method could mitigate the AUC of the strongest likelihood ratio attack. Experimental results support that mixup training successfully defends the state-of-the-art MIAs for generative models, yet without model performance degradation or any additional training efforts, showing great promise to be deployed in practice.
Qiansiqi Hu, Liyao Xiang, Chenghu Zhou
INFOCOM3
2023 Grace: Graph Self-Distillation and Completion to Mitigate Degree-Related Biases
abstract
Due to the universality of graph data, node classification shows its great importance in a wide range of real-world applications. Despite the successes of Graph Neural Networks (GNNs), GNN based methods rely heavily on rich connections and perform poorly on low-degree nodes. Since many real-world graphs follow a long-tailed distribution in node degrees, they suffer from a substantial performance bottleneck as a significant fraction of nodes is of low degree. In this paper, we point out that under-represented self-representations and low neighborhood homophily ratio of low-degree nodes are two main culprits. Based on that, we propose a novel method Grace which improves the node representation by self-distillation, and increases neighborhood homophily ratio of low-degree nodes by graph completion. To avoid error propagation of graph completion, label propagation is further leveraged. Experimental evidence has shown that our method well supports real-world graphs, and is superior in balancing degree-related bias and overall performance on node classification tasks.
Hui Xu 0011, Liyao Xiang, Femke Huang, Yuting Weng, Ruijie Xu 0005, Xinbing Wang, Chenghu Zhou
KDD2
2023 Adaptive Beamforming for Non-Line-of-Sight IRS-Assisted Communications without CSI
abstract
Channel acquisition is a major bottleneck in fully exploiting the potential of intelligent reflecting surfaces (IRSs) to improve the wireless environment. In order to bypass such difficulty, an alternative is to optimize IRS based on the received signal statistics rather than channel state information (CSI), namely blind beamforming. The two recent methods, RFocus and conditional sample mean (CSM), fall into this category, both of which have been shown highly effective in practice. Nevertheless, we find a subtle drawback with the existing blind beamforming methods that they may not work well for the non-line-of-sight (NLoS) case for two reasons. First, many more signal samples are needed when the direct propagation diminishes. Second, if the direct propagation is completely blocked then the existing blind beamforming methods cannot work whatsoever. To address this issue, we propose an adaptive strategy for blind beamforming, which guarantees an approximation ratio of the global optimum. Field tests and simulations show that the proposed blind beamforming method is much more suited for NLoS environment than the existing ones.
Wenhai Lai, Shuyi Ren, Liyao Xiang, Xin Li 0112, Shaobo Niu, Kaiming Shen
PIMRC4
2023 A New Zero Knowledge Argument for General Circuits and Its Application
abstract
Verifying the correctness of computation without revealing the input is a critical issue intensively studied in real-world applications. The recent surge of zero knowledge arguments has been focusing on its efficiency and practicality. Among them, GKR-based arguments have received wide attention and become the foundation of many zero-knowledge proof protocols. However, GKR-based protocols are restricted to layered arithmetic circuits. We proposeTerrace, a new, efficient zero-knowledge argument system for general circuits, based on GKR. By dynamically patching cross-layer claims to the original circuit for verification instead of verifying those claims separately,Terraceis able to reduce the total circuit size and thus enjoys a logarithmic factor less verification time and proof size.Terraceis further extended to include the verification of non-arithmetic operations by rewriting those claims in the multilinear extension form. Experimental results demonstrate that Terrace enjoys a competitive performance on efficiency, and shows great promise in enabling low-cost verification of neural networks.
Haohua Duan, Liyao Xiang, Xinbing Wang, Pengzhi Chu, Chenghu Zhou
IEEE Trans. Inf. Forensics Secur.2
2023 DPlanner: A Privacy Budgeting System for Utility
abstract
Differential mymargin privacy has been deployed to machine learning platforms to preserve the privacy of data in use. A long neglected but important fact is that data privacy is a non-replenishable resource and should be carefully scheduled to maximize its utility gain. In this work, we propose a new privacy budgeting system—DPlanner, which estimates data blocks’ importance to queries and assigns fractional privacy budget to those data blocks contributing most to a query. The scheduler is novelly designed to include two-fold randomness, which satisfies differential privacy with tight budgets, at the same time guarantees the expected utility in the worst-case query sequence when queries arrive in an online fashion. Experiments in a variety of machine learning settings have shown that our DPlanner outperforms the state-of-the-art schedulers by serving at least 25% more queries, or reducing the total privacy consumption by over 50%.
Weiting Li, Liyao Xiang, Bin Guo 0001, Zhetao Li, Xinbing Wang
IEEE Trans. Inf. Forensics Secur.2
2023 Differentially-Private Deep Learning With Directional Noise
abstract
With the popularity of deep learning applications, the privacy of training data has become a major concern as the data sources may be sensitive. Recent studies have found that deep learning models are vulnerable to privacy attacks, which are able to infer private training data from model parameters. To mitigate such attacks, differential privacy has been proposed to preserve data privacy by adding randomized noise to these models. However, since deep learning models usually consist of a large number of parameters and complicated layered structures, an overwhelming amount of noise is often inserted, which significantly degrades model accuracy. We seek a better tradeoff between model utility and data privacy, by choosing directions of noise w.r.t. the utility subspace. We propose an optimized mechanism for differentially-private stochastic gradient descent, and derive a closed-form solution. The form of the solution makes the mechanism ready to be deployed in real-world deep learning systems. Experimental results on a variety of models, datasets, and privacy settings show that our proposed mechanism achieves higher accuracies at the same privacy guarantee compared to the state-of-the-art methods. Further, we extend the privacy guarantee to a mutual information bound, and propose a general form to the utility-privacy problem.
Liyao Xiang, Weiting Li, Jungang Yang 0002, Xinbing Wang, Baochun Li
IEEE Trans. Mob. Comput.1
2023 Matrix Gaussian Mechanisms for Differentially-Private Learning
abstract
The wide deployment of machine learning algorithms has become a severe threat to user data privacy. As the learning data is of high dimensionality and high orders, preserving its privacy is intrinsically hard. Conventional differential privacy mechanisms often incur significant utility decline as they are designed for scalar values from the start. We recognize that it is because conventional approaches do not take the data structural information into account, and fail to provide sufficient privacy or utility. As the main novelty of this work, we proposeMatrix Gaussian Mechanism(MGM), a new$ (\epsilon,\delta)$-differential privacy mechanism for preserving learning data privacy. By imposing the unimodal distributions on the noise, we introduce two mechanisms based on MGM with an improved utility. We further show that with the utility space available, the proposed mechanisms can be instantiated with optimized utility, and has a closed-form solution scalable to large-scale problems. We experimentally show that our mechanisms, applied to privacy-preserving federated learning, are superior than the state-of-the-art differential privacy mechanisms in utility.
Jungang Yang 0002, Liyao Xiang, Jiahao Yu 0001, Xinbing Wang, Bin Guo 0001, Zhetao Li, Baochun Li
IEEE Trans. Mob. Comput.2
2022 Privacy-Preserving Split Learning via Patch Shuffling over Transformers
abstract
We focus on the privacy-preserving problem in split learning in this work. In vanilla split learning, a neural network is split to different devices to be trained, risking leaking the private training data in the process. We novelly propose a patch shuffling scheme on transformers to preserve training data privacy, yet without degrading overall model performance. Formal privacy guarantees are provided and we further introduce the batch shuffling and the spectral shuffling schemes to enhance the guarantee. We show through experiments that our methods successfully defend the black-box, white-box, and adaptive attacks in split learning, with superior performance over baselines, and are efficient to deploy with negligible overhead compared to the vanilla split learning.
Dixi Yao, Liyao Xiang, Hengyuan Xu, Hangyu Ye
ICDM2
2022 CAQ: Toward Context-Aware and Self-Adaptive Deep Model Computation for AIoT Applications
abstract
Artificial Intelligence of Things (AIoT) has recently accepted significant interests. Remarkably, embedded artificial intelligence (e.g., deep learning) on-device transforms IoT devices into intelligent systems that robustly and privately process data. Quantization technique is widely used to compress deep models for narrowing the resource gap between computation demands and platform supply. However, existing quantization schemes induce unsatisfaction for IoT scenarios since they are oblivious to dynamic changes of application context (e.g., battery and hierarchical memory availability) during the long-term operation. Subsequently, they will mismatch the user-desired resource efficiency and application lifetime. Also, to adapt to the dynamic context, we can neither accept the latency for model retraining with existing hand-crafted quantization nor the overhead for quantization bit width researching with prior on-demand quantization. This article presents a context-aware and self-adaptive deep model quantization (CAQ) system for IoT application scenarios. CAQ integrates a novel switchable multigate quantization framework, optimizing the quantized model accuracy and energy efficiency in diverse contexts. Based on the learned model, CAQ can switch among different gating networks in a context-aware manner and then adopt it to automatically capture the representation importance of various layers for optimal quantization bit-width selection. The experimental results show that CAQ achieves up to 50% storage savings with even 2.61% higher accuracy than the state-of-the-art baselines.
Sicong Liu 0005, Yungang Wu, Bin Guo 0001, Yuzhan Wang, Liyao Xiang, Zhetao Li, Zhiwen Yu 0001
IEEE Internet Things J.6
2022 Achieving adversarial robustness via sparsity
Ningyi Liao, Shufan Wang, Liyao Xiang, Nanyang Ye 0001, Pengzhi Chu
Mach. Learn.3
2022 Differential Privacy for Tensor-Valued Queries
abstract
Private individual information are increasingly exposed through high-dimensional and high-order data, with the wide deployment of learning techniques. These data are typically expressed in form of tensors, but there is no principled way to guarantee privacy for tensor-valued queries. Conventional differential privacy is typically applied to scalar values without a precise definition on the shape of the queried data. Realizing that the conventional mechanisms do not take the data structural information into account, we proposeTensor Variate Gaussian(TVG), a new$(\epsilon,\delta) $-differential privacy mechanism for tensor-valued queries. We further introduce two mechanisms based on TVG with an improved utility by imposing the unimodal differentially-private noise. With the utility space available, the proposed mechanisms can be instantiated with an optimized utility, and the optimization problem has a closed-form solution scalable to large-scale problems. Finally, we experimentally test our mechanisms on a variety of datasets and models, demonstrating that TVG is superior than other state-of-the-art mechanisms on tensor-valued queries.
Jungang Yang 0002, Liyao Xiang, Rui-dong Chen, Weiting Li, Baochun Li
IEEE Trans. Inf. Forensics Secur.2
2021 Speedup Robust Graph Structure Learning with Low-Rank Information
abstract
Recent studies have shown that graph neural networks (GNNs) are vulnerable to unnoticeable adversarial perturbations, which largely confines their deployment in many safety-critical domains. Robust graph structure learning has been proposed to improve the GNN performance in the face of adversarial attacks. In particular, the low-rank methods are utilized to purify the perturbed graphs. However, these methods are mostly computationally expensive with O(n3) time complexity and O(n2) space complexity. We propose LRGNN, a fast and robust graph structure learning framework, which exploits the low-rank property as prior knowledge to speed up optimization. To eliminate adversarial perturbation, LRGNN decouples the adjacency matrix into a low-rank component and a sparse one, and learns by minimizing the rank of the first part while suppressing the second part. Its sparse variant is formed to reduce the memory footprint further. Experimental results on various attack settings have shown LRGNN acquires comparable robustness with the state-of-the-art much more efficiently, boasting a significant advantage on large-scale graphs.
Hui Xu 0011, Liyao Xiang, Jiahao Yu 0001, Xinbing Wang
CIKM2
2021 Spatiotemporal Graph Neural Network for Traffic Prediction Exploiting Cascading Behavior
abstract
As a critical part of the intelligent transportation system, traffic prediction is challenging due to the time-evolving cascading behavior, i.e., the fluctuation of traffic conditions on one road will affect neighboring roads in the future. To address this issue, we propose a novel learning framework, which is able to extract the most relevant historical information for prediction by capturing the underlying cascading behavior. An encoder-decoder architecture is adopted, where the historical contextual information of each road is encoded into a sequence of historical embeddings. A spatiotemporal attention mechanism is devised to model the cascading behavior in the embedding space so that the most relevant information for prediction is concentrated. Extensive experiments on a real-world large-scale highway dataset verify the effectiveness of our proposed approach, observing 3% ~ 5% improvement over state-of-the-art methods.
Xiaoying Gan, Luoyi Fu, Liyao Xiang, Haiming Jin
GLOBECOM4
2021 Federated Model Search via Reinforcement Learning
abstract
Federated Learning (FL) framework enables training over distributed datasets while keeping the data local. However, it is difficult to customize a model fitting for all unknown local data. A pre-determined model is most likely to lead to slow convergence or low accuracy, especially when the distributed data is non-i.i.d.. To resolve the issue, we propose a model searching method in the federated learning scenario, and the method automatically searches a model structure fitting for the unseen local data. We novelly design a reinforcement learning-based framework that samples and distributes sub-models to the participants and updates its model selection policy by maximizing the reward. In practice, the model search algorithm takes a long time to converge, and hence we adaptively assign sub-models to participants according to the transmission condition. We further propose delay-compensated synchronization to mitigate loss over late updates to facilitate convergence. Extensive experiments show that our federated model search algorithm produces highly accurate models efficiently, particularly on non-i.i.d. data.
Dixi Yao, Lingdong Wang, Liyao Xiang, Yanjun Tong
ICDCS4
2021 Privacy Budgeting for Growing Machine Learning Datasets
abstract
The wide deployment of machine learning (ML) models and service APIs exposes the sensitive training data to untrusted and unknown parties, such as end-users and corporations. It is important to preserve data privacy in the released ML models. An essential issue with today's privacy-preserving ML platforms is a lack of concern on the tradeoff between data privacy and model utility: a private datablock can only be accessed a finite number of times as each access is privacy-leaking. However, it has never been interrogated whether such privacy leaked in the training brings good utility. We propose a differentially-private access control mechanism on the ML platform to assign datablocks to queries. Each datablock arrives at the platform with a privacy budget, which would be consumed at each query access. We aim to make the most use of the data under the privacy budget constraints. In practice, both datablocks and queries arrive continuously so that each access decision has to be made without knowledge about the future. Hence we propose online algorithms with a worst-case performance guarantee. Experiments on a variety of settings show our privacy budgeting scheme yields high utility on ML platforms.
Weiting Li, Liyao Xiang
INFOCOM2
2021 Decentralized Multi-AGV Task Allocation based on Multi-Agent Reinforcement Learning with Information Potential Field Rewards
abstract
Automated Guided Vehicles (AGVs) have been widely used for material handling in flexible shop floors. Each product requires various raw materials to complete the assembly in production process. AGVs are used to realize the automatic handling of raw materials in different locations. Efficient AGVs task allocation strategy can reduce transportation costs and improve distribution efficiency. However, the traditional centralized approaches make high demands on the control center’s computing power and real-time capability. In this paper, we present decentralized solutions to achieve flexible and self-organized AGVs task allocation. In particular, we propose two improved multi-agent reinforcement learning algorithms, MAD-DPG-IPF (Information Potential Field) and BiCNet-IPF, to realize the coordination among AGVs adapting to different scenarios. To address the reward-sparsity issue, we propose a reward shaping strategy based on information potential field, which provides stepwise rewards and implicitly guides the AGVs to different material targets. We conduct experiments under different settings (3 AGVs and 6 AGVs), and the experiment results indicate that, compared with baseline methods, our work obtains up to 47% task response improvement and 22% training iterations reduction.
Bin Guo 0001, Jiangshan Zhang, Jiaqi Liu 0002, Sicong Liu 0005, Zhiwen Yu 0001, Zhetao Li, Liyao Xiang
MASS8
2021 A secure data collection strategy using mobile vehicles joint UAVs in smart city
Qingyong Deng, Shaobo Huang, Zhetao Li, Bin Guo 0001, Liyao Xiang, Rong Ran
Comput. Networks5
2020 Context-Aware Deep Model Compression for Edge Cloud Computing
abstract
While deep neural networks (DNNs) have led to a paradigm shift, its exorbitant computational requirement has always been a roadblock in its deployment to the edge, such as wearable devices and smartphones. Hence a hybrid edge-cloud computational framework is proposed to transfer part of the computation to the cloud, by naively partitioning the DNN operations under the constant network condition assumption. However, real-world network state varies greatly depending on the context, and DNN partitioning only has limited strategy space. In this paper, we explore the structural flexibility of DNN to fit the edge model to varying network contexts and different deployment platforms. Specifically, we designed a reinforcement learning-based decision engine to search for model transformation strategies in response to a combined objective of model accuracy and computation latency. The engine generates a context-aware model tree so that the DNN can decide the model branch to switch to at runtime. By the emulation and field experimental results, our approach enjoys a 30% − 50% latency reduction while retaining the model accuracy.
Lingdong Wang, Liyao Xiang, Jiaju Chen, Dixi Yao, Xinbing Wang, Baochun Li
ICDCS2
2020 Achieving Consensus in Privacy-Preserving Decentralized Learning
abstract
Machine learning algorithms have been widely deployed on decentralized systems so that users with private, local data can jointly contribute to a better generalized model. One promising approach is Aggregation of Teacher Ensembles, which transfers knowledge of locally trained models to a global one without releasing any private data. However, previous methods largely focus on privately aggregating the local results without concerning their validity, which easily leads to erroneous aggregation results especially when data is unbalanced across different users. Hence, we propose a private consensus protocol - which reveals nothing else but the label with the highest votes, in the condition that the number of votes exceeds a given threshold. The purpose is to filter out undesired aggregation results that could hurt the aggregator model performance. Our protocol also guarantees differential privacy such that any adversary with auxiliary information cannot gain any additional knowledge from the results. We show that with our protocol, we achieve the same privacy level with an improved accuracy compared to previous works.
Liyao Xiang, Lingdong Wang, Shufan Wang, Baochun Li
ICDCS1
2020 Interpretable Complex-Valued Neural Networks for Privacy Protection
Liyao Xiang, Hao Zhang 0063, Jie Ren 0018, Quanshi Zhang
ICLR1
2019 Differentially-Private Deep Learning from an optimization Perspective
abstract
With the amount of user data crowdsourced for data mining dramatically increasing, there is an urgent need to protect the privacy of individuals. Differential privacy mechanisms are conventionally adopted to add noise to the user data, so that an adversary is not able to gain any additional knowledge about individuals participating in the crowdsourcing, by inferring from the learned model. However, such protection is usually achieved with significantly degraded learning results. We have observed that the fundamental cause of this problem is that the relationship between model utility and data privacy is not accurately characterized, leading to privacy constraints that are overly strict. In this paper, we address this problem from an optimization perspective, and formulate the problem as one that minimizes the accuracy loss given a set of privacy constraints. We use sensitivity to describe the impact of perturbation noise to the model utility, and propose a new optimized additive noise mechanism that improves overall learning accuracy while conforming to individual privacy constraints. As a highlight of our privacy mechanism, it is highly robust in the high privacy regime (when ∈ → 0), and against any changes in the model structure and experimental settings.
Liyao Xiang, Baochun Li
INFOCOM1
2017 $Tack: $ Learning Towards Contextual and Ephemeral Indoor Localization With Crowdsourcing
abstract
At events, such as conferences, indoor localization is both contextual and ephemeral, in that localization is only needed within the context of and for the duration of the event. As such, the costs and requirements of providing such services need to be minimal. In this paper, we design, implement, and evaluate Tack, a new mobile application framework that is specifically engineered to support such contextual and ephemeral indoor localization during an event. To provide location-based services with Tack, an event organizer only needs to bring and place a small number of (reusable) beacons around the venue before the event begins. As a system framework, Tack uses a combination of known beacon locations, contacts over bluetooth low energy, crowdsourcing, and dead-reckoning to estimate and refine user locations. To make our location estimates more accurate, we embrace the inherent nature of beacons, design crowdsourcing-based inference algorithms, and present an extensive evaluation by running real-world experiments with iOS devices and beacons. Tack has been implemented as an open-source framework on the iOS platform and can be used by mobile applications designed for events with location-based services.
Liyao Xiang, Tzu-Yin Tai, Baochun Li, Bo Li 0001
IEEE J. Sel. Areas Commun.1
2015 Coalition Formation Towards Energy-Efficient Collaborative Mobile Computing
abstract
With mobile offloading, computation-intensive tasks can be offloaded from mobile devices to the cloud to conserve energy. In principle, the idea is to trade the relatively low communication energy expense for high computation power consumption. In this paper, we propose that computation-intensive tasks can be distributed among nearby mobile devices, and focus on the case that a group of mobile users may collaborate with one another with one common target job. In particular, a user can reduce its own energy consumption by delegating a portion of the job to nearby users in a coalition. We propose distributed collaboration strategies based on game theory, and formulate the problem as a non-transferable utility coalition formation game in which users join or split from coalitions depending on the local preference. The stability of the resulting partition is studied. We show through simulation that the proposed algorithm reduces up to 22% of the average energy costs compared to the non-cooperative case, and the running time scales well as the number of users grows.
Liyao Xiang, Baochun Li, Bo Li 0001
ICCCN1
2014 Ready, Set, Go: Coalesced offloading from mobile devices to the cloud
abstract
With an abundance of computing resources, cloud computing systems have been widely used to elastically offload the execution of computation-intensive applications on mobile devices, leading to performance gains and better power efficiency. However, existing works have so far focused on one application only, and multiple applications are not coordinated when sending their offloading requests to the cloud. In this paper, we propose the new technique of coalesced offloading, which exploits the potential for multiple applications to coordinate their offloading requests with the objective of saving additional energy on mobile devices. The intuition is that, by sending these requests in “bundles,” the period of time that the network interface stays in the high-power state can be reduced. We present two online algorithms, collectively referred to as Ready, Set, Go (RSG), that make near-optimal decisions on how offloading requests from multiple applications are to be best coalesced. We show, both analytically and experimentally using actual smartphones, that RSG is able to achieve additional energy savings while maintaining satisfactory performance.
Liyao Xiang, Shiwen Ye, Yuan Feng 0005, Baochun Li, Bo Li 0001
INFOCOM1
2012 A discriminatory pricing double auction for spectrum allocation
abstract
Cognitive radio is promising in improving spectrum efficiency by enabling unlicensed users to access to the licensed spectrum. Spectrum auction is perceived as a potential way to realize it. Primary users (PUs) act as sellers selling unused spectrum bands and secondary users (SUs) act as buyers who intend to get spectrum bands from PUs. In situations where multiple PUs and SUs exist, double auction is a paradigm to assign spectrum. Efficiency and economic robustness are considered two essential properties in the model. Previous work often employ bid-independent uniform pricing to maintain economic robust at the substantial cost of efficiency. In this paper, we investigate the tradeoff between efficiency and robustness. We propose DIPA, a DIscriminatory Pricing double Auction for spectrum, in which bidders are charged of varying prices for the same item they purchase. We demonstrate that DIPA is robust and improves efficiency largely over the previous design.
Liyao Xiang, Gaofei Sun, Jing Liu 0023, Xinbing Wang
WCNC1