Tao Xiang 0001

dblp:22/4460-1 · DBLP profile ↗
← Back
26ranked-venue papers in the field
3as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 9 (2 first)Data Mining & Knowledge Discovery · 6 (1 first)Database Systems & Data Management · 4Other / Interdisciplinary · 4Information Retrieval & Web Search · 3
YearPublicationVenuePosition
2026 HiFi-WF: Toward Realistic Website Fingerprinting with Multi-tab and Subpage Recognition
abstract
Website Fingerprinting (WF) is an emerging traffic analysis technique that enables a passive adversary to infer which websites a user visits. However, most existing studies, whether in single-tab or multi-tab settings, rely on the unrealistic assumption that users only access website homepages, diverging significantly from real-world browsing behavior. Even recent works extending WF to subpages primarily focus on website-level identification, without distinguishing which specific subpages are visited, thereby limiting the attack's granularity and scope. In this paper, we propose HiFi-WF (Hierarchical Fine-grained Website Fingerprinting), a novel framework that breaks the homepage-only assumption and extends WF to multi-tab recognition and fine-grained subpage identification. We formulate the task as a hierarchical multi-label classification problem, jointly modeling the distinctions and correlations between homepages and subpages. To this end, HiFi-WF integrates a unified CNN-based extractor and layered encoder with a Feature Interaction Module based on multi-head cross-attention to capture inter-level dependencies. An Enhanced SubHead enforces hierarchical constraints to suppress invalid subpage predictions, while a cascaded channel–spatial attention mechanism refines discriminative features for precise hierarchical identification. Experimental results demonstrate that HiFi-WF achieves state-of-the-art performance at both hierarchical levels, attaining F1-scores of 92.1% (homepage) and 81.9% (subpage), thereby validating its effectiveness in advancing WF attacks toward realistic, fine-grained, and multi-tab browsing scenarios. Related codes and datasets can be found in https://github.com/wusongyang02-blip/HiFi-WF.
Chuan Ma 0001, Ming Ding 0001, Long Yuan 0001, Biwen Chen, Yuwen Qian, Tao Xiang 0001
WWW7
2026 HP2: Hybrid and precision-guided filter pruning for CNN compression
Shangwei Guo, Jialing He, Run Wang 0001, Tao Xiang 0001
Inf. Sci.6
2025 Semantic Gaussian Mixture Variational Autoencoder for Sequential Recommendation
Beibei Li 0001, Tao Xiang 0001, Beihong Jin, Yiyuan Zheng
DASFAA (5)2
2025 Beyond Single Tabs: A Transformative Few-Shot Approach to Multi-Tab Website Fingerprinting Attacks
abstract
Website Fingerprinting (WF) attacks allow passive eavesdroppers to deduce the websites a user visits by analyzing encrypted traffic, threatening user privacy. While current WF attacks achieve high accuracy, they typically assume single-tab browsing, which is unrealistic as users often open multiple tabs, creating mixed traffic. Existing multi-tab WF approaches require large datasets and frequent retraining due to evolving website content, limiting their practicality. In this paper, we introduce Few-shot Multi-tab Website Fingerprinting (FMWF), a novel approach designed to address the limitations of existing multi-tab WF attacks. FMWF directly tackles the challenges of mixed, overlapping traffic traces generated from multi-tab browsing, leveraging two key innovations: (1) an advanced data augmentation technique that synthesizes realistic multi-tab traffic sequences from easily collected single-tab traces, thereby dramatically reducing the need for large-scale real-world traffic data; and (2) a powerful fine-tuning algorithm based on transfer learning that adapts pre-trained models to new, multi-tab environments with minimal additional data. This two-stage framework enables FMWF to capture the complex effectively, overlapping traffic patterns inherent in multi-tab browsing while maintaining a high level of flexibility and significantly lowering computational and data collection burdens. Our experiments, conducted using real traffic traces collected from three widely-used browsers-Microsoft Edge, Google Chrome, and Tor Browser-highlight the superior performance of FMWF in both closed-world and open-world scenarios. Notably, FMWF achieves a minimum 12.3% improvement in accuracy compared to ARES (SP'23) [7], TMWF (CCS'23) [13], and BAPM (ACSAC'21) [10] in the open-world scenario. The code with related datasets is available at https://github.com/WW-Meng/FMWF.
Wenwen Meng, Chuan Ma 0001, Ming Ding 0001, Chunpeng Ge 0001, Yuwen Qian, Tao Xiang 0001
WWW6
2025 Model Supply Chain Poisoning: Backdooring Pre-trained Models via Embedding Indistinguishability
abstract
Pre-trained models (PTMs) are widely adopted across various downstream tasks in the machine learning supply chain. Adopting untrustworthy PTMs introduces significant security risks, where adversaries can poison the model supply chain by embedding hidden malicious behaviors (backdoors) into PTMs. However, existing backdoor attacks to PTMs can only achieve partially task-agnostic and the embedded backdoors are easily erased during the fine-tuning process. This makes it challenging for the backdoors to persist and propagate through the supply chain. In this paper, we propose a novel and severer backdoor attack, TransTroj, which enables the backdoors embedded in PTMs to efficiently transfer in the model supply chain. In particular, we first formalize this attack as an indistinguishability problem between poisoned and clean samples in the embedding space. We decompose embedding indistinguishability into pre- and post-indistinguishability, representing the similarity of the poisoned and reference embeddings before and after the attack. Then, we propose a two-stage optimization that separately optimizes triggers and victim PTMs to achieve embedding indistinguishability. We evaluate TransTroj on four PTMs and six downstream tasks. Experimental results show that our method significantly outperforms SOTA task-agnostic backdoor attacks -- achieving nearly 100% attack success rate on most downstream tasks -- and demonstrates robustness under various system settings. Our findings underscore the urgent need to secure the model supply chain against such transferable backdoor attacks. The code is available at https://github.com/haowang-cqu/TransTroj
Hao Wang 0227, Shangwei Guo, Jialing He, Hangcheng Liu, Tianwei Zhang 0004, Tao Xiang 0001
WWW6
2025 Corrigendum: An Unbiased Risk Estimator for Partial Label Learning with Augmented Classes
abstract
This is a corrigendum for the article “An Unbiased Risk Estimator for Partial Label Learning with Augmented Classes” published in ACM Trans. Intell. Syst. Technol. 15(6): 131:1-131:22 (2024).
Senlin Shu, Beibei Li 0001, Tao Xiang 0001, Zhongshi He
ACM Trans. Intell. Syst. Technol.4
2024 Reducing Interaction Noise for Sequential Recommendation via Robust Interests
Yiyuan Zheng, Beihong Jin, Beibei Li 0001, Weijiang Lai, Tao Xiang 0001
DASFAA (3)5
2024 Multi-intent Driven Contrastive Sequential Recommendation
Yiyuan Zheng, Beibei Li 0001, Beihong Jin, Weijiang Lai, Tao Xiang 0001
ECML/PKDD (9)6
2024 NLPSweep: A comprehensive defense scheme for mitigating NLP backdoor attacks
Tao Xiang 0001, Fei Ouyang, Di Zhang 0011, Chunlong Xie, Hao Wang 0227
Inf. Sci.1
2024 An Unbiased Risk Estimator for Partial Label Learning with Augmented Classes
abstract
Partial Label Learning (PLL) is a typical weakly supervised learning task, which assumes each training instance is annotated with a set of candidate labels containing the ground-truth label. Recent PLL methods adopt identification-based disambiguation to alleviate the influence of false positive labels and achieve promising performance. However, they require all classes in the test set to have appeared in the training set, ignoring the fact that new classes will keep emerging in real applications. To address this issue, in this article, we focus on the problem of Partial Label Learning with Augmented Class (PLLAC), where one or more augmented classes are not visible in the training stage but appear in the inference stage. Specifically, we propose an unbiased risk estimator with theoretical guarantees for PLLAC, which estimates the distribution of augmented classes by differentiating the distribution of known classes from unlabeled data and can be equipped with arbitrary PLL loss functions. Besides, we provide a theoretical analysis of the estimation error bound of the estimator, which guarantees the convergence of the empirical risk minimizer to the true risk minimizer as the number of training data tends to infinity. Furthermore, we add a risk-penalty regularization term in the optimization objective to alleviate the influence of the over-fitting issue caused by negative empirical risk. Extensive experiments on benchmark, UCI, and real-world datasets demonstrate the effectiveness of the proposed approach.
Senlin Shu, Beibei Li 0001, Tao Xiang 0001, Zhongshi He
ACM Trans. Intell. Syst. Technol.4
2024 Multiple-Instance Learning from Pairwise Comparison Bags
abstract
Multiple-instance learning (MIL) is a significant weakly supervised learning problem, where the training data consist of bags containing multiple instances and bag-level labels. Most previous MIL research required fully labeled bags. However, collecting such data is challenging due to the labeling costs or privacy concerns. Fortunately, we can easily collect pairwise comparison information, indicating one bag is more likely to be positive than the other. Therefore, we investigate a novel MIL problem about learning a bag-level binary classifier only from pairwise comparison bags. To solve this problem, we display the data generation process and provide a baseline method to train an instance-level classifier based on unlabeled-unlabeled learning. To achieve better performance, we propose a convex formulation to train a bag-level classifier and give a generalization error bound. Comprehensive experiments show that both the baseline method and the convex formulation achieve satisfactory performance, while the convex formulation performs better. 1
Senlin Shu, Haobo Wang 0001, Hongxin Wei, Tao Xiang 0001, Beibei Li 0001
ACM Trans. Intell. Syst. Technol.5
2023 Towards Query-Efficient Black-Box Attacks: A Universal Dual Transferability-Based Framework
abstract
Adversarial attacks have threatened the application of deep neural networks in security-sensitive scenarios. Most existing black-box attacks fool the target model by interacting with it many times and producing global perturbations. However, all pixels are not equally crucial to the target model; thus, indiscriminately treating all pixels will increase query overhead inevitably. In addition, existing black-box attacks take clean samples as start points, which also limits query efficiency. In this article, we propose a novel black-box attack framework, constructed on a strategy of dual transferability (DT), to perturb the discriminative areas of clean examples within limited queries. The first kind of transferability is the transferability of model interpretations. Based on this property, we identify the discriminative areas of clean samples for generating local perturbations. The second is the transferability of adversarial examples, which helps us to produce local pre-perturbations for further improving query efficiency. We achieve the two kinds of transferability through an independent auxiliary model and do not incur extra query overhead. After identifying discriminative areas and generating pre-perturbations, we use the pre-perturbed samples as better start points and further perturb them locally in a black-box manner to search the corresponding adversarial examples. The DT strategy is general; thus, the proposed framework can be applied to different types of black-box attacks. We conduct extensive experiments to show that, under various system settings, our framework can significantly improve the query efficiency of existing black-box attacks and attack success rates.
Tao Xiang 0001, Hangcheng Liu, Shangwei Guo, Yan Gan, Wenjian He, Xiaofeng Liao 0001
ACM Trans. Intell. Syst. Technol.1
2023 Multiple-Instance Learning From Unlabeled Bags With Pairwise Similarity
abstract
Inmultiple-instance learning(MIL), each training example is represented by a bag of instances. A training bag is either negative if it contains no positive instances or positive if it has at least one positive instance. Previous MIL methods generally assume that training bags are fully labeled. However, the exact labels of training examples may not be accessible, due to security, confidentiality, and privacy concerns. Fortunately, it could be easier for us to access the pairwise similarity between two bags (indicating whether two bags share the same label or not) and unlabeled bags, as we do not need to know the underlying label of each bag. In this paper, we provide the first attempt to investigate MIL from only similar-dissimilar-unlabeled bags. To solve this new MIL problem, we first propose a strong baseline method that trains an instance-level classifier by employing an unlabeled-unlabeled learning strategy. Then, we also propose to train a bag-level classifier based on a convex formulation and theoretically derive a generalization error bound for this method. Comprehensive experimental results show that our instance-level classifier works well, while our bag-level classifier even has better performance.
Lei Feng 0006, Senlin Shu, Yuzhou Cao, Lue Tao, Hongxin Wei, Tao Xiang 0001, Bo An 0001, Gang Niu 0001
IEEE Trans. Knowl. Data Eng.6
2022 ELAA: An efficient local adversarial attack using model interpreters
abstract
Modern deep neural networks are highly vulnerable to adversarial examples, which attracts more and more researchers' attention to craft powerful adversarial examples. Most of these generation algorithms create global perturbations that would affect the visual quality of adversarial examples. To mitigate such drawbacks, some attacks attempt to generate local perturbations. However, existing local adversarial attacks are time-consuming and the generated adversarial examples are still distinguishable from clean images. In this paper, we propose a novel efficient local adversarial attack (ELAA) using model interpreters to generate severe local perturbations and improve the imperceptibly of the generated adversarial examples. Specifically, we take advantage of model interpretation methods to search the discriminative regions of clean images. Then, we generate local adversarial examples by adding masks to original clean images. We also propose a new optimization method to reduce the redundancy of local perturbations. Through extensive experiments, we show our ELAA can maintain a high attack ability while preserving the visual quality of clean images. Experimental results also demonstrate our local attack outperforms state-of-the-art local attack methods under various system settings.
Shangwei Guo, Siyuan Geng, Tao Xiang 0001, Hangcheng Liu, Ruitao Hou
Int. J. Intell. Syst.3
2021 Multiple-Instance Learning from Similar and Dissimilar Bags
abstract
Multiple-instance learning (MIL) is an important weakly supervised binary classification problem, where training instances are arranged in bags, and each bag is assigned a positive or negative label. Most of the previous studies for MIL assume that training bags are fully labeled. However, in some real-world scenarios, it could be difficult to collect fully labeled bags, due to the expensive time and labor consumption of the labeling task. Fortunately, it could be much easier for us to collect similar and dissimilar bags (indicating whether two bags share the same label or not), because we do not need to figure out the underlying label of each bag in this case. Therefore, in this paper, we for the first time investigate MIL from only similar and dissimilar bags. To solve this new MIL problem, we propose a convex formulation to train a bag-level classifier based on empirical risk minimization and theoretically derive a generalization error bound. In addition, we also propose a strong baseline for this new MIL problem, which aims to train an instance-level classifier by minimizing the instance-level empirical risk. Extensive experimental results clearly demonstrate that our proposed baseline works well, while our proposed convex formulation is even better.
Lei Feng 0006, Senlin Shu, Yuzhou Cao, Lue Tao, Hongxin Wei, Tao Xiang 0001, Bo An 0001, Gang Niu 0001
KDD6
2021 A novel hybrid augmented loss discriminator for text-to-image synthesis
abstract
For the text-to-image synthesis task, most discriminators in existing generative adversarial networks based methods tend to fall into a local suboptimal state too early in the training process, resulting in the poor quality of generated images. To address the above problems, a hybrid augmented loss discriminator is designed. In this designed discriminator, to reduce the sensitivity of the discriminator classification recognition, make it pay attention to the semantic and structural changes, we add the loss value of the fake sample to the loss value of the real sample. Moreover, to indirectly guide the generator to generate samples, the loss value of the real sample is added to the fake sample. The loss value mixed with real and fake samples actually augments signal transmission. It perturbs parameter update of the discriminator during optimization and prevents the discriminator from falling into the local suboptimal state prematurely. Whereafter, we apply the proposed discriminator to two kinds of text-to-image synthesis tasks. Experimental results show that the proposed method can help the baseline models to improve performance.
Yan Gan, Mao Ye 0001, Shangming Yang, Tao Xiang 0001
Int. J. Intell. Syst.5
2021 Exploring the redaction mechanisms of mutable blockchains: A comprehensive survey
abstract
Blockchain technology has attracted tremendous interest from both industry and academia. It is typically used to record a public history of transactions (e.g., payment/smart contract data), but storing nonpayment/contract data in transactions has been common. The ability to store data unrelated to payment/contract such as illicit data on blockchain may be abused for malicious purposes. For example, one may use blockchain to store the data related to child pornography and copyright violations, which are publicly visible and immutable. Moreover, an immutable blockchain is not suitable for all blockchain-based applications. So far, numerous redaction mechanisms for the mutable blockchain have been developed. In this paper, we aim at conducting a comprehensive survey that reviews and analyzes the state-of-the-art redaction mechanisms. We start by giving a general presentation of blockchain and summarize the typical methods of inserting data in blockchain. Next, we discuss the challenges of designing the redaction mechanism and propose a list of evaluation criteria. Then, redaction mechanisms of the existing mutable blockchains are systemically reviewed and analyzed based on our evaluation criteria. The analyses include algorithmic overviews, performance limitations, and security vulnerabilities. Finally, the comparisons and analyses provide new insights into these mechanisms. This survey will provide developers and researchers a comprehensive view and facilitate the design of future mutable blockchains.
Di Zhang 0011, Junqing Le, Tao Xiang 0001, Xiaofeng Liao 0001
Int. J. Intell. Syst.4
2021 Access control encryption without sanitizers for Internet of Energy
Tao Xiang 0001, Xiaoguo Li, Hong Xiang
Inf. Sci.2
2020 A training-integrity privacy-preserving federated learning scheme with trusted execution environment
Tong Li 0011, Tao Xiang 0001, Zheli Liu, Jin Li 0002
Inf. Sci.4
2020 An efficient blockchain-based privacy preserving scheme for vehicular social networks
Yuwen Pu, Tao Xiang 0001, Chunqiang Hu, Arwa Alrawais, Hongyang Yan
Inf. Sci.2
2019 ImageProof: Enabling Authentication for Large-Scale Image Retrieval
abstract
With the explosive growth of online images and the popularity of search engines, a great demand has arisen for small and medium-sized enterprises to build and outsource large-scale image retrieval systems to cloud platforms. While reducing storage and retrieval burdens, enterprises are at risk of facing untrusted cloud service providers. In this paper, we take the first step in studying the problem of query authentication for large-scale image retrieval. Due to the large size of image files, the main challenges are to (i) design efficient authenticated data structures (ADSs) and (ii) balance search, communication, and verification complexities. To address these challenges, we propose two novel ADSs, the Merkle randomized k-d tree and the Merkle inverted index with cuckoo filters, to ensure the integrity of query results in each step of image retrieval. For each ADS, we develop corresponding search and verification algorithms on the basis of a series of systemic design strategies. Furthermore, we put together the ADSs and algorithms to design the final authentication scheme for image retrieval, which we name ImageProof. We also propose several optimization techniques to improve the performance of the proposed ImageProof scheme. Security analysis and extensive experiments are performed to show the robustness and efficiency of ImageProof.
Shangwei Guo, Jianliang Xu, Ce Zhang 0007, Cheng Xu 0004, Tao Xiang 0001
ICDE5
2018 Efficient biometric identity-based encryption
Xiaoguo Li, Tao Xiang 0001, Fei Chen 0003, Shangwei Guo
Inf. Sci.2
2017 A Compressive Sensing based privacy preserving outsourcing of image storage and identity authentication service in cloud
Guiqiang Hu, Di Xiao 0001, Tao Xiang 0001, Sen Bai, Yushu Zhang 0001
Inf. Sci.3
2016 Processing secure, verifiable and efficient SQL over outsourced database
Tao Xiang 0001, Xiaoguo Li, Fei Chen 0003, Shangwei Guo, Yuanyuan Yang 0001
Inf. Sci.1
2013 Independent spanning trees in crossed cubes
Yan-Hong Zhang, Tao Xiang 0001
Inf. Process. Lett.3
2011 Security analysis of the public key algorithm based on Chebyshev polynomials over the integer ring ZN
Fei Chen 0003, Xiaofeng Liao 0001, Tao Xiang 0001, Hongying Zheng
Inf. Sci.3