Xiaojun Chen 0004

dblp:20/3215-4 · DBLP profile ↗
← Back
64ranked-venue papers
1as first author
44since 2021 · last 2026
0000-0003-0362-847XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 1 first-author · 21 since 2021Security and privacy · 15 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 12 since 2021Computer networks · 9 · 4 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ARGH-Mark: Anchor-Synchronized Watermarking with Hamming Correction for Robust and Quality-Preserving LLM Attribution
abstract
The proliferation of large language models has intensified demands for reliable content attribution, yet existing watermarking techniques face a fundamental trilemma: they cannot simultaneously optimize for robustness against attacks, minimal text quality degradation, and detection efficiency. To resolve this challenge, we propose ARGH-Mark, a novel watermarking framework that integrates three synergistic innovations: (1) Anchor-synchronized phase recovery for maintaining detection integrity under insertion/deletion attacks, (2) RG-balanced vocabulary modulation that dynamically partitions lexicons via contextual hashing to preserve generation quality, and (3) Hamming-based error correction enabling single-bit error rectification through algebraic coding. Comprehensive evaluations across question answering (ELI5), summarization (CNN/DailyMail), and text generation (C4) demonstrate state-of-the-art performance: the proposed ARGH-Mark framework achieves near-perfect match rate and bit accuracy across diverse configurations, while preserving the quality of the generated text. It significantly reduces detection latency, enabling real-time extraction, and maintains high robustness against token tampering attacks through integrated Hamming error correction, ensuring reliable attribution in adversarial settings. ARGH-Mark achieves a new Pareto frontier in the watermarking design space and advances trustworthy deployment of generative AI in alignment-critical applications.
He Li 0010, Xiaojun Chen 0004, Jingcheng He, Zhendong Zhao, Shuguang Yuan 0003, Yunfei Yang 0001
AAAI2
2026 DeepTracer: Tracing Stolen Model via Deep Coupled Watermarks
abstract
Model watermarking techniques can embed watermark information into the protected model for ownership declaration by constructing specific input-output pairs. However, existing watermarks are easily removed when facing model stealing attacks, and make it difficult for model owners to effectively verify the copyright of stolen models. In this paper, we analyze the root cause of the failure of current watermarking methods under model stealing scenarios and then explore potential solutions. Specifically, we introduce a robust watermarking framework, DeepTracer, which leverages a novel watermark samples construction method and a same-class coupling loss constraint. DeepTracer can incur a high-coupling model between watermark task and primary task that makes adversaries inevitably learn the hidden watermark task when stealing the primary task functionality. Furthermore, we propose an effective watermark samples filtering mechanism that elaborately select watermark key samples used in model ownership verification to enhance the reliability of watermarks. Extensive experiments across multiple datasets and models demonstrate that our method surpasses existing approaches in defending against various model stealing attacks, as well as watermark attacks, and achieves new state-of-the-art effectiveness and robustness.
Yunfei Yang 0001, Xiaojun Chen 0004, Yuexin Xuan, Zhendong Zhao, He Li 0010
AAAI2
2026 Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation
abstract
Generative vision-language models like Stable Diffusion demonstrate remarkable capabilities in creative media synthesis, but they also pose substantial risks of producing unsafe, offensive, or culturally inappropriate content when prompted adversarially. Current defenses struggle to align outputs with human values without sacrificing generation quality or incurring high costs. To address these challenges, we introduce VALOR (Value-Aligned LLM-Overseen Rewriter), a modular, zero-shot agentic framework for safer and more helpful text-to-image generation. VALOR integrates layered prompt analysis with human-aligned value reasoning: a multi-level NSFW detector filters lexical and semantic risks; a cultural value alignment module identifies violations of social norms, legality, and representational ethics; and an intention disambiguator detects subtle or indirect unsafe implications. When unsafe content is detected, prompts are selectively rewritten by a large language model under dynamic, role-specific instructions designed to preserve user intent while enforcing alignment. If the generated image still fails a safety check, VALOR optionally performs a stylistic regeneration to steer the output toward a safer visual domain without altering core semantics. Experiments across adversarial, ambiguous, and value-sensitive prompts show that VALOR significantly reduces unsafe outputs by up to 100.00% while preserving prompt usefulness and creativity. These results highlight VALOR as a scalable and effective approach for deploying safe, aligned, and helpful image generation systems in open-world settings.
Xiaojun Chen 0004, Bingshan Liu, Zeyao Liu, Zhendong Zhao, Xiaoyan Gu 0001
AAAI2
2026 Buster: Implanting Semantic Backdoor Into Text Encoder to Mitigate NSFW Content Generation
Xiaojun Chen 0004, Yuexin Xuan, Zhendong Zhao, Xinfeng Li, Xiaojun Jia, Xiaofeng Wang 0001
DASFAA (5)2
2026 ComMark: Covert and Robust Black-Box Model Watermarking with Compressed Samples
abstract
The rapid advancement of deep learning has turned models into highly valuable assets due to their reliance on massive data and costly training processes. However, these models are increasingly vulnerable to leakage and theft, highlighting the critical need for robust intellectual property protection. Model watermarking has emerged as an effective solution, with black-box watermarking gaining significant attention for its practicality and flexibility. Nonetheless, existing black-box methods often fail to better balance covertness (hiding the watermark to prevent detection and forgery) and robustness (ensuring the watermark resists removal)—two essential properties for real-world copyright verification. In this paper, we propose ComMark, a novel black-box model watermarking framework that leverages frequency-domain transformations to generate compressed, covert, and attack-resistant watermark samples by filtering out high-frequency information. To further enhance watermark robustness, our method incorporates simulated attack scenarios and a similarity loss during training. Comprehensive evaluations across diverse datasets and architectures demonstrate that ComMark achieves state-of-the-art performance in both covertness and robustness.
Yunfei Yang 0001, Xiaojun Chen 0004, Zhendong Zhao, Yu Zhou 0015, Xiaoyan Gu 0001, Juan Cao 0001
ICMR2
2026 JAGUAR: efficient and secure unbalanced PSI under malicious adversaries in the client-server setting
abstract
Abstract In many unbalanced private set intersection (uPSI) applications of the client–server setting, the server needs to perform uPSI with multiple clients. Cong et al. (ACM CCS’21) proposed a state-of-the-art (SOTA) uPSI protocol based on fully homomorphic encryption (FHE), achieving malicious security by employing an oblivious pseudorandom function (OPRF) in the pre-processing phase. However, re-executing existing uPSI protocols with each client imposes significant computational overhead for the server. In this paper, we present JAGUAR, a maliciously secure and efficient uPSI protocol designed for this setting. JAGUAR reduces online computation through a Divide-and-Combine optimization, requiring only $${\mathcal {O}}(\sqrt{|X|})$$ O ( | X | ) homomorphic multiplications. Furthermore, it employs a novel fixed VOLE-based OPRF that enables reusable and lightweight pre-processing across multiple clients. Experimental results demonstrate that JAGUAR achieves up to $$2.7\times$$ 2.7 × improvement in online runtime compared to the SOTA protocol in LAN. In multi-client scenarios, JAGUAR further outperforms existing protocols by a wide margin in terms of scalability and overall performance.
Weizhan Jing, Xiaojun Chen 0004, Ye Dong, Qiang Liu 0060, Tingyu Fan
Cybersecur.2
2026 Personalized Subgraph Federated Learning With Decoupled Data Heterogeneity in Mobile-Edge Computing
abstract
Graph Federated Learning (FL) has attracted extensive attention in recent years due to its ability to train global graph models in a distributed manner without exposing original local data. However, fine-grained data heterogeneity remains largely overlooked in collaborative graph model training. Existing graph FL methods that address heterogeneity are mostly adapted from traditional FL and fail to account for the unique complexity of graph-specific heterogeneity. Specifically, graph heterogeneity can be further decomposed into feature heterogeneity and structural heterogeneity, which are tightly coupled during local training. To address this issue, we propose a novel local graph module, Feature and Structure Decoupling Convolution (FSD-Conv), designed to disentangle the interplay between feature bias and structural bias. With FSD-Conv, clients can learn feature-related yet structure-unbiased representations, thereby alleviating the adverse impact of graph heterogeneity in federated training. Furthermore, we introduce FedFSD, a personalized graph FL framework that achieves effective personalized model aggregation through an explainable neural network operating in a low-dimensional space. Extensive experiments on six graph datasets under both disjoint and overlapping client partitioning schemes demonstrate the effectiveness of FedFSD in handling complex graph data heterogeneity.
Bisheng Tang, Xiaojun Chen 0004, Shaopu Wang, Yuexin Xuan, Zhendong Zhao, Xingyu Gao 0001
IEEE Trans. Mob. Comput.2
2025 Antelope: Potent and Concealed Jailbreak Attack Strategy
abstract
Due to the remarkable generative potential of diffusion-based models, numerous researches have investigated jailbreak attacks targeting these frameworks. A particularly concerning threat within image models is the generation of Not-Safe-for-Work (NSFW) content. Despite the implementation of security filters, numerous efforts continue to explore ways to circumvent these safeguards. Current attack methodologies primarily encompass adversarial prompt engineering or concept obfuscation, yet they frequently suffer from slow search efficiency, conspicuous attack characteristics and poor alignment with targets. To overcome these challenges, we propose Antelope, a more robust and covert jailbreak attack strategy designed to expose security vulnerabilities inherent in generative models. Specifically, Antelope leverages the confusion of sensitive concepts with similar ones, facilitates searches in the semantically adjacent text space of these related concepts and aligns them with the target imagery, thereby generating sensitive images that are consistent with the target and capable of evading detection. Besides, we successfully exploit the transferability of model-based attacks to penetrate online black-box services. Experimental evaluations demonstrate that Antelope outperforms existing baselines across multiple defensive mechanisms, underscoring its efficacy and versatility.
Xiaojun Chen 0004
CIKM2
2025 Take Attention Inside: Neighbor Pair Graph Contrastive Learning
abstract
Graph Contrastive Learning(GCL) is a fundamental pretraining research method in Graph Neural Networks (GNNs), which puts rich graph-level insights into the graph data to augment the data representation. However, since the existing GCLs generally regard the intra-layer node as negative samples, they cannot cope with the diverse coupled neighbor relationships, which can be pre-trained with the combination of negative and positive samples. Coupled relationships can keep the attribute preference in node-level contrast and correctly pass this preference into the downstream tasks. To further prove the effectiveness of coupled neighbor relationships in the pretraining phase, we propose a novel GNN pretraining model Neighbor Pair Contrastive Graph Siamese Networks (NPC-GSN) for graph contrast. NPC-GSN expands the dissimilar neighbor’s representation discrepancy and decreases the representation discrepancy of similar neighbors in the pretraining phase, aiming to promote downstream node classification. Our extensive experiments on five graph datasets against several pretraining GNN models demonstrate the competitive effectiveness of NPC-GSN in node classification, and the frequency domain and ablation experiments also verify the effectiveness of NPC-GSN.
Bisheng Tang, Xiaojun Chen 0004, Shaopu Wang, Yuexin Xuan, Zhendong Zhao
ICASSP2
2025 Model-Guardian: Protecting against Data-Free Model Stealing Using Gradient Representations and Deceptive Predictions
abstract
Model stealing attack is increasingly threatening the confidentiality of machine learning models deployed in the cloud. Recent studies reveal that adversaries can exploit data synthesis techniques to steal machine learning models even in scenarios devoid of real data, leading to data-free model stealing attacks. Existing defenses against such attacks suffer from limitations, including poor effectiveness, insufficient generalization ability, and low comprehensiveness. In response, this paper introduces a novel defense framework named Model-Guardian. Comprising two components, Data-Free Model Stealing Detector (DFMS-Detector) and Deceptive Predictions (DPreds), Model-Guardian is designed to address the shortcomings of current defenses with the help of the artifact properties of synthetic samples and gradient representations of samples. Extensive experiments on seven prevalent data-free model stealing attacks showcase the effectiveness and superior generalization ability of Model-Guardian, outperforming eleven defense methods and establishing a new state-of-the-art performance. Notably, this work pioneers the utilization of various GANs and diffusion models for generating highly realistic query samples in attacks, with Model-Guardian demonstrating accurate detection capabilities.
Yunfei Yang 0001, Xiaojun Chen 0004, Yuexin Xuan, Zhendong Zhao
ICME2
2025 ColorFP: Improving AI-Generated Text Detection via Fixed Vocabulary Partitioning and Half-Bit Fingerprinting
He Li 0010, Xiaojun Chen 0004, Yunfei Yang 0001, Zhendong Zhao, Shuguang Yuan 0003
PRICAI (4)2
2025 Comet: Accelerating Private Inference for Large Language Model by Predicting Activation Sparsity
abstract
With the growing use of large language models (LLMs) hosted on cloud platforms to offer inference services, privacy concerns about the potential leakage of sensitive information are escalating. Secure Multi-Party Computation (MPC) is a promising solution to protect the privacy in LLM inference. However, MPC requires frequent inter-server communication, causing high performance overhead. Inspired by the prevalent activation sparsity of LLMs, where most neuron are not activated after non-linear activation functions, we propose an efficient private inference system, Comet. This system employs an accurate and fast predictor to predict the sparsity distribution of activation function output. Additionally, we introduce a new private inference protocol. It efficiently and securely avoids computations involving zero values by exploiting the spatial locality of the predicted sparsity distribution. While this computation-avoidance approach impacts the spatiotemporal continuity of KV cache entries, we address this challenge with a low-communication overhead cache refilling strategy that merges miss requests and incorporates a prefetching mechanism. Finally, we evaluate Comet on four common LLMs and compare it with six state-of-the-art private inference systems. Comet achieves a$1.87\times-2.63\times$speedup and a$1.94\times-2.64\times$communication reduction.
Guang Yan, Yuhui Zhang 0011, Zimu Guo, Lutan Zhao, Xiaojun Chen 0004, Wenhao Wang 0001, Dan Meng 0002, Rui Hou 0001
SP5
2025 VCR: Fast Private Set Intersection with Improved VOLE and CRT-Batching
abstract
Private set intersection (PSI) allows two participants to compute the intersection of their private sets without revealing any additional information beyond the intersection itself. It is known that oblivious linear evaluation (OLE) can be used to construct the online efficient PSI protocol. However, oblivious transfer (OT) and fully homomorphic encryption (FHE)-based offline OLE generation are expensive, and the online computational complexity is super-linear and still a heavy burden for large-scale sets. In this paper, we propose VCR, an efficient PSI protocol from vector OLE (VOLE) with the offline-online paradigm. Concretely, we first propose the batched short VOLE protocol to reduce offline overhead for generating VOLE tuples. Then, we design a batched private membership test protocol from pre-computed VOLE to accelerate the online computation. Experiments demonstrate that VCR outperforms prior art. Compared to state-of-the-art work, we reduce the total communication costs (resp. running time) by 341× and 9.1× (resp. 6.5× and 2.5×) on average for OT and FHE-based protocols.
Weizhan Jing, Xiaojun Chen 0004, Ye Dong, Yaxi Yang, Qiang Liu 0060
TrustCom2
2025 FedShelter: Efficient privacy-preserving federated learning with poisoning resistance for resource-constrained IoT network
Tingyu Fan, Xiaojun Chen 0004, Ye Dong, Weizhan Jing, Zhendong Zhao
Comput. Networks2
2025 MD-SONIC: Maliciously-Secure Outsourcing Neural Network Inference With Reduced Online Communication
abstract
With the widespread deployment of Deep-Learning-as-a-Service, secure multi-party computation-based outsourcing neural network (NN) inference has garnered significant attention for its high-security guarantee. Nevertheless, under the dishonest-majority setting with malicious adversaries, prior secure inference works are still costly in terms of communication and run-time. Additionally, existing outsourcing frameworks impose a substantial client-side design, which leads to obstacles in resource-constrained devices. To address the above challenges, we propose MD-SONIC, an online efficient and maliciously-secure framework for outsourcing NN inference with a dishonest majority. We first construct communication-efficient n-party protocols for the basic primitives such as fixed-point multiplication and most significant bit extraction by combining mask-sharing and TinyOT-sharing with SPD$\mathbb {Z}_{2^{k}}$seamlessly. Then, we build fast secure blocks for the widely used NN operators, including matrix multiplication, ReLU, and Maxpool, on top of our basic primitives. To enable an arbitrary number of users to outsource the secure inference task to n computing servers, we propose a lightweight-client and fast$\Sigma $paradigm named SPIN, stemming from zero-knowledge proofs. Our SPIN can be instantiated into a set of efficient outsourcing protocols over multiple algebraic structures (e.g., finite field and ring). We also conduct extensive evaluations of MD-SONIC on various neural networks. Compared to the work by Damgård et al. (IEEE S&P’19) and MD-ML (USENIX Security’24), we achieve up to$594.4\times $and$45.1\times $online communication improvements, and improve the online execution time by at most$14.3\times $(resp.$20.5\times $) and$1.8\times $(resp.$2.3\times $) in LAN (resp. WAN).
Xiaojun Chen 0004, Ye Dong, Rui Hou 0001, Qiang Liu 0060
IEEE Trans. Inf. Forensics Secur.2
2024 Lightweight Secure Aggregation for Personalized Federated Learning with Backdoor Resistance
abstract
Existing federated learning (FL) systems are highly vulnerable in terms of security and privacy due to their distributed architecture, facing poisoning attacks and inference attacks from adversaries. Some prior works have combined poisoning defenses with cryptographic tools: Secure Multi-Party Computation, Zero-Knowledge Proof, and Homomorphic Encryption to propose robust secure aggregation methods that provide security and privacy preservation for FL. Recently, Qin et al. (KDD’23) demonstrate that personalized federated learning (pFL) can effectively resist backdoor injection in poisoning attacks. In this paper, we analyze that as the number of malicious attackers increases, pFL remains vulnerable to backdoor attacks. Moreover, we reveal that current robust secure aggregation methods fail to offer efficient and robust backdoor defense for pFL. Therefore, we propose FLIGHT, a robust secure aggregation method for pFL. It implements a lightweight backdoor detection through a two-stage personalized defense mechanism and ensures privacy preservation using communication-efficient two-party secure computation (2PC) protocols. Extensive experiments on diverse datasets and neural networks validate that FLIGHT decreases run-time up to 64× compared by prior work RoFL (S&P’23), and 42× compared to FLAME (USENIX Security’22).
Tingyu Fan, Xiaojun Chen 0004, Ye Dong, Yuexin Xuan, Weizhan Jing
ACSAC2
2024 CipherDM: Secure Three-Party Inference for Diffusion Model Sampling
Xiaojun Chen 0004, He Li 0010, Tingyu Fan, Zhendong Zhao
ECCV (71)2
2024 Roger: A Round Optimized GPU-Friendly Secure Inference Framework
abstract
Secure neural network inference provides a promising solution to preserve the privacy of Deep Learning as a Service (DLaaS), but its substantial communication and computation overhead remain challenging. Recent works such as GForce [1] and Piranha [2] have introduced GPU-friendly secure inference protocols with improved computation efficiency, yet these approaches are either limited to supporting specialized-trained networks or expensive in communication. As a consequence, there remain potential improvements in functionalities and communication efficiency. To address the above challenges, we introduce Roger, a two-party secure inference framework with semi-honest security, designed to support general neural network inference with a reduced number of round complexity. Drawing inspiration from ABY2.0 [3], we propose the Partial-Fix technology, which fixes the share of one participant during the offline phase to improve its computation efficiency. Then, an online communication-free protocol for secure linear layer computation and a constant-round secure comparison protocol are proposed upon Partial-Fix. Implemented on top of Piranha, the experiments demonstrate that for the CIFAR10 dataset, a single inference on VGG16 requires only 0.40 seconds. In comparison to GForce (resp. Piranha), Roger at least achieves 1.20× (resp. 1.94×) improvement in LAN setting in terms of throughput.
Xiaojun Chen 0004, Ye Dong, Weizhan Jing, Tingyu Fan
ICC2
2024 Comet: Communication-Efficient Batch Secure Three-Party Neural Network Inference with Client-Aiding
abstract
Secure neural network inference enables server (model provider) and client to perform neural network inference without leaking their private inputs. Existing SOTA three-party computation (3PC) inference works emerge challenges on two fronts: i) GPU-accelerated CryptGPU (S&P'21) and P-FALCON (USENIX Security'22) face challenges related to high communication overhead. ii) communication-efficient Meteor(www'23) raises more computation burden and GPU memory usage. These challenges result in lower efficiency when handling large-scale batch inference requests on resource-constrained devices. In this work, we propose Comet,a communication-efficient batch secure three-party inference framework with client-aiding, which achieves semi-honest security in honest majority without collusion between the client and the servers. First, we propose client-aided sharing semantics, which leverages client-generated random values to enhance online communication efficiency. We also design efficient 3PC protocols for neural network operators based on GPU, improving the computational efficiency of both linear and nonlinear layers. Furthermore, we address the tradeoff between communication cost and GPU memory utilization, surpassing SOTA by 1.3-1.9× in communication, 1.5-3.8× in runtime on large-scale batch inference tasks.
Tingyu Fan, Xiaojun Chen 0004, Ye Dong, Weizhan Jing
ICC2
2024 DualCOS: Query-Efficient Data-Free Model Stealing with Dual Clone Networks and Optimal Samples
abstract
Although data-free model stealing attacks are free from reliance on real data, they suffer from limitations, including low accuracy and high query budgets, which restrict their practical feasibility. In this paper, we propose a novel data-free model stealing framework called DualCOS. As a whole, DualCOS is divided into two stages: interactive training and semi-supervised boosting. To optimize the usage of query budgets, we use a dual clone model architecture to address the challenge of querying victim model during generator training. We also introduce active learning-based sampling strategy and sample reuse mechanism to achieve an efficient query process. Furthermore, once query budget is exhausted, the semi-supervised boosting is employed to continue improving the final clone accuracy. Through extensive evaluations, we demonstrate the superiority of our proposed method in terms of accuracy and query efficiency, particularly in scenarios involving hard labels and multiple classes.
Yunfei Yang 0001, Xiaojun Chen 0004, Yuexin Xuan, Zhendong Zhao
ICME2
2024 Progtuning: Progressive Fine-Tuning Framework for Transformer-Based Language Models
Xiaoshuang Ji, Zhendong Zhao, Xiaojun Chen 0004, Zeyao Liu
ICONIP (9)3
2024 STMS: An Out-Of-Distribution Model Stealing Method Based on Causality
abstract
Machine learning, particularly deep learning, is extensively applied in various real-life scenarios. However, recent research has highlighted the severe infringement of privacy and intellectual property caused by model stealing attacks. Therefore, more researchers are dedicated to studying the principles and methods of such attacks to promote the security development of artificial intelligence. Most of the existing model stealing attacks rely on prior information of the attacked models and consider ideal conditions. In order to better understand and defend against model stealing in real-world scenarios, we propose a novel model stealing method, named STMS, based on causal inference learning. For the first time, we introduce the problem of out-of-distribution generalization into the model stealing domain. The proposed approach operates under more challenging conditions, where the training and testing data of the target model are unknown, black-box, hard-label outputs, and there is a distribution shift during the testing phase. STMS achieves comparable or better stealing accuracy and generalization performance than prior works on multiple datasets and tasks. Moreover, this universal framework can be applied to improve the effectiveness of other model stealing methods and can also be migrated to other areas of machine learning.
Yunfei Yang 0001, Xiaojun Chen 0004, Zhendong Zhao, Yuexin Xuan, Bisheng Tang
IJCNN2
2024 An Effective Multiple Private Set Intersection
Qiang Liu 0060, Xiaojun Chen 0004, Weizhan Jing, Ye Dong
SecureComm (1)2
2024 OCE-PTree: An Online Communication Efficient Privacy-Preserving Decision Tree Evaluation
Xiaojun Chen 0004, Weizhan Jing, Tingyu Fan
SecureComm (1)2
2023 Unsupervised Graph Structure-Assisted Personalized Federated Learning
abstract
Non-IID data presents a significant challenge for federated learning(FL), and personalized FL is a natural solution to address this challenge. Recently, Graph Neural Network (GNN) has recently emerged to model the complex client relationship using a client graph to refine personalized models. However, this approach depends on an existing client relation graph on the server, making it impractical unless this prerequisite is satisfied. Furthermore, noisy and missing connections in the original graph structures can degrade personalization performance. In this work, we propose an unsupervised structure learning approach to improve personalized FL, where the server learns a dynamic client graph through self-supervision and generates structure-based client representations. These representations are then broadcasted to users, regulating local training using the learned knowledge as an inductive bias. Empirical studies on benchmark datasets demonstrate the significant effectiveness of our approach and the high quality of the client graphs. The code is available at https://github.com/lazyJane/FedSKA.
Xiaojun Chen 0004, Bisheng Tang, Shaopu Wang, Yuexin Xuan, Zhendong Zhao
ECAI2
2023 Lightweight Reference-Less Summary Quality Evaluation via Key Feature Extraction
Shunan Zang, Jingwen Lin, Xiaojun Chen 0004
ICANN (8)4
2023 Practical and General Backdoor Attacks Against Vertical Federated Learning
Yuexin Xuan, Xiaojun Chen 0004, Zhendong Zhao, Bisheng Tang, Ye Dong
ECML/PKDD (2)2
2023 Meteor: Improved Secure 3-Party Neural Network Inference with Reducing Online Communication Costs
abstract
Secure neural network inference has been a promising solution to private Deep-Learning-as-a-Service, which enables the service provider and user to execute neural network inference without revealing their private inputs. However, the expensive overhead of current schemes is still an obstacle when applied in real applications. In this work, we present Meteor, an online communication-efficient and fast secure 3-party computation neural network inference system aginst semi-honest adversary in honest-majority. The main contributions of Meteor are two-fold: i) We propose a new and improved 3-party secret sharing scheme stemming from the linearity of replicated secret sharing, and design efficient protocols for the basic cryptographic primitives, including linear operations, multiplication, most significant bit extraction, and multiplexer. ii) Furthermore, we build efficient and secure blocks for the widely used neural network operators such as Matrix Multiplication, ReLU, and Maxpool, along with exploiting several specific optimizations for better efficiency. Our total communication with the setup phase is a little larger than SecureNN (PoPETs’19) and Falcon (PoPETs’21), two state-of-the-art solutions, but the gap is not significant when the online phase must be optimized as a priority. Using Meteor, we perform extensive evaluations on various neural networks. Compared to SecureNN and Falcon, we reduce the online communication costs by up to 25.6 × and 1.5 ×, and improve the running-time by at most 9.8 × (resp. 8.1 ×) and 1.5 × (resp. 2.1 ×) in LAN (resp. WAN) for the online inference.
Ye Dong, Xiaojun Chen 0004, Weizhan Jing, Kaiyun Li, Weiping Wang 0005
WWW2
2023 Generalized heterophily graph data augmentation for node classification
Bisheng Tang, Xiaojun Chen 0004, Shaopu Wang, Yuexin Xuan, Zhendong Zhao
Neural Networks2
2023 FlexBNN: Fast Private Binary Neural Network Inference With Flexible Bit-Width
abstract
Advancements in deep learning enable neural network (NN) inference to be a service, but service providers and clients want to keep their inputs secret for privacy protection.Private Inferenceis the task of evaluating NN without leaking private inputs. Existing secure multiparty computation (MPC)-based solutions mainly focus on fixed bit-width methodology, such as 32 and 64 bits. Binary Neural Network (BNN) is efficient when evaluated in MPC and has achieved reasonable accuracy for commonly used datasets, but prior private BNN inference solutions, which focus onBoolean Circuits, are still costly in communication and run-time. In this paper, we introduce FLEXBNN, a fast private BNN inference framework using three-party computation (3PC) inArithmetic Circuitsagainst semi-honest adversaries with honest-majority. In FLEXBNN, we propose to employ flexible and small bit-width equipped with a seamless bit-width conversion method and design several specific optimizations towards the basic operations: i) We propose bit-width determination methods for Matrix Multiplication and Sign-based Activation function. ii) We integrate Batch Normalization and Max-Pooling into the Sign-based Activation function for better efficiency. iii) More importantly, we achieve seamless bit-width conversion within the Sign-based Activation function with no additional cost. Extensive experiments illustrate that FLEXBNN outperforms state-of-the-art solutions in communication, run-time, and scalability. On average, FLEXBNN is 11× faster than XONN (USENIX Security’ 19) in LAN, 46× (resp. 9.3×) faster than QUOTIENT (ACM CCS’19) in LAN (resp. WAN), 10× faster than BANNERS (ACM IH&MMSec’21) in LAN, and 1.1-2.9× (resp. 1.5-2.7×) faster than FALCON (semi-honest, PoPETs’21) in LAN (resp. WAN), and improves the respective communication by 500×, 127×, and 1.3-1.5× compared to XONN, BANNERS, and FALCON.
Ye Dong, Xiaojun Chen 0004, Xiangfu Song, Kaiyun Li
IEEE Trans. Inf. Forensics Secur.2
2022 Noise Suppression with Label Graph in Distantly Supervised Relation Extraction
abstract
Distantly supervised relation extraction suffers from the influence of noise data. Some works solved this issue with relation-aware attention based on multi-instance learning to reduce the weights of noise data in sentence bag, which achieved remarkable results. However, the relationships of label-label and label-sentence in semantic space that can enhance label representation effectively was ignored. The enhanced label representation can keep the representation of noise data that does not contain the corresponding relation away from the label in semantic space. In view of this, we propose a novel method to capture the two relationships above to suppress noise in distantly supervised relation extraction with a label graph. To be specific, the single label of each sentence is first expanded to multi-label and the label graph is built based on the label hierarchy to capture implicit relation among them. Then the relation-aware attention in semantic space is deployed to assign low weights to noise sentences for the connection of label-sentence. Finally, the distantly supervised relation extraction task is optimized as a multi-label multi-classification problem. Experiment results on New York Times indicate that our method shows significantly improvement compared with the strong baseline methods.
Dakui Wang, Yangyang Ding, Xiaojun Chen 0004
CSCWD5
2022 DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints
abstract
Backdoor attack is a type of serious security threat to deep learning models. An adversary can provide users with a model trained on poisoned data to manipulate prediction behavior in test stage using a backdoor. The backdoored models behave normally on clean images, yet can be activated and output incorrect prediction if the input is stamped with a specific trigger pattern. Most existing backdoor attacks focus on manually defining imperceptible triggers in input space without considering the abnormality of triggers' latent representations in the poisoned model. These attacks are susceptible to backdoor detection algorithms and even visual inspection. In this paper, We propose a novel and stealthy backdoor attack - DEFEAT. It poisons the clean data using adaptive imperceptible perturbation and restricts latent representation during training process to strengthen our attack's stealthiness and resistance to defense algorithms. We conduct extensive experiments on multiple image classifiers using real-world datasets to demonstrate that our attack can 1) hold against the state-of-the-art defenses, 2) deceive the victim model with high attack success without jeopardizing model utility, and 3) provide practical stealthiness on image data.
Zhendong Zhao, Xiaojun Chen 0004, Yuexin Xuan, Ye Dong, Dakui Wang, Kaitai Liang
CVPR2
2022 Multi-initial-Center Federated Learning with Data Distribution Similarity-Aware Constraint
Xiaojun Chen 0004, Shaopu Wang, Yangyang Ding, Kaiyun Li
ICA3PP2
2022 CLTS+: A New Chinese Long Text Summarization Dataset with Abstractive Summaries
Shunan Zang, Xiaojun Chen 0004, Yangyang Ding
ICANN (1)4
2022 NEWSFARM: A Large-Scale Chinese Corpus of Long News Summarization
abstract
Recently, the field of natural language processing (NLP) has grown rapidly, driven by massive datasets. At the same time, the need for automatic summarization systems has been rapidly increasing as the amount of textual information on the web and in large data centers became intractable for human readers. However, the lack of large-scale and high-quality Chinese datasets remain a critical bottleneck for further research on automatic text summarization. To close this gap, we searched domestic and foreign Chinese news websites and designed the FCFS (Format and Content Filtering System) to crawl and filter these records to construct NEWSFARM. The NEWSFARM is a large-scale Chinese long news summarization corpus, containing more than 220K Chinese long news and summaries written by professional editors or authors, all of which are released to the public1. We calculated the static metrics and designed many experiments with the baseline models to evaluate the dataset. By comparing with the common datasets, the results not only demonstrate the usefulness and challenges of the proposed corpus for automatic text summarization but also validate the superiority of FCFS.
Shunan Zang, Xiaojun Chen 0004, Peng Zhang 0001
ICPR4
2022 KAFNN: A Knowledge Augmentation Framework to Graph Neural Networks
abstract
The semi-supervised node classification task is a basic problem in graph neural networks(GNNs). GNNs have shown their superiority in graph datasets over traditional neural networks such as Multilayer Perceptron. However, due to the limitation of Weisfeiler-Lehman, the existing GNNs will discard some prior knowledge, which is hard to be coped with, such as Dropout skill, etc. In this paper, we proposed a framework called KAFNN to introduce knowledge discarded obliviously to enhance data representation. KAFNN, based on the Siamese network, introduces the framework of combining GNNs and deep neural networks(DNNs) to capture the data presentation as whole as possible, which will inject more knowledge into GNNs. Extensive experiments based on seven public datasets and seven GNN models have shown that KAFNN has promoted presentation of several state-of-the-art GNN models in a competitive performance.
Bisheng Tang, Xiaojun Chen 0004, Dakui Wang, Zhendong Zhao
IJCNN2
2022 Rethinking the Feature Iteration Process of Graph Convolution Networks
abstract
Node classification is a fundamental research problem in graph neural networks(GNNs), which uses node's feature and label to capture node embedding in a low dimension. The existing graph node classification approaches mainly focus on GNNs from global and local perspectives. The relevant research is relatively insufficient for the micro perspective, which refers to the feature itself. In this paper, we prove that deeper GCNs' features will be updated with the same coefficient in the same dimension, limiting deeper GCNs' expression. To overcome the limits of the deeper GCN model, we propose a zero feature (k-ZF) method to train GCNs. Specifically, k-ZF randomly sets the initial k feature value to zero, acting as a data rectifier and augmenter, and is also a skill equipped with GCNs models and other GCNs skills. Extensive experiments based on three public datasets show that k-ZF significantly improves GCNs in the feature aspect and achieves competitive accuracy.
Bisheng Tang, Xiaojun Chen 0004, Dakui Wang, Zhendong Zhao
IJCNN2
2022 ACTSS: Input Detection Defense against Backdoor Attacks via Activation Subset Scanning
abstract
Deep neural networks are vulnerable to backdoor attacks where adversaries inject the trigger into partial training data to manipulate the trained model misclassification. In addition, the poisoned model behaves normally on clean inputs, and the malicious behavior only occurs when the secret trigger is present, making backdoor attacks hard to be detected. Most existing input detection methods leverage the link between triggers and outputs to reveal the poisoned inputs, which suffer from the trigger-size or the “all-to-all” attack scenario. We show that the internal activations produced by benign and poisoned inputs are significantly different in the poisoned model. In this paper, we propose a novel and run-time input detection algorithm, Activation Subset Scanning (ACTSS), which extracts the activations of incoming inputs and leverages an anomaly detection algorithm to identify malicious inputs. We search and score for the abnormal activation subset according to the statistical difference of activations between benign and poisoned data using nonparametric statistics technology. Extensive experiments are conducted on three public datasets: CIFAR10, GTSRB, and ImageNet, with three representative models. The results verify our approach's effectiveness and state-of-the-art performance, which achieve over 98% false rejection rate for different types of triggers.
Yuexin Xuan, Xiaojun Chen 0004, Zhendong Zhao, Yangyang Ding, Jianming Lv
IJCNN2
2022 PrUE: Distilling Knowledge from Sparse Teacher Networks
Shaopu Wang, Xiaojun Chen 0004, Mengzhen Kou, Jinqiao Shi
ECML/PKDD (3)2
2022 Improved Network Pruning via Similarity-Based Regularization
Shaopu Wang, Jiaxin Zhang 0026, Xiaojun Chen 0004, Jinqiao Shi
PRICAI (2)4
2021 FLOD: Oblivious Defender for Private Byzantine-Robust Federated Learning with Dishonest-Majority
Ye Dong, Xiaojun Chen 0004, Kaiyun Li, Dakui Wang
ESORICS (1)2
2021 Enhancing Label Representations with Relational Inductive Bias Constraint for Fine-Grained Entity Typing
abstract
Fine-Grained Entity Typing (FGET) is a task that aims at classifying an entity mention into a wide range of entity label types. Recent researches improve the task performance by imposing the label-relational inductive bias based on the hierarchy of labels or label co-occurrence graph. However, they usually overlook explicit interactions between instances and labels which may limit the capability of label representations. Therefore, we propose a novel method based on a two-phase graph network for the FGET task to enhance the label representations, via imposing the relational inductive biases of instance-to-label and label-to-label. In the phase 1, instance features will be introduced into label representations to make the label representations more representative. In the phase 2, interactions of labels will capture dependency relationships among them thus make label representations more smooth. During prediction, we introduce a pseudo-label generator for the construction of the two-phase graph. The input instances differ from batch to batch so that the label representations are dynamic. Experiments on three public datasets verify the effectiveness and stability of our proposed method and achieve state-of-the-art results on their testing sets.
Xiaojun Chen 0004, Dakui Wang
IJCAI2
2021 A Differential Privacy Collaborative Deep Learning Algorithm in Pervasive Edge Computing Environment
abstract
With the development of 5G technology and intelligent terminals, the future direction of the Industrial Internet of Things (IIoT) evolution is Pervasive Edge Computing (PEC). In the pervasive edge computing environment, intelligent terminals can perform calculations and data processing. By migrating part of the original cloud computing model's calculations to intelligent terminals, the intelligent terminal can complete model training without uploading local data to a remote server. Pervasive edge computing solves the problem of data islands and is also successfully applied in scenarios such as vehicle interconnection and video surveillance. However, pervasive edge computing is facing great security problems. Suppose the remote server is honest but curious. In that case, it can still design algorithms for the intelligent terminal to execute and infer sensitive content such as their identity data and private pictures through the information returned by the intelligent terminal. In this paper, we research the problem of honest but curious remote servers infringing intelligent terminal privacy and propose a differential privacy collaborative deep learning algorithm in the pervasive edge computing environment. We use a Gaussian mechanism that meets the differential privacy guarantee to add noise on the first layer of the neural network to protect the data of the intelligent terminal and use analytical moments accountant technology to track the cumulative privacy loss. Experiments show that with the Gaussian mechanism, the training data of intelligent terminals can be protected reduction inaccuracy.
Dayin Zhang, Xiaojun Chen 0004, Jinqiao Shi, Dakui Wang
TrustCom2
2021 Robust node embedding against graph structural perturbations
Zhendong Zhao, Xiaojun Chen 0004, Dakui Wang, Yuexin Xuan
Inf. Sci.2
2020 A Shared-Word Sensitive Sequence-to-Sequence Features Extractor for Sentences Matching
abstract
Sentences matching is a basic task in Natural Language Processing (NLP). Interaction-based methods, which employ interactions between words of two sentences and construct word-level matching features to classify, are generally used due to their fine-grained features. However, they have many invalid interactions that may affect matching precision. In this paper, we limit the objects of interacting to shared words4 of two sentences. On the one hand, they can reduce invalid interactions. On the other hand, because of the different context semantics, the representation of the same word may be quite different, conversely, the representation difference can also be used to reflect the semantic difference of different contexts. To better extract global features of shared words, we introduce a sequence-to-sequence features extractor to force decoder to learn more contextual information from encoder. We implement the method based on Transformer[28], with syntactic parsing as additional knowledge. Our proposed method achieved better performance than strong baselines and the experiment results also demonstrate the efficiency of sequence-to-sequence features extractor and significance of the shared words.
Dakui Wang, Xiaojun Chen 0004, Pencheng Liao, Shujuan Chen
ECAI3
2020 An Efficient 3-Party Framework for Privacy-Preserving Neural Network Inference
Liyan Shen, Xiaojun Chen 0004, Jinqiao Shi, Ye Dong, Binxing Fang
ESORICS (1)2
2020 SoMem: A Self-optimizing Memory Network for Distributed Person Re-identification
abstract
Person Re-Identification (Re-ID) aims to match the persons contained in surveillance videos, and is usually run on powerful servers in a supervised mode. However, centralized processing of massive video from thousands of cameras in a city is very costly and causes serious problems of privacy protection. Moreover, the labeling of numerous data for supervised training is also infeasible in this scenario. To address this problem, we propose a novel Self-optimizing Memory Network model, namely SoMem, which runs person Re-ID on edge devices in a totally unsupervised and distributed way. Specifically, SoMem adopts a random walk based collaborative training procedure to optimize the visual model on each camera based on locally collected images, and builds a distributed memory network to memorize and match the observed persons by using a distributed mutual ranking algorithm. Based on the cross-camera person matching results learned by the memory network, the visual models on edge devices are further optimized in a self-organized manner. Comprehensive experiments are conducted on several real person Re-ID datasets and deployed on edge devices to show the effectiveness and efficiency of this novel distributed Re-ID model.
Jianming Lv, Chaojie Hu, Yipeng Zhou, Xiaojun Chen 0004
ICTAI4
2020 Improving Abstractive Summarization with Iterative Representation
abstract
In the neural abstractive summarization field, comprehensive document representation and summary embellishment are two major challenges. To tackle the above problems, we propose an Iterative Abstractive Summarization (IAS) model through iterating the document and summary representation. Specifically, (1) we design a selective gated strategy to constantly update the input representation in the encoder, which is consistent with the repeated updating of human memory information in human writing. (2) We design an iterative unit to revise the comprehensive representation iteratively for polishing the summary. Moreover, we utilize reinforcement learning to optimize our model for the non-differentiable metric ROUGE, which can alleviate the exposure bias during predicting words effectively. Experiments on the CNN/Daily Mail, Gigaword and DUC-2004 datasets show that the IAS model can generate high-quality summaries with varied length, and outperforms baseline methods significantly in terms of ROUGE and Human metrics.
Jinpeng Li 0003, Xiaojun Chen 0004, Yanan Cao 0001, Ruipeng Jia
IJCNN3
2020 Improving Abstractive Text Summarization with History Aggregation
abstract
Recent neural sequence to sequence models have provided feasible solutions for abstractive summarization. However, such models are still hard to tackle long text dependency in the summarization task. A high-quality summarization system usually depends on strong encoder which can refine important information from long input texts so that the decoder can generate salient summaries from the encoder's memory. In this paper, we propose an aggregation mechanism based on the Transformer model to address the challenge of long text representation. Our model can review history information to make encoder hold more memory capacity. Empirically, we apply our aggregation mechanism to the Transformer model and experiment on CNN/DailyMail dataset to achieve higher quality summaries compared to several strong baseline models on the ROUGE metrics.
Pengcheng Liao, Xiaojun Chen 0004, Xiaofei Zhou 0002
IJCNN3
2020 Generate Images with Obfuscated Attributes for Private Image Classification
Dakui Wang, Xiaojun Chen 0004
MMM (2)3
2020 CLTS: A New Chinese Long Text Summarization Dataset
Xiaojun Chen 0004, Yanan Cao 0001, Jinpeng Li 0003
NLPCC (1)3
2020 EaSTFLy: Efficient and secure ternary federated learning
Ye Dong, Xiaojun Chen 0004, Liyan Shen, Dakui Wang
Comput. Secur.2
2019 A More Efficient Private Set Intersection Protocol Based on Random OT and Balance Hash
abstract
Private set intersection (PSI) is a specific application problem in the field of secure multi-party computation. It allows the participants to compute set intersection collaboratively without learning any additional information about the sets. It is a building block for many real-world privacy-related applications. The state-of-the-art PSI protocol is mainly based on Random Oblivious Transfer (ROT) and cuckoo hash. It is efficient enough in computation but needs high communication cost and high network bandwidth. In this paper, we describe a novel PSI protocol based on ROT and balance hash under semi-honest adversary model. Experiments demonstrate that our protocol has better performance than the cuckoo hash based PSI at the expense of leaking a part of the indexes of hash bucket. We have proved that the impact of information leakage on security is almost negligible. In addition, for the purpose of reducing the communication overhead, we utilize cuckoo filter to optimize our protocol.
Liyan Shen, Xiaojun Chen 0004, Jinqiao Shi, Binxing Fang
ICC2
2019 Privacy-Preserving Distributed Machine Learning Based on Secret Sharing
Ye Dong, Xiaojun Chen 0004, Liyan Shen, Dakui Wang
ICICS2
2019 Abstractive Text Summarization with Multi-Head Attention
abstract
In this paper, we present a novel sequence-to-sequence architecture with multi-head attention for automatic summarization of long text. Summaries generated by previous abstractive methods have the problems of duplicate and missing original information commonly. To address these problems, we propose a multi-head attention summarization (MHAS) model, which uses multi-head attention mechanism to learn relevant information in different representation subspaces. The MHAS model can consider the previously predicted words when generating new words to avoid generating a summary of redundant repetition words. And it can learn the internal structure of the article by adding self-attention layer to the traditional encoder and decoder and make the model better preserve the original information. We also integrate the multi-head attention distribution into pointer network creatively to improve the performance of the model. Experiments are conducted on CNN/Daily Mail dataset, which is a long text English corpora. Experimental results show that our proposed model outperforms the previous extractive and abstractive models.
Jinpeng Li 0003, Xiaojun Chen 0004, Yanan Cao 0001, Pengcheng Liao, Peng Zhang 0001
IJCNN3
2018 Efficient and Private Set Intersection of Human Genomes
Liyan Shen, Xiaojun Chen 0004, Dakui Wang, Binxing Fang, Ye Dong
BIBM2
2018 Attention-Based RNN Model for Joint Extraction of Intent and Word Slot Based on a Tagging Strategy
Zheng Fang 0002, Yanan Cao 0001, Yanbing Liu 0007, Xiaojun Chen 0004, Jianlong Tan
ICANN (3)5
2018 Improve Word Mover's Distance with Part-of-Speech Tagging
abstract
Word Mover's Distance (WMD) is a document distance metric with free parameter, intelligible interpretation and unprecedented accuracy on document classification. WMD is on the basis of word embedding and largely focuses on semantic relationships rather than syntactic relationships, which would bring some limitations on measuring document distance. To enhance the impact of syntactic information, we proposed a new method called WMD with Part-of-Speech (PWMD) that integrates part-of-speech (POS) into the original WMD model. POS is a kind of syntactic information, providing more valuable features combined with WMD in document distance metric. Two combination strategies of the POS tagging are provided in “WMD, “word level” and “document level”. The results of contrastive experiments have shown that the PWMD is able to get better document distance than WMD.
Xiaojun Chen 0004, Li Bai 0004, Dakui Wang, Jinqiao Shi
ICPR1
2017 Efficient and Scalable Privacy-Preserving Similar Document Detection
abstract
Similar document detection has been well studied for many applications, such as file management systems, plagiarism and double submission detection. Traditional detection algorithms are challenged by the privacy-preserving problems. Recently, privacy-preserving similar document detection between two parties gains more attention. However, most of the existing works mainly focus on computing similarity between two documents, and they are inefficient with O(n2) computation complexity when processing secure comparison between two n-document sets. Focusing on this problem, this paper presents a new efficient and scalable privacy-preserving similar document detection protocol based on oblivious multi-garbled Bloom filter intersection and MinHash algorithm. Experimental evaluation shows that when processing large document sets, our protocol still remains linear computation complexity with the scale of document sets increasing and achieves overwhelming computational performance improvement against other major approaches.
Xiaojie Yu, Xiaojun Chen 0004, Jinqiao Shi, Liyan Shen, Dakui Wang
GLOBECOM2
2017 Improving Password Guessing Using Byte Pair Encoding
Dakui Wang, Xiaojun Chen 0004, Jinqiao Shi, Li Guo 0001
ISC3
2015 Towards misdirected email detection based on multi-attributes
abstract
Email has become widely used in recent years bringing with it new problems. Although this event doesn't happen often, misdirected emails can bring out great information leakage. It is not easy to detect these misdirected emails from legitimate ones since they may be only distinguishable in the sender's perspective. Existing methods discover misdirected emails from user agent or gateway but are not appropriate for varied application environment. This paper proposes a misdirected mail detection method based on multi-attributes which can be deployed on server side. Three type of attributes including email content fingerprinting, social relationship and meta information are considered in this method. Based on SVM classification algorithm, experiments show that it can detect misdirected emails with up to 91.6% accuracy.
Yiguo Pu, Jinqiao Shi, Xiaojun Chen 0004, Li Guo 0001, Tingwen Liu
ISCC3
2014 Towards misdirected email detection for preventing information leakage
abstract
With the widespread usage of emails, information leakage via misdirected emails becomes a practical and disastrous problem, which should be addressed at all costs. Prior methods have two limitations: privacy issue as relying on email contents to work, and high cost as building too many targeted models. In this paper, we reduce the detection of misdirected emails to a binary classification problem, and build only a universal model to detect misdirected emails. We introduce some representative features that can vividly describe the characteristics of misdirected emails while not infringe users' privacy. Then we design novel algorithms to get these features. The random forest classifier is chosen to perform the detecting task. Experimental results show that our work is able to detect misdirected emails with 89% precision rate and 82% recall rate in average.
Tingwen Liu, Yiguo Pu, Jinqiao Shi, Quangang Li, Xiaojun Chen 0004
ISCC5
2014 Winnowing Double Structure for Wildcard Query in Payload Attribution
Xiaojun Chen 0004, Yiguo Pu, Jinqiao Shi, Sihan Qing
ISC3
2007 Parallelizing Protocol Processing on SMT Processor Efficiently: A FSM Decomposition Approach
abstract
With the increase of network bandwidth, high performance protocol processing plays more and more important role in high speed network security. Recent studies show that current computer architecture advances and CPU performance improvements have limited impact on network protocol processing performance. Some studies find that in real SMT processor like Intel Xeon processor with hyper-threadings, the sharing resources (like cache) contention between threads can hurt the processing performance of network applications like servers or IDS. How to make protocol processing cope with the advances in computer architecture has been widely studied. In this paper, we put our focus on the processing performance of TCP automata phases, using execution based simulations to model the relationship between each phase performance and cache size, and then measuring the cache contention between threads. We find (1) the load/store units can be the bottleneck of protocol processing; and (2) in connection establishing phase of TCP processing, cache contention between threads is more aggressive than any other phase. We also suggest a FSM decomposition based parallel processing approach to use sharing cache of SMT processors effectively.
Li Guo 0001, Binxing Fang, Xiaojun Chen 0004
IPCCC4