Xuanli He

dblp:182/1859 · DBLP profile ↗
← Back
25ranked-venue papers
10as first author
18since 2021 · last 2025
0000-0002-0955-841XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 10 first-author · 18 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 GRADA: Graph-based Reranking against Adversarial Documents Attack
abstract
Retrieval Augmented Generation (RAG) frameworks can improve the factual accuracy of large language models (LLMs) by integrating external knowledge from retrieved documents, which is useful for overcoming the limitations of models' static intrinsic knowledge.However, these systems are susceptible to adversarial attacks that manipulate the retrieval process by introducing documents that are adversarial yet semantically similar to the query.Notably, while these adversarial documents resemble the query, they exhibit weak similarity to benign documents in the retrieval set.Thus, we propose a simple yet effective Graph-based Reranking against Adversarial Document Attacks (GRADA) framework aimed at preserving retrieval quality while significantly reducing the success of adversaries.Our study evaluates the effectiveness of our approach through experiments conducted on six LLMs: GPT-3.5-Turbo,GPT-4o, Llama3.1-8b-Instruct,Llama3.1-70b-Instruct,Qwen2.5-7b-Instruct, and Qwen2.5-14b-Instruct.We use three datasets to assess performance, with results from the Natural Questions dataset showing up to an 80% reduction in attack success rates while maintaining minimal loss in accuracy.
Jingjie Zheng, Aryo Pradipta Gema, Giwon Hong, Xuanli He, Pasquale Minervini, Youcheng Sun, Qiongkai Xu
EMNLP4
2025 An Auditing Test to Detect Behavioral Shift in Language Models
abstract
As language models (LMs) approach human-level performance, a comprehensive understanding of their behavior becomes crucial. This includes evaluating capabilities, biases, task performance, and alignment with societal values. Extensive initial evaluations, including red teaming and diverse benchmarking, can establish a model’s behavioral profile. However, subsequent fine-tuning or deployment modifications may alter these behaviors in unintended ways. We present an efficient statistical test to tackle Behavioral Shift Auditing (BSA) in LMs, which we define as detecting distribution shifts in qualitative properties of the output distributions of LMs. Our test compares model generations from a baseline model to those of the model under scrutiny and provides theoretical guarantees for change detection while controlling false positives. The test features a configurable tolerance parameter that adjusts sensitivity to behavioral changes for different use cases. We evaluate our approach using two case studies: monitoring changes in (a) toxicity and (b) translation performance. We find that the test is able to detect meaningful changes in behavior distributions using just hundreds of examples.
Leo Richter, Xuanli He, Pasquale Minervini, Matt J. Kusner
ICLR2
2025 IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models
abstract
David Ifeoluwa Adelani, Jessica Ojo, Israel Abebe Azime, Jian Yun Zhuang, Jesujoba Oluwadara Alabi, Xuanli He, Millicent Ochieng, Sara Hooker, Andiswa Bukula, En-Shiun Annie Lee, Chiamaka Ijeoma Chukwuneke, Happy Buzaaba, Blessing Kudzaishe Sibanda, Godson Koffi Kalipe, Jonathan Mukiibi, Salomon Kabongo Kabenamualu, Foutse Yuehgoh, Mmasibidi Setaka, Lolwethu Ndolela, Nkiruka Odu, Rooweither Mabuya, Salomey Osei, Shamsuddeen Hassan Muhammad, Sokhar Samb, Tadesse Kebede Guge, Tombekai Vangoni Sherman, Pontus Stenetorp. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
David Ifeoluwa Adelani, Jessica Ojo, Israel Abebe Azime, Jian Yun Zhuang, Jesujoba O. Alabi, Xuanli He, Millicent Ochieng, Sara Hooker, Andiswa Bukula, Annie En-Shiun Lee, Chiamaka Ijeoma Chukwuneke, Happy Buzaaba, Blessing K. Sibanda, Godson Kalipe, Jonathan Mukiibi, Salomon Kabongo, Foutse Yuehgoh, Mmasibidi Setaka, Lolwethu Ndolela, Nkiruka Odu, Rooweither Mabuya, Salomey Osei, Shamsuddeen Hassan Muhammad, Sokhar Samb, Tadesse Kebede Guge, Tombekai Vangoni Sherman, Pontus Stenetorp
NAACL (Long Papers)6
2025 Are We Done with MMLU?
abstract
Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile Van Krieken, Pasquale Minervini. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao 0043, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile van Krieken, Pasquale Minervini
NAACL (Long Papers)7
2025 Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
abstract
Yu Zhao, Alessio Devoto, Giwon Hong, Xiaotang Du, Aryo Pradipta Gema, Hongru Wang, Xuanli He, Kam-Fai Wong, Pasquale Minervini. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yu Zhao 0043, Alessio Devoto, Giwon Hong, Xiaotang Du, Aryo Pradipta Gema, Hongru Wang 0003, Xuanli He, Kam-Fai Wong, Pasquale Minervini
NAACL (Long Papers)7
2024 Using Natural Language Explanations to Improve Robustness of In-context Learning
abstract
Recent studies demonstrated that large language models (LLMs) can excel in many tasks via in-context learning (ICL).However, recent works show that ICL-prompted models tend to produce inaccurate results when presented with adversarial inputs.In this work, we investigate whether augmenting ICL with natural language explanations (NLEs) improves the robustness of LLMs on adversarial datasets covering natural language inference and paraphrasing identification.We prompt LLMs with a small set of human-generated NLEs to produce further NLEs, yielding more accurate results than both a zero-shot-ICL setting and using only human-generated NLEs.Our results on five popular LLMs (GPT3.5-turbo,Llama2, Vicuna, Zephyr, and Mistral) show that our approach yields over 6% improvement over baseline approaches for eight adversarial datasets: HANS, ISCS, NaN, ST, PICD, PISP, ANLI, and PAWS.Furthermore, previous studies have demonstrated that prompt selection strategies significantly enhance ICL on in-distribution test sets.However, our findings reveal that these strategies do not match the efficacy of our approach for robustness evaluations, resulting in an accuracy drop of 8% compared to the proposed approach.1
Xuanli He, Yuxiang Wu, Oana-Maria Camburu, Pasquale Minervini, Pontus Stenetorp
ACL (1)1
2024 AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African Languages
abstract
Jiayi Wang, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Anuoluwapo Aremu, Jessica Ojo, Shamsuddeen Hassan Muhammad, Salomey Osei, Abdul-Hakeem Omotayo, Chiamaka Chukwuneke, Perez Ogayo, Oumaima Hourrane, Salma El Anigri, Lolwethu Ndolela, Thabiso Mangwana, Shafie Abdi Mohamed, Hassan Ayinde, Oluwabusayo Olufunke Awoyomi, Lama Alkhaled, Sana Al-azzawi, Naome A. Etori, Millicent Ochieng, Clemencia Siro, Njoroge Kiragu, Eric Muchiri, Wangari Kimotho, Lyse Naomi Wamba Momo, Daud Abolade, Simbiat Ajao, Iyanuoluwa Shode, Ricky Macharm, Ruqayya Nasir Iro, Saheed S. Abdullahi, Stephen E. Moore, Bernard Opoku, Zainab Akinjobi, Abeeb Afolabi, Nnaemeka Obiefuna, Onyekachi Raphael Ogbu, Sam Ochieng’, Verrah Akinyi Otiende, Chinedu Emmanuel Mbonu, Sakayo Toadoum Sari, Yao Lu, Pontus Stenetorp. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jiayi Wang 0010, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin P. Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Aremu Anuoluwapo, Jessica Ojo, Shamsuddeen Hassan Muhammad, Salomey Osei, Abdul-Hakeem Omotayo, Chiamaka Ijeoma Chukwuneke, Perez Ogayo, Oumaima Hourrane, Salma El Anigri, Lolwethu Ndolela, Thabiso Mangwana, Shafie Abdi Mohamed, Ayinde Hassan, Oluwabusayo Olufunke Awoyomi, Lama Alkhaled, Sana Sabah Al-Azzawi, Naome A. Etori, Millicent Ochieng, Clemencia Siro, Njoroge Kiragu, Eric Muchiri, Wangari Kimotho, Sakayo Toadoum Sari, Lyse Naomi Wamba Momo, Daud Abolade, Simbiat Ajao, Iyanuoluwa Shode, Ricky Macharm, Ruqayya Nasir Iro, Saheed S. Abdullahi, Stephen E. Moore, Bernard Opoku, Zainab Akinjobi, Afolabi Abeeb, Nnaemeka C. Obiefuna, Onyekachi Raphael Ogbu, Sam Ochieng', Verrah Otiende, Chinedu E. Mbonu, Pontus Stenetorp
NAACL-HLT8
2024 Backdoor Attacks on Multilingual Machine Translation
abstract
Jun Wang, Qiongkai Xu, Xuanli He, Benjamin Rubinstein, Trevor Cohn. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jun Wang 0126, Qiongkai Xu, Xuanli He, Benjamin I. P. Rubinstein, Trevor Cohn
NAACL-HLT3
2024 SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks
abstract
Abstract Modern NLP models are often trained on public datasets drawn from diverse sources, rendering them vulnerable to data poisoning attacks. These attacks can manipulate the model’s behavior in ways engineered by the attacker. One such tactic involves the implantation of backdoors, achieved by poisoning specific training instances with a textual trigger and a target class label. Several strategies have been proposed to mitigate the risks associated with backdoor attacks by identifying and removing suspected poisoned examples. However, we observe that these strategies fail to offer effective protection against several advanced backdoor attacks. To remedy this deficiency, we propose a novel defensive mechanism that first exploits training dynamics to identify poisoned samples with high precision, followed by a label propagation step to improve recall and thus remove the majority of poisoned instances. Compared with recent advanced defense methods, our method considerably reduces the success rates of several backdoor attacks while maintaining high classification accuracy on clean test sets.
Xuanli He, Qiongkai Xu, Jun Wang 0126, Benjamin I. P. Rubinstein, Trevor Cohn
Trans. Assoc. Comput. Linguistics1
2023 Can Knowledge Graphs Simplify Text?
abstract
Knowledge Graph (KG)-to-Text Generation has seen recent improvements in generating fluent and informative sentences which describe a given KG. As KGs are widespread across multiple domains and contain important entity-relation information, and as text simplification aims to reduce the complexity of a text while preserving the meaning of the original text, we propose KGSimple, a novel approach to unsupervised text simplification which infuses KG-established techniques in order to construct a simplified KG path and generate a concise text which preserves the original input's meaning. Through an iterative and sampling KG-first approach, our model is capable of simplifying text when starting from a KG by learning to keep important information while harnessing KG-to-text generation to output fluent and descriptive sentences. We evaluate various settings of the KGSimple model on currently-available KG-to-text datasets, demonstrating its effectiveness compared to unsupervised text simplification models which start with a given complex text. Our code is available on GitHub.
Anthony M. Colas, Haodi Ma, Xuanli He, Daisy Zhe Wang
CIKM3
2023 Mitigating Backdoor Poisoning Attacks through the Lens of Spurious Correlation
abstract
Modern NLP models are often trained over large untrusted datasets, raising the potential for a malicious adversary to compromise model behaviour.For instance, backdoors can be implanted through crafting training instances with a specific textual trigger and a target label.This paper posits that backdoor poisoning attacks exhibit spurious correlation between simple text features and classification labels, and accordingly, proposes methods for mitigating spurious correlation as means of defence.Our empirical study reveals that the malicious triggers are highly correlated to their target labels; therefore such correlations are extremely distinguishable compared to those scores of benign features, and can be used to filter out potentially problematic instances.Compared with several existing defences, our defence method significantly reduces attack success rates across backdoor attacks, and in the case of insertion-based attacks, our method provides a near-perfect defence. 1
Xuanli He, Qiongkai Xu, Jun Wang 0126, Benjamin I. P. Rubinstein, Trevor Cohn
EMNLP1
2022 Protecting Intellectual Property of Language Generation APIs with Lexical Watermark
abstract
Nowadays, due to the breakthrough in natural language generation (NLG), including machine translation, document summarization, image captioning, etc NLG models have been encapsulated in cloud APIs to serve over half a billion people worldwide and process over one hundred billion word generations per day. Thus, NLG APIs have already become essential profitable services in many commercial companies. Due to the substantial financial and intellectual investments, service providers adopt a pay-as-you-use policy to promote sustainable market growth. However, recent works have shown that cloud platforms suffer from financial losses imposed by model extraction attacks, which aim to imitate the functionality and utility of the victim services, thus violating the intellectual property (IP) of cloud APIs. This work targets at protecting IP of NLG APIs by identifying the attackers who have utilized watermarked responses from the victim NLG APIs. However, most existing watermarking techniques are not directly amenable for IP protection of NLG APIs. To bridge this gap, we first present a novel watermarking method for text generation APIs by conducting lexical modification to the original outputs. Compared with the competitive baselines, our watermark approach achieves better identifiable performance in terms of p-value, with fewer semantic losses. In addition, our watermarks are more understandable and intuitive to humans than the baselines. Finally, the empirical studies show our approach is also applicable to queries from different domains, and is effective on the attacker trained on a mixture of the corpus which includes less than 10% watermarked samples.
Xuanli He, Qiongkai Xu, Lingjuan Lyu, Fangzhao Wu, Chenguang Wang 0001
AAAI1
2022 Student Surpasses Teacher: Imitation Attack for Black-Box NLP APIs
abstract
Machine-learning-as-a-service (MLaaS) has attracted millions of users to their splendid large-scale models. Although published as black-box APIs, the valuable models behind these services are still vulnerable to imitation attacks. Recently, a series of works have demonstrated that attackers manage to steal or extract the victim models. Nonetheless, none of the previous stolen models can outperform the original black-box APIs. In this work, we conduct unsupervised domain adaptation and multi-victim ensemble to showing that attackers could potentially surpass victims, which is beyond previous understanding of model extraction. Extensive experiments on both benchmark datasets and real-world APIs validate that the imitators can succeed in outperforming the original black-box models on transferred domains. We consider our work as a milestone in the research of imitation attack, especially on NLP APIs, as the superior performance could influence the defense or even publishing strategy of API providers.
Qiongkai Xu, Xuanli He, Lingjuan Lyu, Lizhen Qu, Gholamreza Haffari
COLING2
2022 Extracted BERT Model Leaks More Information than You Think!
abstract
The collection and availability of big data, combined with advances in pre-trained models (e.g.BERT), have revolutionized the predictive performance of natural language processing tasks.This allows corporations to provide machine learning as a service (MLaaS) by encapsulating fine-tuned BERT-based models as APIs.Due to significant commercial interest, there has been a surge of attempts to steal remote services via model extraction.Although previous works have made progress in defending against model extraction attacks, there has been little discussion on their performance in preventing privacy leakage.This work bridges this gap by launching an attribute inference attack against the extracted BERT model.Our extensive experiments reveal that model extraction can cause severe privacy leakage even when victim models are facilitated with advanced defensive strategies.
Xuanli He, Lingjuan Lyu, Chen Chen 0043, Qiongkai Xu
EMNLP1
2022 CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks
abstract
Previous works have validated that text generation APIs can be stolen through imitation attacks, causing IP violations. In order to protect the IP of text generation APIs, recent work has introduced a watermarking algorithm and utilized the null-hypothesis test as a post-hoc ownership verification on the imitation models. However, we find that it is possible to detect those watermarks via sufficient statistics of the frequencies of candidate watermarking words. To address this drawback, in this paper, we propose a novel Conditional wATERmarking framework (CATER) for protecting the IP of text generation APIs. An optimization method is proposed to decide the watermarking rules that can minimize the distortion of overall word distributions while maximizing the change of conditional word selections. Theoretically, we prove that it is infeasible for even the savviest attacker (they know how CATER works) to reveal the used watermarks from a large pool of potential word pairs based on statistical inspection. Empirically, we observe that high-order conditions lead to an exponential growth of suspicious (unused) watermarks, making our crafted watermarks more stealthy. In addition, CATER can effectively identify IP infringement under architectural mismatch and cross-domain imitation attacks, with negligible impairments on the generation quality of victim APIs. We envision our work as a milestone for stealthily protecting the IP of text generation APIs.
Xuanli He, Qiongkai Xu, Yi Zeng 0005, Lingjuan Lyu, Fangzhao Wu, Jiwei Li 0001, Ruoxi Jia 0001
NeurIPS1
2022 Generate, Annotate, and Learn: NLP with Synthetic Text
abstract
Abstract This paper studies the use of language models as a source of synthetic unlabeled text for NLP. We formulate a general framework called “generate, annotate, and learn (GAL)” to take advantage of synthetic text within knowledge distillation, self-training, and few-shot learning applications. To generate high-quality task-specific text, we either fine-tune LMs on inputs from the task of interest, or prompt large LMs with few examples. We use the best available classifier to annotate synthetic text with soft pseudo labels for knowledge distillation and self-training, and use LMs to obtain hard labels for few-shot learning. We train new supervised models on the combination of labeled and pseudo-labeled data, which results in significant gains across several applications. We investigate key components of GAL and present theoretical and empirical arguments against the use of class-conditional LMs to generate synthetic labeled text instead of unlabeled text. GAL achieves new state-of-the-art knowledge distillation results for 6-layer transformers on the GLUE leaderboard.
Xuanli He, Islam Nassar, Jamie Kiros, Gholamreza Haffari, Mohammad Norouzi 0002
Trans. Assoc. Comput. Linguistics1
2021 Generalised Unsupervised Domain Adaptation of Neural Machine Translation with Cross-Lingual Data Selection
abstract
This paper considers the unsupervised domain adaptation problem for neural machine translation (NMT), where we assume the access to only monolingual text in either the source or target language in the new domain.We propose a cross-lingual data selection method to extract in-domain sentences in the missing language side from a large generic monolingual corpus.Our proposed method trains an adaptive layer on top of multilingual BERT by contrastive learning to align the representation between the source and target language.This then enables the transferability of the domain classifier between the languages in a zero-shot manner.Once the in-domain data is detected by the classifier, the NMT model is then adapted to the new domain by jointly learning translation and domain discrimination tasks.We evaluate our cross-lingual data selection method on NMT across five diverse domains in three language pairs, as well as a real-world scenario of translation for COVID-19.The results show that our proposed method outperforms other selection baselines up to +1.5 BLEU score.
Thuy-Trang Vu, Xuanli He, Dinh Q. Phung, Gholamreza Haffari
EMNLP (1)2
2021 Model Extraction and Adversarial Transferability, Your BERT is Vulnerable!
abstract
Natural language processing (NLP) tasks, ranging from text classification to text generation, have been revolutionised by the pretrained language models, such as BERT.This allows corporations to easily build powerful APIs by encapsulating fine-tuned BERT models for downstream tasks.However, when a fine-tuned BERT model is deployed as a service, it may suffer from different attacks launched by the malicious users.In this work, we first present how an adversary can steal a BERT-based API service (the victim/target model) on multiple benchmark datasets with limited prior knowledge and queries.We further show that the extracted model can lead to highly transferable adversarial attacks against the victim model.Our studies indicate that the potential vulnerabilities of BERT-based API services still hold, even when there is an architectural mismatch between the victim model and the attack model.Finally, we investigate two defence strategies to protect the victim model, and find that unless the performance of the victim model is sacrificed, both model extraction and adversarial transferability can effectively compromise the target models.
Xuanli He, Lingjuan Lyu, Lichao Sun 0001, Qiongkai Xu
NAACL-HLT1
2020 Dynamic Programming Encoding for Subword Segmentation in Neural Machine Translation
abstract
This paper introduces Dynamic Programming Encoding (DPE), a new segmentation algorithm for tokenizing sentences into subword units.We view the subword segmentation of output sentences as a latent variable that should be marginalized out for learning and inference.A mixed character-subword transformer is proposed, which enables exact log marginal likelihood estimation and exact MAP inference to find target segmentations with maximum posterior probability.DPE uses a lightweight mixed character-subword transformer as a means of pre-processing parallel data to segment output sentences using dynamic programming.Empirical results on machine translation suggest that DPE is effective for segmenting output sentences and can be combined with BPE dropout for stochastic segmentation of source sentences.DPE achieves an average improvement of 0.9 BLEU over BPE (Sennrich et al., 2016) and an average improvement of 0.55 BLEU over BPE dropout (Provilkov et al., 2019) on several WMT datasets including English ↔ (German, Romanian, Estonian, Finnish, Hungarian).
Xuanli He, Gholamreza Haffari, Mohammad Norouzi 0002
ACL1
2020 Towards Differentially Private Text Representations
abstract
Most deep learning frameworks require users to pool their local data or model updates to a trusted server to train or maintain a global model. The assumption of a trusted server who has access to user information is ill-suited in many applications. To tackle this problem, we develop a new deep learning framework under an untrusted server setting, which includes three modules: (1) embedding module, (2) randomization module, and (3) classifier module. For the randomization module, we propose a novel local differentially private (LDP) protocol to reduce the impact of privacy parameter ε on accuracy, and provide enhanced flexibility in choosing randomization probabilities for LDP. Analysis and experiments show that our framework delivers comparable or even better performance than the non-private framework and existing LDP protocols, demonstrating the advantages of our LDP protocol.
Lingjuan Lyu, Yitong Li 0002, Xuanli He, Tong Xiao 0001
SIGIR3
2019 Fog-Embedded Deep Learning for the Internet of Things
abstract
In current deep learning models, centralized architecture forces participants to pool their data to the central Cloud to train a global model, while distributed architecture requires a parameter server to mediate the training process. However, privacy issues, response delays, and computation and communication bottlenecks prevent these architectures from working well at the scale of Internet of Things devices. To counter these problems, in this paper we build a Fog-embedded privacy-preserving deep learning framework (FPPDL), which moves computation from the centralized Cloud to Fog nodes near the end devices. The experimental results on benchmark image datasets under different settings demonstrate that FPPDL achieves comparable accuracy to the centralized stochastic gradient descent (SGD) framework, and delivers better accuracy than the standalone SGD framework. Our evaluations also show that both computation and communication cost are greatly reduced by FPPDL, hence achieving the desired tradeoff between privacy and performance.
Lingjuan Lyu, James C. Bezdek, Xuanli He, Jiong Jin
IEEE Trans. Ind. Informatics3
2018 Sequence to Sequence Mixture Model for Diverse Machine Translation
abstract
Sequence to sequence (SEQ2SEQ) models often lack diversity in their generated translations.This can be attributed to the limitation of SEQ2SEQ models in capturing lexical and syntactic variations in a parallel corpus resulting from different styles, genres, topics, or ambiguity of the translation process.In this paper, we develop a novel sequence to sequence mixture (S2SMIX) model that improves both translation diversity and quality by adopting a committee of specialized translation models rather than a single translation model.Each mixture component selects its own training dataset via optimization of the marginal loglikelihood, which leads to a soft clustering of the parallel corpus.Experiments on four language pairs demonstrate the superiority of our mixture model compared to a SEQ2SEQ baseline with standard or diversity-boosted beam search.Our mixture model uses negligible additional parameters and incurs no extra computation cost during decoding.
Xuanli He, Gholamreza Haffari, Mohammad Norouzi 0002
CoNLL1
2018 Privacy-preserving collaborative fuzzy clustering
Lingjuan Lyu, James C. Bezdek, Yee Wei Law, Xuanli He, Marimuthu Palaniswami
Data Knowl. Eng.4
2017 Privacy-Preserving Collaborative Deep Learning with Application to Human Activity Recognition
abstract
The proliferation of wearable devices has contributed to the emergence of mobile crowdsensing, which leverages the power of the crowd to collect and report data to a third party for large-scale sensing and collaborative learning. However, since the third party may not be honest, privacy poses a major concern. In this paper, we address this concern with a two-stage privacy-preserving scheme called RG-RP: the first stage is designed to mitigate maximum a posteriori (MAP) estimation attacks by perturbing each participant's data through a nonlinear function called repeated Gompertz (RG); while the second stage aims to maintain accuracy and reduce transmission energy by projecting high-dimensional data to a lower dimension, using a row-orthogonal random projection (RP) matrix. The proposed RG-RP scheme delivers better recovery resistance to MAP estimation attacks than most state-of-the-art techniques on both synthetic and real-world datasets. For collaborative learning, we proposed a novel LSTM-CNN model combining the merits of Long Short-Term Memory (LSTM) and Convolutional Neural Networks (CNN). Our experiments on two representative movement datasets captured by wearable sensors demonstrate that the proposed LSTM-CNN model outperforms standalone LSTM, CNN and Deep Belief Network. Together, RG+RP and LSTM-CNN provide a privacy-preserving collaborative learning framework that is both accurate and privacy-preserving.
Lingjuan Lyu, Xuanli He, Yee Wei Law, Marimuthu Palaniswami
CIKM2
2017 Fog-Empowered Anomaly Detection in IoT Using Hyperellipsoidal Clustering
abstract
Anomaly detection is important for time-critical Internet of Things (IoT) applications, such as healthcare and emergency management. The recent introduction of Fog computing architecture provides an efficient platform for delay sensitive IoT applications. Exploiting the advantages of Fog computing for anomaly detection provides the ability to detect abnormal patterns in an accurate and timely manner. Use of Centralized and Distributed anomaly detection methods suffer from significant latency and energy consumption issues. Hence, we propose a novel anomaly detection method, called Fog-Empowered anomaly detection, by harnessing the processing power of the Fog computing platform and using an efficient hyperellipsoidal clustering algorithm. The end nodes in the Fog computing architecture do not perform any processing or clustering on the data. The Fog layer and the Cloud layer nodes perform the clustering and anomaly detection process, thus helping to achieve anomaly detection in a timely manner. The evaluation using synthetic and real datasets demonstrates that our proposed approach achieves a significant reduction in latency and energy consumption compared to the Distributed and Centralized schemes, while achieving a comparable detection accuracy compared to a Centralized scheme.
Lingjuan Lyu, Jiong Jin, Sutharshan Rajasegarar, Xuanli He, Marimuthu Palaniswami
IEEE Internet Things J.4