Qiongkai Xu

dblp:127/0174 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0003-3312-6825ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 7 first-author · 19 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Activation Decomposition and Steering for LLM Backdoor Remediation
abstract
Existing works on defending against LLM backdoor attacks rely on either auxiliary models or safety-related datasets for defending against backdoor attacks on large language models, which are not always available.To address these challenges, we propose our Contrastive-Selective Activation Decomposition and Steering (CS-ADS), which contrasts relatively more benign and poisoned settings to decompose the feature vectors for steering without relying on additional auxiliary models or datasets.With such disentangled vectors for remediation, our method can achieve feasible defense qualities even better than datasetbased contrastive steering strategies.This novel decomposition-based solution is motivated by the key insight that feature representations of prompt pairs can encode the same benign semantics in different proportions, even when both prompt pairs are similarly backdoored.Such discrepancies allow our method to identify effective remediation directions for steering the generation process, thereby preventing undesired outputs.We evaluate CS-ADS against multiple state-of-the-art backdoor attacks, and experimental results show that CS-ADS provides effective defense across settings.
Lingfeng Zhong, Qiongkai Xu, Usman Naseem
ACL (1)2
2025 ALGEN: Few-shot Inversion Attacks on Textual Embeddings via Cross-Model Alignment and Generation
abstract
With the growing popularity of Large Language Models (LLMs) and vector databases, private textual data is increasingly processed and stored as numerical embeddings. However, recent studies have proven that such embeddings are vulnerable to inversion attacks, where original text is reconstructed to reveal sensitive information. Previous research has largely assumed access to millions of sentences to train attack models, e.g., through data leakage or nearly unrestricted API access. With our method, a single data point is sufficient for a partially successful inversion attack. With as little as 1k data samples, performance reaches an optimum across a range of black-box encoders, without training on leaked data. We present a Few-shot Textual Embedding Inversion Attack using Cross-Model ALignment and GENeration (ALGEN), by aligning victim embeddings to the attack space and using a generative model to reconstruct text. We find that ALGEN attacks can be effectively transferred across domains and languages, revealing key information. We further examine a variety of defense mechanisms against ALGEN, and find that none are effective, highlighting the vulnerabilities posed by inversion attacks. By significantly lowering the cost of inversion and proving that embedding spaces can be aligned through one-step optimization, we establish a new textual embedding inversion paradigm with broader applications for embedding alignment in NLP.
Yiyi Chen 0002, Qiongkai Xu, Johannes Bjerva
ACL (1)2
2025 WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation Watermarks
abstract
Embeddings-as-a-Service (EaaS) is a service offered by large language model (LLM) developers to supply embeddings generated by LLMs.Previous research suggests that EaaS is prone to imitation attacks-attacks that clone the underlying EaaS model by training another model on the queried embeddings.As a result, EaaS watermarks are introduced to protect the intellectual property of EaaS providers.In this paper, we first show that existing EaaS watermarks can be removed by paraphrasing when attackers clone the model.Subsequently, we propose a novel watermarking technique that involves linearly transforming the embeddings, and show that it is empirically and theoretically robust against paraphrasing.1 Training Dataset ` Verification Dataset 0.4 0.6 0 0 0.4 0.6 0.6 0 0.4 0.57 -0.85 1.28 1.28 0.57 -0.85 -0.85 1.28 0.57 Secret Linear TransformationInverse Linear Transformation
Anudeex Shetty, Qiongkai Xu, Jey Han Lau
ACL (1)2
2025 PatentScore: Multi-dimensional Evaluation of LLM-Generated Patent Claims
abstract
High-stakes texts such as patent claims, medical records, and technical reports are structurally complex and demand a high degree of reliability and precision.While large language models (LLMs) have recently been applied to automate their generation in high-stakes domains, reliably evaluating such outputs remains a major challenge.Conventional natural language generation (NLG) metrics are effective for generic documents but fail to capture the structural and legal characteristics essential to evaluating complex high-stakes documents.To address this gap, we propose PatentScore, a multi-dimensional evaluation framework specifically designed for one of the most intricate and rigorous domains, patent claims.PatentScore integrates hierarchical decomposition of claim elements, validation patterns grounded in legal and technical standards, and scoring across structural, semantic, and legal dimensions.In experiments on our dataset which consists of 400 Claim1, PatentScore achieved the highest correlation with expert annotations (r = 0.819), significantly outperforming widely used NLG metrics.This work establishes a new standard for evaluating LLMgenerated patent claims, providing a solid foundation for research on patent generation and validation.
Yongmin Yoo, Qiongkai Xu, Longbing Cao
EMNLP2
2025 GRADA: Graph-based Reranking against Adversarial Documents Attack
abstract
Retrieval Augmented Generation (RAG) frameworks can improve the factual accuracy of large language models (LLMs) by integrating external knowledge from retrieved documents, which is useful for overcoming the limitations of models' static intrinsic knowledge.However, these systems are susceptible to adversarial attacks that manipulate the retrieval process by introducing documents that are adversarial yet semantically similar to the query.Notably, while these adversarial documents resemble the query, they exhibit weak similarity to benign documents in the retrieval set.Thus, we propose a simple yet effective Graph-based Reranking against Adversarial Document Attacks (GRADA) framework aimed at preserving retrieval quality while significantly reducing the success of adversaries.Our study evaluates the effectiveness of our approach through experiments conducted on six LLMs: GPT-3.5-Turbo,GPT-4o, Llama3.1-8b-Instruct,Llama3.1-70b-Instruct,Qwen2.5-7b-Instruct, and Qwen2.5-14b-Instruct.We use three datasets to assess performance, with results from the Natural Questions dataset showing up to an 80% reduction in attack success rates while maintaining minimal loss in accuracy.
Jingjie Zheng, Aryo Pradipta Gema, Giwon Hong, Xuanli He, Pasquale Minervini, Youcheng Sun, Qiongkai Xu
EMNLP7
2024 WARDEN: Multi-Directional Backdoor Watermarks for Embedding-as-a-Service Copyright Protection
abstract
Embedding as a Service (EaaS) has become a widely adopted solution, which offers feature extraction capabilities for addressing various downstream tasks in Natural Language Processing (NLP).Prior studies have shown that EaaS can be prone to model extraction attacks; nevertheless, this concern could be mitigated by adding backdoor watermarks to the text embeddings and subsequently verifying the attack models post-publication.Through the analysis of the recent watermarking strategy for EaaS, EmbMarker, we design a novel CSE (Clustering, Selection, Elimination) attack that removes the backdoor watermark while maintaining the high utility of embeddings, indicating that the previous watermarking approach can be breached.In response to this new threat, we propose a new protocol to make the removal of watermarks more challenging by incorporating multiple possible watermark directions.Our defense approach, WARDEN, notably increases the stealthiness of watermarks and has been empirically shown to be effective against CSE attack.
Anudeex Shetty, Qiongkai Xu
ACL (1)4
2024 Seeing the Forest through the Trees: Data Leakage from Partial Transformer Gradients
abstract
Recent studies have shown that distributed machine learning is vulnerable to gradient inversion attacks, where private training data can be reconstructed by analyzing the gradients of the models shared in training.Previous attacks established that such reconstructions are possible using gradients from all parameters in the entire models.However, we hypothesize that most of the involved modules, or even their sub-modules, are at risk of training data leakage, and we validate such vulnerabilities in various intermediate layers of language models.Our extensive experiments reveal that gradients from a single Transformer layer, or even a single linear component with 0.54% parameters, are susceptible to training data leakage.Additionally, we show that applying differential privacy on gradients during training offers limited protection against the novel vulnerability of data disclosure. 1
Weijun Li 0003, Qiongkai Xu, Mark Dras
EMNLP2
2024 Backdoor Attacks on Multilingual Machine Translation
abstract
Jun Wang, Qiongkai Xu, Xuanli He, Benjamin Rubinstein, Trevor Cohn. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jun Wang 0126, Qiongkai Xu, Xuanli He, Benjamin I. P. Rubinstein, Trevor Cohn
NAACL-HLT2
2024 SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks
abstract
Abstract Modern NLP models are often trained on public datasets drawn from diverse sources, rendering them vulnerable to data poisoning attacks. These attacks can manipulate the model’s behavior in ways engineered by the attacker. One such tactic involves the implantation of backdoors, achieved by poisoning specific training instances with a textual trigger and a target class label. Several strategies have been proposed to mitigate the risks associated with backdoor attacks by identifying and removing suspected poisoned examples. However, we observe that these strategies fail to offer effective protection against several advanced backdoor attacks. To remedy this deficiency, we propose a novel defensive mechanism that first exploits training dynamics to identify poisoned samples with high precision, followed by a label propagation step to improve recall and thus remove the majority of poisoned instances. Compared with recent advanced defense methods, our method considerably reduces the success rates of several backdoor attacks while maintaining high classification accuracy on clean test sets.
Xuanli He, Qiongkai Xu, Jun Wang 0126, Benjamin I. P. Rubinstein, Trevor Cohn
Trans. Assoc. Comput. Linguistics2
2023 Fingerprint Attack: Client De-Anonymization in Federated Learning
abstract
Federated Learning allows collaborative training without data sharing in settings where participants do not trust the central server and one another. Privacy can be further improved by ensuring that communication between the participants and the server is anonymized through a shuffle; decoupling the participant identity from their data. This paper seeks to examine whether such a defense is adequate to guarantee anonymity, by proposing a novel fingerprinting attack over gradients sent by the participants to the server. We show that clustering of gradients can easily break the anonymization in an empirical study of learning federated language models on two language corpora. We then show that training with differential privacy can provide a practical defense against our fingerprint attack.
Qiongkai Xu, Trevor Cohn, Olga Ohrimenko
ECAI1
2023 Mitigating Backdoor Poisoning Attacks through the Lens of Spurious Correlation
abstract
Modern NLP models are often trained over large untrusted datasets, raising the potential for a malicious adversary to compromise model behaviour.For instance, backdoors can be implanted through crafting training instances with a specific textual trigger and a target label.This paper posits that backdoor poisoning attacks exhibit spurious correlation between simple text features and classification labels, and accordingly, proposes methods for mitigating spurious correlation as means of defence.Our empirical study reveals that the malicious triggers are highly correlated to their target labels; therefore such correlations are extremely distinguishable compared to those scores of benign features, and can be used to filter out potentially problematic instances.Compared with several existing defences, our defence method significantly reduces attack success rates across backdoor attacks, and in the case of insertion-based attacks, our method provides a near-perfect defence. 1
Xuanli He, Qiongkai Xu, Jun Wang 0126, Benjamin I. P. Rubinstein, Trevor Cohn
EMNLP2
2023 Humanly Certifying Superhuman Classifiers
Qiongkai Xu, Christian Walder
ICLR1
2023 Training-free Lexical Backdoor Attacks on Language Models
abstract
Large-scale language models have achieved tremendous success across various natural language processing (NLP) applications. Nevertheless, language models are vulnerable to backdoor attacks, which inject stealthy triggers into models for steering them to undesirable behaviors. Most existing backdoor attacks, such as data poisoning, require further (re)training or fine-tuning language models to learn the intended backdoor patterns. The additional training process however diminishes the stealthiness of the attacks, as training a language model usually requires long optimization time, a massive amount of data, and considerable modifications to the model parameters.
Yujin Huang, Terry Yue Zhuo, Qiongkai Xu, Han Hu 0011, Xingliang Yuan, Chunyang Chen 0001
WWW3
2022 Protecting Intellectual Property of Language Generation APIs with Lexical Watermark
abstract
Nowadays, due to the breakthrough in natural language generation (NLG), including machine translation, document summarization, image captioning, etc NLG models have been encapsulated in cloud APIs to serve over half a billion people worldwide and process over one hundred billion word generations per day. Thus, NLG APIs have already become essential profitable services in many commercial companies. Due to the substantial financial and intellectual investments, service providers adopt a pay-as-you-use policy to promote sustainable market growth. However, recent works have shown that cloud platforms suffer from financial losses imposed by model extraction attacks, which aim to imitate the functionality and utility of the victim services, thus violating the intellectual property (IP) of cloud APIs. This work targets at protecting IP of NLG APIs by identifying the attackers who have utilized watermarked responses from the victim NLG APIs. However, most existing watermarking techniques are not directly amenable for IP protection of NLG APIs. To bridge this gap, we first present a novel watermarking method for text generation APIs by conducting lexical modification to the original outputs. Compared with the competitive baselines, our watermark approach achieves better identifiable performance in terms of p-value, with fewer semantic losses. In addition, our watermarks are more understandable and intuitive to humans than the baselines. Finally, the empirical studies show our approach is also applicable to queries from different domains, and is effective on the attacker trained on a mixture of the corpus which includes less than 10% watermarked samples.
Xuanli He, Qiongkai Xu, Lingjuan Lyu, Fangzhao Wu, Chenguang Wang 0001
AAAI2
2022 Student Surpasses Teacher: Imitation Attack for Black-Box NLP APIs
abstract
Machine-learning-as-a-service (MLaaS) has attracted millions of users to their splendid large-scale models. Although published as black-box APIs, the valuable models behind these services are still vulnerable to imitation attacks. Recently, a series of works have demonstrated that attackers manage to steal or extract the victim models. Nonetheless, none of the previous stolen models can outperform the original black-box APIs. In this work, we conduct unsupervised domain adaptation and multi-victim ensemble to showing that attackers could potentially surpass victims, which is beyond previous understanding of model extraction. Extensive experiments on both benchmark datasets and real-world APIs validate that the imitators can succeed in outperforming the original black-box models on transferred domains. We consider our work as a milestone in the research of imitation attack, especially on NLP APIs, as the superior performance could influence the defense or even publishing strategy of API providers.
Qiongkai Xu, Xuanli He, Lingjuan Lyu, Lizhen Qu, Gholamreza Haffari
COLING1
2022 Extracted BERT Model Leaks More Information than You Think!
abstract
The collection and availability of big data, combined with advances in pre-trained models (e.g.BERT), have revolutionized the predictive performance of natural language processing tasks.This allows corporations to provide machine learning as a service (MLaaS) by encapsulating fine-tuned BERT-based models as APIs.Due to significant commercial interest, there has been a surge of attempts to steal remote services via model extraction.Although previous works have made progress in defending against model extraction attacks, there has been little discussion on their performance in preventing privacy leakage.This work bridges this gap by launching an attribute inference attack against the extracted BERT model.Our extensive experiments reveal that model extraction can cause severe privacy leakage even when victim models are facilitated with advanced defensive strategies.
Xuanli He, Lingjuan Lyu, Chen Chen 0043, Qiongkai Xu
EMNLP4
2022 Variational Autoencoder with Disentanglement Priors for Low-Resource Task-Specific Natural Language Generation
abstract
In this paper, we propose a variational autoencoder with disentanglement priors, VAE-DPRIOR, for task-specific natural language generation with none or a handful of taskspecific labeled examples.In order to tackle compositional generalization across tasks, our model performs disentangled representation learning by introducing a conditional prior for the latent content space and another conditional prior for the latent label space.Both types of priors satisfy a novel property called ϵ-disentangled.We show both empirically and theoretically that the novel priors can disentangle representations even without specific regularizations as in the prior work.The content prior enables directly sampling diverse content representations from the content space learned from the seen tasks, and fuse them with the representations of novel tasks for generating semantically diverse texts in the low-resource settings.Our extensive experiments demonstrate the superior performance of our model over competitive baselines in terms of i) data augmentation in continuous zero/few-shot learning, and ii) text style transfer in the few-shot setting.The code is available at https://github. com/zhuang-li/VAE-DPrior.
Zhuang Li 0001, Lizhen Qu, Qiongkai Xu, Tongtong Wu, Tianyang Zhan, Gholamreza Haffari
EMNLP3
2022 CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks
abstract
Previous works have validated that text generation APIs can be stolen through imitation attacks, causing IP violations. In order to protect the IP of text generation APIs, recent work has introduced a watermarking algorithm and utilized the null-hypothesis test as a post-hoc ownership verification on the imitation models. However, we find that it is possible to detect those watermarks via sufficient statistics of the frequencies of candidate watermarking words. To address this drawback, in this paper, we propose a novel Conditional wATERmarking framework (CATER) for protecting the IP of text generation APIs. An optimization method is proposed to decide the watermarking rules that can minimize the distortion of overall word distributions while maximizing the change of conditional word selections. Theoretically, we prove that it is infeasible for even the savviest attacker (they know how CATER works) to reveal the used watermarks from a large pool of potential word pairs based on statistical inspection. Empirically, we observe that high-order conditions lead to an exponential growth of suspicious (unused) watermarks, making our crafted watermarks more stealthy. In addition, CATER can effectively identify IP infringement under architectural mismatch and cross-domain imitation attacks, with negligible impairments on the generation quality of victim APIs. We envision our work as a milestone for stealthily protecting the IP of text generation APIs.
Xuanli He, Qiongkai Xu, Yi Zeng 0005, Lingjuan Lyu, Fangzhao Wu, Jiwei Li 0001, Ruoxi Jia 0001
NeurIPS2
2021 Model Extraction and Adversarial Transferability, Your BERT is Vulnerable!
abstract
Natural language processing (NLP) tasks, ranging from text classification to text generation, have been revolutionised by the pretrained language models, such as BERT.This allows corporations to easily build powerful APIs by encapsulating fine-tuned BERT models for downstream tasks.However, when a fine-tuned BERT model is deployed as a service, it may suffer from different attacks launched by the malicious users.In this work, we first present how an adversary can steal a BERT-based API service (the victim/target model) on multiple benchmark datasets with limited prior knowledge and queries.We further show that the extracted model can lead to highly transferable adversarial attacks against the victim model.Our studies indicate that the potential vulnerabilities of BERT-based API services still hold, even when there is an architectural mismatch between the victim model and the attack model.Finally, we investigate two defence strategies to protect the victim model, and find that unless the performance of the victim model is sacrificed, both model extraction and adversarial transferability can effectively compromise the target models.
Xuanli He, Lingjuan Lyu, Lichao Sun 0001, Qiongkai Xu
NAACL-HLT4
2021 Privacy Monitoring Service for Conversations
abstract
Leakage of personal information in conversations raises serious privacy concerns. Malicious people or bots could pry into sensitive personal information of vulnerable people, such as juveniles, through conversations with them or their digital personal assistants. To address the problem, we present a privacy-leakage warning system that monitors conversations in social media and intercepts the outgoing text messages from a user or a digital assistant, if they impose potential privacy leakage risks. Such messages are redirected to authorized users for approval, before they are sent out. We demonstrate how our system is deployed and used on a social media conversation platform, e.g., Facebook Messenger.
Qiongkai Xu, Lizhen Qu
WSDM1
2020 Adhering, Steering, and Queering: Treatment of Gender in Natural Language Generation
abstract
Natural Language Generation (NLG) supports the creation of personalized, contextualized, and targeted content. However, the algorithms underpinning NLG have come under scrutiny for reinforcing gender, racial, and other problematic biases. Recent research in NLG seeks to remove these biases through principles of fairness and privacy. Drawing on gender and queer theories from sociology and Science and Technology studies, we consider how NLG can contribute towards the advancement of gender equity in society. We propose a conceptual framework and technical parameters for aligning NLG with feminist HCI qualities. We present three approaches: (1) adhering to current approaches of removing sensitive gender attributes, (2) steering gender differences away from the norm, and (3) queering gender by troubling stereotypes. We discuss the advantages and limitations of these approaches across three hypothetical scenarios; newspaper headlines, job advertisements, and chatbots. We conclude by discussing considerations for implementing this framework and related ethical and equity agendas.
Yolande A. A. Strengers, Lizhen Qu, Qiongkai Xu, Jarrod Knibbe
CHI3
2020 Personal Information Leakage Detection in Conversations
abstract
The global market size of conversational assistants (chatbots) is expected to grow to USD 9.4 billion by 2024, according to Marketsand-Markets.Despite the wide use of chatbots, leakage of personal information through chatbots poses serious privacy concerns for their users.In this work, we propose to protect personal information by warning users of detected suspicious sentences generated by conversational assistants.The detection task is formulated as an alignment optimization problem and a new dataset PERSONA-LEAKAGE is collected for evaluation.In this paper, we propose two novel constrained alignment models, which consistently outperform baseline methods on PERSONA-LEAKAGE 1 .Moreover, we conduct analysis on the behavior of recently proposed personalized chit-chat dialogue systems.The empirical results show that those systems suffer more from personal information disclosure than the widely used Seq2Seq model and the language model.In those cases, a significant number of information leaking utterances can be detected by our models with high precision.
Qiongkai Xu, Lizhen Qu, Gholamreza Haffari
EMNLP (1)1
2019 Privacy-Aware Text Rewriting
abstract
Biased decisions made by automatic systems have led to growing concerns in research communities. Recent work from the NLP community focuses on building systems that make fair decisions based on text. Instead of relying on unknown decision systems or human decision-makers, we argue that a better way to protect data providers is to remove the trails of sensitive information before publishing the data. In light of this, we propose a new privacy-aware text rewriting task and explore two privacy-aware back-translation methods for the task, based on adversarial training and approximate fairness risk. Our extensive experiments on three real-world datasets with varying demo-graphical attributes show that our methods are effective in obfuscating sensitive attributes. We have also observed that the fairness risk method retains better semantics and fluency, while the adversarial training method tends to leak less sensitive information.
Qiongkai Xu, Lizhen Qu, Ran Cui
INLG1
2017 Attentive Graph-based Recursive Neural Network for Collective Vertex Classification
abstract
Vertex classification is a critical task in graph analysis, where both contents and linkage of vertices are incorporated during classification. Recently, researchers proposed using deep neural network to build an end-to-end framework, which can capture both local content and structure information. These approaches were proved effective in incorporating semantic meanings of neighbouring vertices, while the usefulness of this information was not properly considered. In this paper, we propose an Attentive Graph-based Recursive Neural Network (AGRNN), which exerts attention on neural network to make our model focus on vertices with more relevant semantic information. We evaluated our approach on three real-world datasets and also datasets with synthetic noise. Our experimental results show that AGRNN achieves the state-of-the-art performance, in terms of effectiveness and robustness. We have also illustrated some attention weight samples to demonstrate the rationality of our model.
Qiongkai Xu, Qing Wang 0002, Lizhen Qu
CIKM1
2016 Deep Neural Networks for Learning Graph Representations
abstract
In this paper, we propose a novel model for learning graph representations, which generates a low-dimensional vector representation for each vertex by capturing the graph structural information. Different from other previous research efforts, we adopt a random surfing model to capture graph structural information directly, instead of using the sampling-based method for generating linear sequences proposed by Perozzi et al. (2014). The advantages of our approach will be illustrated from both theorical and empirical perspectives. We also give a new perspective for the matrix factorization method proposed by Levy and Goldberg (2014), in which the pointwise mutual information (PMI) matrix is considered as an analytical solution to the objective function of the skip-gram model with negative sampling proposed by Mikolov et al. (2013). Unlike their approach which involves the use of the SVD for finding the low-dimensitonal projections from the PMI matrix, however, the stacked denoising autoencoder is introduced in our model to extract complex features and model non-linearities. To demonstrate the effectiveness of our model, we conduct experiments on clustering and visualization tasks, employing the learned vertex representations as features. Empirical results on datasets of varying sizes show that our model outperforms other stat-of-the-art models in such tasks.
Shaosheng Cao, Wei Lu 0011, Qiongkai Xu
AAAI3
2016 Semantic Documents Relatedness using Concept Graph Representation
abstract
We deal with the problem of document representation for the task of measuring semantic relatedness between documents. A document is represented as a compact concept graph where nodes represent concepts extracted from the document through references to entities in a knowledge base such as DBpedia. Edges represent the semantic and structural relationships among the concepts. Several methods are presented to measure the strength of those relationships. Concepts are weighted through the concept graph using closeness centrality measure which reflects their relevance to the aspects of the document. A novel similarity measure between two concept graphs is presented. The similarity measure first represents concepts as continuous vectors by means of neural networks. Second, the continuous vectors are used to accumulate pairwise similarity between pairs of concepts while considering their assigned weights. We evaluate our method on a standard benchmark for document similarity. Our method outperforms state-of-the-art methods including ESA (Explicit Semantic Annotation) while our concept graphs are much smaller than the concept vectors generated by ESA. Moreover, we show that by combining our concept graph with ESA, we obtain an even further improvement.
Yuan Ni, Qiongkai Xu, Yosi Mass, Dafna Sheinwald, Huijia Zhu, Shao Sheng Cao
WSDM2
2015 GraRep: Learning Graph Representations with Global Structural Information
abstract
In this paper, we present {GraRep}, a novel model for learning vertex representations of weighted graphs. This model learns low dimensional vectors to represent vertices appearing in a graph and, unlike existing work, integrates global structural information of the graph into the learning process. We also formally analyze the connections between our work and several previous research efforts, including the DeepWalk model of Perozzi et al. as well as the skip-gram model with negative sampling of Mikolov et al.
Shaosheng Cao, Wei Lu 0011, Qiongkai Xu
CIKM3