VLDB 2026 Research / reviewers in the wild / expert
Wasi Uddin Ahmad
dblp:183/0576
· DBLP profile ↗
29ranked-venue papers
13as first author
21since 2021 · last 2026
0000-0001-8171-9583ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 10 first-author · 20 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight ModelsabstractMehrzad Samadi, Aleksander Ficek, Sean Narenthiran, Siddhartha Jain, Wasi Uddin Ahmad, Somshubra Majumdar, Vahid Noroozi, Boris Ginsburg. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Mehrzad Samadi, Aleksander Ficek, Sean Narenthiran, Siddhartha Jain 0001, Wasi Uddin Ahmad, Somshubra Majumdar, Vahid Noroozi, Boris Ginsburg |
ACL (1) | 5 |
| 2025 | LibEvolutionEval: A Benchmark and Study for Version-Specific Code GenerationabstractSachit Kuhar, Wasi Uddin Ahmad, Zijian Wang, Nihal Jain, Haifeng Qian, Baishakhi Ray, Murali Krishna Ramanathan, Xiaofei Ma, Anoop Deoras. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sachit Kuhar, Wasi Uddin Ahmad, Zijian Wang 0002, Nihal Jain, Haifeng Qian, Baishakhi Ray, Murali Krishna Ramanathan, Xiaofei Ma 0001, Anoop Deoras |
NAACL (Long Papers) | 2 |
| 2024 | On Leveraging Encoder-only Pre-trained Language Models for Effective Keyphrase GenerationabstractThis study addresses the application of encoder-only Pre-trained Language Models (PLMs) in keyphrase generation (KPG) amidst the broader availability of domain-tailored encoder-only models compared to encoder-decoder models. We investigate three core inquiries: (1) the efficacy of encoder-only PLMs in KPG, (2) optimal architectural decisions for employing encoder-only PLMs in KPG, and (3) a performance comparison between in-domain encoder-only and encoder-decoder PLMs across varied resource settings. Our findings, derived from extensive experimentation in two domains reveal that with encoder-only PLMs, although keyphrase extraction with Conditional Random Fields slightly excels in identifying present keyphrases, the KPG formulation renders a broader spectrum of keyphrase predictions. Additionally, prefix-LM fine-tuning of encoder-only PLMs emerges as a strong and data-efficient strategy for KPG, outperforming general-domain seq2seq PLMs. We also identify a favorable parameter allocation towards model depth rather than width when employing encoder-decoder architectures initialized with encoder-only PLMs. The study sheds light on the potential of utilizing encoder-only PLMs for advancing KPG systems and provides a groundwork for future KPG methods. Our code and pre-trained checkpoints are released at https://github.com/uclanlp/DeepKPG. Di Wu 0054, Wasi Uddin Ahmad, Kai-Wei Chang 0001 |
LREC/COLING | 2 |
| 2024 | CoCoMIC: Code Completion by Jointly Modeling In-file and Cross-file ContextabstractWhile pre-trained language models (LM) for code have achieved great success in code completion, they generate code conditioned only on the contents within the file, i.e., in-file context, but ignore the rich semantics in other files within the same project, i.e., project-level cross-file context, a critical source of information that is especially useful in modern modular software development. Such overlooking constrains code LMs’ capacity in code completion, leading to unexpected behaviors such as generating hallucinated class member functions or function calls with unexpected arguments. In this work, we propose CoCoMIC, a novel framework that jointly learns the in-file and cross-file context on top of code LMs. To empower CoCoMIC, we develop CCFinder, a static-analysis-based tool that locates and retrieves the most relevant project-level cross-file context for code completion. CoCoMIC successfully improves the existing code LM with a 33.94% relative increase in exact match and 28.69% in identifier matching for code completion when the cross-file context is provided. Finally, we perform a series of ablation studies and share valuable insights for future research on integrating cross-file context into code LMs. Yangruibo Ding, Zijian Wang 0002, Wasi Uddin Ahmad, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth 0001, Bing Xiang |
LREC/COLING | 3 |
| 2024 | Code Representation Learning at ScaleabstractRecent studies have shown that code language model at scale demonstrate significant performance gains on downstream tasks, i.e., code generation. However, most of the existing works on code representation learning train models at a hundred million parameter scale using very limited pretraining corpora. In this work, we fuel code representation learning with a vast amount of code data via a two-stage pretraining scheme. We first train the encoders via a mix that leverages both randomness in masking language modeling and implicit structure and semantic aspects of programming language. We then enhance the representations via contrastive learning with hard negative and hard positive constructed in an unsupervised manner. We establish an off-the-shelf encoder model that persistently outperforms the existing models on a wide variety of downstream tasks by large margins. To comprehend the factors contributing to successful code representation learning, we conduct detailed ablations and share our findings on (i) a customized and effective token-level denoising scheme for source code; (ii) the importance of hard negatives and hard positives; (iii) how the proposed bimodal contrastive learning boost the cross-lingual semantic search performance; and (iv) how the pretraining schemes decide the downstream task performance scales with the model size. Dejiao Zhang, Wasi Uddin Ahmad, Hantian Ding, Ramesh Nallapati, Dan Roth 0001, Xiaofei Ma 0001, Bing Xiang |
ICLR | 2 |
| 2024 | Repoformer: Selective Retrieval for Repository-Level Code CompletionabstractRecent advances in retrieval-augmented generation (RAG) have initiated a new era in repository-level code completion. However, the invariable use of retrieval in existing methods exposes issues in both efficiency and robustness, with a large proportion of the retrieved contexts proving unhelpful or harmful to code language models (code LMs). In this paper, we propose a selective RAG framework to avoid retrieval when unnecessary. To power this framework, we design a self-supervised learning approach to enable a code LM to accurately self-evaluate whether retrieval can improve its output quality and robustly leverage the potentially noisy retrieved contexts. Using this LM as both the selective RAG policy and the generation model, our framework achieves state-of-the-art repository-level code completion performance on diverse benchmarks including RepoEval, CrossCodeEval, and CrossCodeLongEval, a new long-form code completion benchmark. Meanwhile, our analyses show that selectively retrieving brings as much as 70% inference speedup in the online serving setting without harming the performance. We further demonstrate that our framework is able to accommodate different generation models, retrievers, and programming languages. These advancements position our framework as an important step towards more accurate and efficient repository-level code completion. Di Wu 0054, Wasi Uddin Ahmad, Dejiao Zhang, Murali Krishna Ramanathan, Xiaofei Ma 0001 |
ICML | 2 |
| 2023 | CrossSum: Beyond English-Centric Cross-Lingual Summarization for 1, 500+ Language PairsabstractAbhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, Yuan-Fang Li, Yong-Bin Kang, Rifat Shahriyar. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Abhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, Yuan-Fang Li, Yong-Bin Kang, Rifat Shahriyar |
ACL (1) | 3 |
| 2023 | ContraCLM: Contrastive Learning For Causal Language ModelabstractNihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang, Feng Nan, Xiaopeng Li, Ming Tan, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Xiaofei Ma, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Nihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang 0002, Feng Nan, Xiaopeng Li 0002, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Xiaofei Ma 0001, Bing Xiang |
ACL (1) | 3 |
| 2023 | Summarize and Generate to Back-translate: Unsupervised Translation of Programming LanguagesabstractBack-translation is widely known for its effectiveness in neural machine translation when there is little to no parallel data.In this approach, a source-to-target model is coupled with a target-to-source model trained in parallel.The target-to-source model generates noisy sources, while the source-to-target model is trained to reconstruct the targets and vice versa.Recent developments of multilingual pre-trained sequence-to-sequence models for programming languages have been very effective for a broad spectrum of downstream software engineering tasks.Hence, training them to build programming language translation systems via back-translation is compelling.However, these models cannot be further trained via back-translation since they learn to output sequences in the same language as the inputs during pre-training.As an alternative, we propose performing back-translation via code summarization and generation.In code summarization, a model learns to generate natural language (NL) summaries given code snippets.In code generation, the model learns to do the opposite.Therefore, target-to-source generation in back-translation can be viewed as a target-to-NL-to-source generation.We show that our proposed approach performs competitively with state-of-the-art methods.We have made the code publicly available.1 Wasi Uddin Ahmad, Saikat Chakraborty 0001, Baishakhi Ray, Kai-Wei Chang 0001 |
EACL | 1 |
| 2023 | Retrieval Enhanced Data Augmentation for Question Answering on Privacy PoliciesabstractPrior studies in privacy policies frame the question answering (QA) task as identifying the most relevant text segment or a list of sentences from a policy document given a user query.Existing labeled datasets are heavily imbalanced (only a few relevant segments), limiting the QA performance in this domain.In this paper, we develop a data augmentation framework based on ensembling retriever models that captures the relevant text segments from unlabeled policy documents and expand the positive examples in the training set.In addition, to improve the diversity and quality of the augmented data, we leverage multiple pre-trained language models (LMs) and cascade them with noise reduction filter models.Using our augmented data on the PrivacyQA benchmark, we elevate the existing baseline by a large margin (10% F1) and achieve a new state-of-the-art F1 score of 50%.Our ablation studies provide further insights into the effectiveness of our approach. Md. Rizwan Parvez, Jianfeng Chi, Wasi Uddin Ahmad, Yuan Tian 0001, Kai-Wei Chang 0001 |
EACL | 3 |
| 2023 | Rethinking Model Selection and Decoding for Keyphrase Generation with Pre-trained Sequence-to-Sequence ModelsabstractKeyphrase Generation (KPG) is a longstanding task in NLP with widespread applications.The advent of sequence-to-sequence (seq2seq) pre-trained language models (PLMs) has ushered in a transformative era for KPG, yielding promising performance improvements.However, many design decisions remain unexplored and are often made arbitrarily.This paper undertakes a systematic analysis of the influence of model selection and decoding strategies on PLM-based KPG.We begin by elucidating why seq2seq PLMs are apt for KPG, anchored by an attention-driven hypothesis.We then establish that conventional wisdom for selecting seq2seq PLMs lacks depth: (1) merely increasing model size or performing task-specific adaptation is not parameter-efficient; (2) although combining in-domain pre-training with task adaptation benefits KPG, it does partially hinder generalization.Regarding decoding, we demonstrate that while greedy search achieves strong F1 scores, it lags in recall compared with samplingbased methods.Based on these insights, we propose DESEL, a likelihood-based decodeselect algorithm for seq2seq PLMs.DESEL improves greedy search by an average of 4.7% semantic F1 across five datasets.Our collective findings pave the way for deeper future investigations into PLM-based KPG. Di Wu 0054, Wasi Uddin Ahmad, Kai-Wei Chang 0001 |
EMNLP | 2 |
| 2023 | Multi-lingual Evaluation of Code Generation Models
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang 0002, Xiaopeng Li 0002, Wasi Uddin Ahmad, Shiqi Wang 0002, Qing Sun 0013, Mingyue Shang, Sujan K. Gonugondla, Hantian Ding, Nathan Fulton, Arash Farahani, Siddhartha Jain 0001, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, Ramesh Nallapati |
ICLR | 7 |
| 2023 | CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code CompletionabstractCode completion models have made significant progress in recent years, yet current popular evaluation datasets, such as HumanEval and MBPP, predominantly focus on code completion tasks within a single file. This over-simplified setting falls short of representing the real-world software development scenario where repositories span multiple files with numerous cross-file dependencies, and accessing and understanding cross-file context is often required to complete the code correctly. To fill in this gap, we propose CrossCodeEval, a diverse and multilingual code completion benchmark that necessitates an in-depth cross-file contextual understanding to complete the code accurately. CrossCodeEval is built on a diverse set of real-world, open-sourced, permissively-licensed repositories in four popular programming languages: Python, Java, TypeScript, and C#. To create examples that strictly require cross-file context for accurate completion, we propose a straightforward yet efficient static-analysis-based approach to pinpoint the use of cross-file context within the current file. Extensive experiments on state-of-the-art code language models like CodeGen and StarCoder demonstrate that CrossCodeEval is extremely challenging when the relevant cross-file context is absent, and we see clear improvements when adding these context into the prompt. However, despite such improvements, the pinnacle of performance remains notably unattained even with the highest-performing model, indicating that CrossCodeEval is also capable of assessing model's capability in leveraging extensive context to make better code completion. Finally, we benchmarked various methods in retrieving cross-file context, and show that CrossCodeEval can also be used to measure the capability of code retrievers. Yangruibo Ding, Zijian Wang 0002, Wasi Uddin Ahmad, Hantian Ding, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth 0001, Bing Xiang |
NeurIPS | 3 |
| 2023 | Towards Greener Yet Powerful Code Generation via Quantization: An Empirical StudyabstractML-powered code generation aims to assist developers to write code in a more productive manner by intelligently generating code blocks based on natural language prompts. Recently, large pretrained deep learning models have pushed the boundary of code generation and achieved impressive performance. However, the huge number of model parameters poses a significant challenge to their adoption in a typical software development environment, where a developer might use a standard laptop or mid-size server to develop code. Such large models cost significant resources in terms of memory, latency, dollars, as well as carbon footprint. Xiaokai Wei, Sujan K. Gonugondla, Shiqi Wang 0002, Wasi Uddin Ahmad, Baishakhi Ray, Haifeng Qian, Xiaopeng Li 0002, Zijian Wang 0002, Qing Sun 0013, Ben Athiwaratkun, Mingyue Shang, Murali Krishna Ramanathan, Parminder Bhatia, Bing Xiang |
ESEC/SIGSOFT FSE | 4 |
| 2021 | GATE: Graph Attention Transformer Encoder for Cross-lingual Relation and Event ExtractionabstractRecent progress in cross-lingual relation and event extraction use graph convolutional networks (GCNs) with universal dependency parses to learn language-agnostic sentence representations such that models trained on one language can be applied to other languages. However, GCNs struggle to model words with long-range dependencies or are not directly connected in the dependency tree. To address these challenges, we propose to utilize the self-attention mechanism where we explicitly fuse structural information to learn the dependencies between words with different syntactic distances. We introduce GATE, a Graph Attention Transformer Encoder, and test its cross-lingual transferability on relation and event extraction tasks. We perform experiments on the ACE05 dataset that includes three typologically different languages: English, Chinese, and Arabic. The evaluation results show that GATE outperforms three recently proposed methods by a large margin. Our detailed analysis reveals that due to the reliance on syntactic dependencies, GATE produces robust representations that facilitate transfer across languages. Wasi Uddin Ahmad, Nanyun Peng 0001, Kai-Wei Chang 0001 |
AAAI | 1 |
| 2021 | Simple or Complex? Learning to Predict Readability of Bengali TextsabstractDetermining the readability of a text is the first step to its simplification. In this paper, we present a readability analysis tool capable of analyzing text written in the Bengali language to provide in-depth information on its readability and complexity. Despite being the 7th most spoken language in the world with 230 million native speakers, Bengali suffers from a lack of fundamental resources for natural language processing. Readability related research of the Bengali language so far can be considered to be narrow and sometimes faulty due to the lack of resources. Therefore, we correctly adopt document-level readability formulas traditionally used for U.S. based education system to the Bengali language with a proper age-to-age comparison. Due to the unavailability of large-scale human-annotated corpora, we further divide the document-level task into sentence-level and experiment with neural architectures, which will serve as a baseline for the future works of Bengali readability prediction. During the process, we present several human-annotated corpora and dictionaries such as a document-level dataset comprising 618 documents with 12 different grade levels, a large-scale sentence-level dataset comprising more than 96K sentences with simple and complex labels, a consonant conjunct count algorithm and a corpus of 341 words to validate the effectiveness of the algorithm, a list of 3,396 easy words, and an updated pronunciation dictionary with more than 67K words. These resources can be useful for several other tasks of this low-resource language. Susmoy Chakraborty, Mir Tafseer Nayeem, Wasi Uddin Ahmad |
AAAI | 3 |
| 2021 | Select, Extract and Generate: Neural Keyphrase Generation with Layer-wise Coverage AttentionabstractWasi Ahmad, Xiao Bai, Soomin Lee, Kai-Wei Chang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wasi Uddin Ahmad, Xiao Bai 0002, Kai-Wei Chang 0001 |
ACL/IJCNLP (1) | 1 |
| 2021 | Intent Classification and Slot Filling for Privacy PoliciesabstractWasi Ahmad, Jianfeng Chi, Tu Le, Thomas Norton, Yuan Tian, Kai-Wei Chang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wasi Uddin Ahmad, Jianfeng Chi, Tu Le, Thomas Norton, Yuan Tian 0001, Kai-Wei Chang 0001 |
ACL/IJCNLP (1) | 1 |
| 2021 | Syntax-augmented Multilingual BERT for Cross-lingual TransferabstractIn recent years, we have seen a colossal effort in pre-training multilingual text encoders using large-scale corpora in many languages to facilitate cross-lingual transfer learning. However, due to typological differences across languages, the cross-lingual transfer is challenging. Nevertheless, language syntax, e.g., syntactic dependencies, can bridge the typological gap. Previous works have shown that pre-trained multilingual encoders, such as mBERT (CITATION), capture language syntax, helping cross-lingual transfer. This work shows that explicitly providing language syntax and training mBERT using an auxiliary objective to encode the universal dependency tree structure helps cross-lingual transfer. We perform rigorous experiments on four NLP tasks, including text classification, question answering, named entity recognition, and task-oriented semantic parsing. The experiment results show that syntax-augmented mBERT improves cross-lingual transfer on popular benchmarks, such as PAWS-X and MLQA, by 1.4 and 1.6 points on average across all languages. In the generalized transfer setting, the performance boosted significantly, with 3.9 and 3.1 points on average in PAWS-X and MLQA. Wasi Uddin Ahmad, Haoran Li 0007, Kai-Wei Chang 0001, Yashar Mehdad |
ACL/IJCNLP (1) | 1 |
| 2021 | Improving Zero-Shot Cross-Lingual Transfer Learning via Robust TrainingabstractPre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potential for zero-shot cross-lingual transfer.However, these multilingual encoders do not precisely align words and phrases across languages.Especially, learning alignments in the multilingual embedding space usually requires sentence-level or word-level parallel corpora, which are expensive to be obtained for low-resource languages.An alternative is to make the multilingual encoders more robust; when fine-tuning the encoder using downstream task, we train the encoder to tolerate noise in the contextual embedding spaces such that even if the representations of different languages are not aligned well, the model can still achieve good performance on zero-shot cross-lingual transfer.In this work, we propose a learning strategy for training robust models by drawing connections between adversarial examples and the failure cases of zero-shot cross-lingual transfer.We adopt two widely used robust training methods, adversarial training and randomized smoothing, to train the desired robust model.The experimental results demonstrate that robust training improves zero-shot cross-lingual transfer on text classification tasks.The improvement is more significant in the generalized crosslingual transfer setting, where the pair of input sentences belong to two different languages. Kuan-Hao Huang, Wasi Uddin Ahmad, Nanyun Peng 0001, Kai-Wei Chang 0001 |
EMNLP (1) | 2 |
| 2021 | Unified Pre-training for Program Understanding and GenerationabstractWasi Ahmad, Saikat Chakraborty, Baishakhi Ray, Kai-Wei Chang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Wasi Uddin Ahmad, Saikat Chakraborty 0001, Baishakhi Ray, Kai-Wei Chang 0001 |
NAACL-HLT | 1 |
| 2020 | A Transformer-based Approach for Source Code SummarizationabstractGenerating a readable summary that describes the functionality of a program is known as source code summarization.In this task, learning code representation by modeling the pairwise relationship between code tokens to capture their long-range dependencies is crucial.To learn code representation for summarization, we explore the Transformer model that uses a self-attention mechanism and has shown to be effective in capturing long-range dependencies.In this work, we show that despite the approach is simple, it outperforms the state-of-the-art techniques by a significant margin.We perform extensive analysis and ablation studies that reveal several important findings, e.g., the absolute encoding of source code tokens' position hinders, while relative encoding significantly improves the summarization performance.We have made our code publicly available 1 to facilitate future research. Wasi Uddin Ahmad, Saikat Chakraborty 0001, Baishakhi Ray, Kai-Wei Chang 0001 |
ACL | 1 |
| 2019 | Cross-Lingual Dependency Parsing with Unlabeled Auxiliary LanguagesabstractCross-lingual transfer learning has become an important weapon to battle the unavailability of annotated resources for low-resource languages.One of the fundamental techniques to transfer across languages is learning language-agnostic representations, in the form of word embeddings or contextual encodings.In this work, we propose to leverage unannotated sentences from auxiliary languages to help learning language-agnostic representations.Specifically, we explore adversarial training for learning contextual encoders that produce invariant representations across languages to facilitate cross-lingual transfer.We conduct experiments on cross-lingual dependency parsing where we train a dependency parser on a source language and transfer it to a wide range of target languages.Experiments on 28 target languages demonstrate that adversarial training significantly improves the overall transfer performances under several different settings.We conduct a careful analysis to evaluate the language-agnostic representations resulted from adversarial training. Wasi Uddin Ahmad, Zhisong Zhang, Xuezhe Ma, Kai-Wei Chang 0001, Nanyun Peng 0001 |
CoNLL | 1 |
| 2019 | Context Attentive Document Ranking and Query SuggestionabstractWe present a context-aware neural ranking model to exploit users' on-task search activities and enhance retrieval performance. In particular, a two-level hierarchical recurrent neural network is introduced to learn search context representation of individual queries, search tasks, and corresponding dependency structure by jointly optimizing two companion retrieval tasks: document ranking and query suggestion. To identify variable dependency structure between search context and users' ongoing search activities, attention at both levels of recurrent states are introduced. Extensive experiment comparisons against a rich set of baseline methods and an in-depth ablation analysis confirm the value of our proposed approach for modeling search context buried in search tasks. Wasi Uddin Ahmad, Kai-Wei Chang 0001, Hongning Wang |
SIGIR | 1 |
| 2018 | Multi-Task Learning for Document Ranking and Query Suggestion
Wasi Uddin Ahmad, Kai-Wei Chang 0001, Hongning Wang |
ICLR (Poster) | 1 |
| 2018 | A Corpus to Learn Refer-to-as Relations for Nominals
Wasi Uddin Ahmad, Kai-Wei Chang 0001 |
LREC | 1 |
| 2018 | Intent-aware Query Obfuscation for Privacy Protection in Personalized Web SearchabstractModern web search engines exploit users' search history to personalize search results, with a goal of improving their service utility on a per-user basis. But it is this very dimension that leads to the risk of privacy infringement and raises serious public concerns. In this work, we propose a client-centered intent-aware query obfuscation solution for protecting user privacy in a personalized web search scenario. In our solution, each user query is submitted with l additional cover queries and corresponding clicks, which act as decoys to mask users' genuine search intent from a search engine. The cover queries are sequentially sampled from a set of hierarchically organized language models to ensure the coherency of fake search intents in a cover search task. Our approach emphasizes the plausibility of generated cover queries, not only to the current genuine query but also to previous queries in the same task, to increase the complexity for a search engine to identify a user's true intent. We also develop two new metrics from an information theoretic perspective to evaluate the effectiveness of provided privacy protection. Comprehensive experiment comparisons with state-of-the-art query obfuscation techniques are performed on the public AOL search log, and the propitious results substantiate the effectiveness of our solution. Wasi Uddin Ahmad, Kai-Wei Chang 0001, Hongning Wang |
SIGIR | 1 |
| 2018 | Hide-n-Seek: An Intent-aware Privacy Protection Plugin for Personalized Web SearchabstractWe develop Hide-n-Seek, an intent-aware privacy protection plugin for personalized web search. In addition to users' genuine search queries, Hide-n-Seek submits k cover queries and corresponding clicks to an external search engine to disguise a user's search intent grounded and reinforced in a search session by mimicking the true query sequence. The cover queries are synthesized and randomly sampled from a topic hierarchy, where each node represents a coherent search topic estimated by both n-gram and neural language models constructed over crawled web documents. Hide-n-Seek also personalizes the returned search results by re-ranking them based on the genuine user profile developed and maintained on the client side. With a variety of graphical user interfaces, we present the topic-based query obfuscation mechanism to the end users for them to digest how their search privacy is protected. Puxuan Yu, Wasi Uddin Ahmad, Hongning Wang |
SIGIR | 2 |
| 2016 | Topic Model based Privacy Protection in Personalized Web SearchabstractModern search engines utilize users' search history for personalization, which provides more effective, useful and relevant search results. However, it also has the potential risk of revealing users' privacy by identifying their underlying intention from their logged search behaviors. To address this privacy issue, we proposed a Topic-based Privacy Protection solution on client side. In our solution, each user query will be submitted with k additional cover queries, which will act as a proxy to disguise users' intent from a search engine. The set of cover queries are generated in a controlled way so that each query carries similar uncertainty to randomize a user's search history while still providing necessary utility for the search engine to perform personalization. We used statistical topic models to infer topics from the original user query and generated cover queries of similar entropy but from unrelated topics. Extensive experiments are performed on AOL search log and the promising results demonstrated the effectiveness of our solution. Wasi Uddin Ahmad, Md. Masudur Rahman 0001, Hongning Wang |
SIGIR | 1 |