VLDB 2026 Research / reviewers in the wild / expert
Ahmed Awadallah 0001
dblp:147/9148 · also Ahmed Hassan Awadallah
· DBLP profile ↗
107ranked-venue papers
19as first author
34since 2021 · last 2025
0000-0001-6426-3537ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 73 · 17 first-author · 32 since 2021Databases, data management, data science and information retrieval · 58 · 13 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-authorSoftware engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sweeping Heterogeneity with Smart MoPs: Mixture of Prompts for LLM Task AdaptationabstractPrompt instruction tuning is a popular approach to better adjust pretrained LLMs for specific downstream tasks. How to extend this approach to simultaneously handle multiple tasks and data distributions is an interesting question. We propose Mixture of Prompts (MoPs) with smart gating functionality. Our proposed system identifies relevant skills embedded in different groups of prompts and dynamically weighs experts (i.e., collection of prompts) based on the target task. Experiments show that MoPs are resilient to model compression, data source, and task composition, making them highly versatile and applicable in various contexts. In practice, MoPs can simultaneously mitigate prompt training ``interference'' in multi-task, multi-source scenarios (e.g., task and data heterogeneity across sources) and possible implications from model approximations. Empirically, MoPs show particular effectiveness in compressed model scenarios, while maintaining favorable performance in uncompressed settings: MoPs can reduce final perplexity from 9% up to 70% in non-i.i.d. distributed cases and from 3% up to 30% in centralized cases, compared to baselines. Chen Dun, Mirian Hipolito Garcia, Guoqing Zheng, Ahmed Awadallah 0001, Robert Sim, Anastasios Kyrillidis |
AAAI | 4 |
| 2025 | Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHFabstractThis paper investigates a basic question in reinforcement learning from human feedback (RLHF) from a theoretical perspective: how to efficiently explore in an online manner under preference feedback and general function approximation. We take the initial step towards a theoretical understanding of this problem by proposing a novel algorithm, *Exploratory Preference Optimization* (XPO). This algorithm is elegantly simple---requiring only a one-line modification to (online) Direct Preference Optimization (DPO; Rafailov et al., 2023)---yet provides the strongest known provable guarantees. XPO augments the DPO objective with a novel and principled *exploration bonus*, enabling the algorithm to strategically explore beyond the support of the initial model and preference feedback data. We prove that XPO is provably sample-efficient and converges to a near-optimal policy under natural exploration conditions, regardless of the initial model's coverage. Our analysis builds on the observation that DPO implicitly performs a form of *Bellman error minimization*. It synthesizes previously disparate techniques from language modeling and theoretical reinforcement learning in a serendipitous fashion through the lens of *KL-regularized Markov decision processes*. Tengyang Xie, Dylan J. Foster, Akshay Krishnamurthy, Corby Rosset, Ahmed Awadallah 0001, Alexander Rakhlin |
ICLR | 5 |
| 2025 | Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized SettingsabstractMixture-of-Experts (MoEs) achieve scalability by dynamically activating subsets of their components.
Yet, understanding how expertise emerges through joint training of gating mechanisms and experts remains incomplete, especially in scenarios without clear task partitions. Motivated by inference costs and data heterogeneity, we study how joint training of gating functions and experts can dynamically allocate domain-specific expertise across multiple underlying data distributions.
As an outcome of our framework, we develop an instance tailored specifically to decentralized training scenarios, introducing *Dynamically Decentralized Orchestration of MoEs* or *DDOME*. *DDOME* leverages heterogeneity emerging from distributional shifts across decentralized data sources to specialize experts dynamically. By integrating a pretrained common expert to inform a gating function, *DDOME* achieves personalized expert subset selection on-the-fly, facilitating just-in-time personalization.
We empirically validate *DDOME* within a Federated Learning (FL) context: *DDOME* attains from 4\% up to an 24\% accuracy improvement over state-of-the-art FL baselines in image and text classification tasks, while maintaining competitive zero-shot generalization capabilities. Furthermore, we provide theoretical insights confirming that the joint gating-experts training is critical for achieving meaningful expert specialization. Yehya Farhat, Hamza ElMokhtar Shili, Fangshuo Liao, Chen Dun, Mirian Hipolito Garcia, Guoqing Zheng, Ahmed Awadallah 0001, Robert Sim, Dimitrios Dimitriadis, Anastasios Kyrillidis |
NeurIPS | 7 |
| 2025 | Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for Deep ResearchabstractExisting question answering (QA) datasets are no longer challenging to most powerful Large Language Models (LLMs). Traditional QA benchmarks like TriviaQA, NaturalQuestions, ELI5 and HotpotQA mainly study ''known unknowns'' with clear indications of both what information is missing, and how to find it to answer the question. A yet unmet need of the NLP community is a bank of non-factoid, multi-perspective questions involving a great deal of unclear information needs, i.e. ''unknown unknowns''. We claim we can find such questions in search engine logs, which is surprising because most question-intent queries are indeed factoid. Furthermore, recent products like Google's DeepResearch (announced a year after this resource was released publicly) specifically address such queries, retrieving hundreds of documents to synthesize report-style responses. We present Researchy Questions, the world's first, only and largest public dataset of ''Deep Research'' questions filtered from real search engine logs to be non-factoid, ''decompositional'' and multi-perspective. We show that users spend substantial ''effort'' on these questions in terms of signals like clicks and session length. We also show that ''slow thinking'' answering techniques, like decomposition into sub-questions shows benefit over answering directly. We release (at https://huggingface.co/datasets/corbyrosset/researchy_questions) about 100k Researchy Questions with a permissive CDLA-2.0 license, along with click histograms on over 350k Clueweb22 URLs that were clicked for each question. Corbin Rosset, Ho-Lam Chung, Guanghui Qin, Ethan C. Chau, Ahmed Awadallah 0001, Jennifer Neville, Nikhil Rao 0001 |
SIGIR | 6 |
| 2024 | Assessing and Verifying Task Utility in LLM-Powered ApplicationsabstractNegar Arabzadeh, Siqing Huo, Nikhil Mehta, Qingyun Wu, Chi Wang, Ahmed Hassan Awadallah, Charles L. A. Clarke, Julia Kiseleva. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Negar Arabzadeh, Siqing Huo, Nikhil Mehta 0003, Qingyun Wu, Chi Wang 0001, Ahmed Awadallah 0001, Charles L. A. Clarke, Julia Kiseleva |
EMNLP | 6 |
| 2024 | Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingabstractLarge language models (LLMs) excel in most NLP tasks but also require expensive cloud servers for deployment due to their size, while smaller models that can be deployed on lower cost (e.g., edge) devices, tend to lag behind in terms of response quality. Therefore in this work we propose a hybrid inference approach which combines their respective strengths to save cost and maintain quality. Our approach uses a router that assigns queries to the small or large model based on the predicted query difficulty and the desired quality level. The desired quality level can be tuned dynamically at test time to seamlessly trade quality for cost as per the scenario requirements. In experiments our approach allows us to make up to 40% fewer calls to the large model, with no drop in response quality. Dujian Ding, Ankur Mallick, Chi Wang 0001, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks V. S. Lakshmanan, Ahmed Awadallah 0001 |
ICLR | 8 |
| 2024 | Teaching Language Models to Hallucinate Less with Synthetic TasksabstractLarge language models (LLMs) frequently hallucinate on abstractive summarization tasks such as document-based question-answering, meeting summarization, and clinical report generation, even though all necessary information is included in context. However, optimizing to make LLMs hallucinate less is challenging, as hallucination is hard to efficiently, cheaply, and reliably evaluate at each optimization step. In this work, we show that reducing hallucination on a _synthetic task_ can also reduce hallucination on real-world downstream tasks. Our method, SynTra, first designs a synthetic task where hallucinations are easy to elicit and measure. It next optimizes the LLM's system message via prefix tuning on the synthetic task, then uses the system message on realistic, hard-to-optimize tasks. Across three realistic abstractive summarization tasks, we reduce hallucination for two 13B-parameter LLMs using supervision signal from only a synthetic retrieval task. We also find that optimizing the system message rather than the model weights can be critical; fine-tuning the entire model on the synthetic task can counterintuitively _increase_ hallucination. Overall, SynTra demonstrates that the extra flexibility of working with synthetic data can help mitigate undesired behaviors in practice. Erik Jones, Hamid Palangi, Clarisse Simões, Varun Chandrasekaran, Subhabrata Mukherjee, Arindam Mitra, Ahmed Awadallah 0001, Ece Kamar |
ICLR | 7 |
| 2023 | ADMoE: Anomaly Detection with Mixture-of-Experts from Noisy LabelsabstractExisting works on anomaly detection (AD) rely on clean labels from human annotators that are expensive to acquire in practice. In this work, we propose a method to leverage weak/noisy labels (e.g., risk scores generated by machine rules for detecting malware) that are cheaper to obtain for anomaly detection. Specifically, we propose ADMoE, the first framework for anomaly detection algorithms to learn from noisy labels. In a nutshell, ADMoE leverages mixture-of-experts (MoE) architecture to encourage specialized and scalable learning from multiple noisy sources. It captures the similarities among noisy labels by sharing most model parameters, while encouraging specialization by building "expert" sub-networks. To further juice out the signals from noisy labels, ADMoE uses them as input features to facilitate expert learning. Extensive results on eight datasets (including a proprietary enterprise security dataset) demonstrate the effectiveness of ADMoE, where it brings up to 34% performance improvement over not using it. Also, it outperforms a total of 13 leading baselines with equivalent network parameters and FLOPS. Notably, ADMoE is model-agnostic to enable any neural network-based detection methods to handle noisy labels, where we showcase its results on both multiple-layer perceptron (MLP) and the leading AD method DeepSAD. Yue Zhao 0016, Guoqing Zheng, Subhabrata Mukherjee, Robert McCann, Ahmed Awadallah 0001 |
AAAI | 5 |
| 2023 | DSEE: Dually Sparsity-embedded Efficient Tuning of Pre-trained Language ModelsabstractXuxi Chen, Tianlong Chen, Weizhu Chen, Ahmed Hassan Awadallah, Zhangyang Wang, Yu Cheng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Xuxi Chen, Tianlong Chen 0001, Weizhu Chen, Ahmed Awadallah 0001, Zhangyang Wang, Yu Cheng 0001 |
ACL (1) | 4 |
| 2023 | On Improving Summarization Factual Consistency from Natural Language FeedbackabstractYixin Liu, Budhaditya Deb, Milagro Teruel, Aaron Halfaker, Dragomir Radev, Ahmed Hassan Awadallah. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yixin Liu 0003, Budhaditya Deb, Milagro Teruel, Aaron Halfaker, Dragomir R. Radev, Ahmed Awadallah 0001 |
ACL (1) | 6 |
| 2023 | Robustness Challenges in Model Distillation and Pruning for Natural Language UnderstandingabstractMengnan Du, Subhabrata Mukherjee, Yu Cheng, Milad Shokouhi, Xia Hu, Ahmed Hassan Awadallah. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Mengnan Du, Subhabrata Mukherjee, Yu Cheng 0001, Milad Shokouhi, Xia Ben Hu, Ahmed Awadallah 0001 |
EACL | 6 |
| 2023 | Axiomatic Preference Modeling for Longform Question AnsweringabstractThe remarkable abilities of large language models (LLMs) like GPT-4 partially stem from posttraining processes like Reinforcement Learning from Human Feedback (RLHF) involving human preferences encoded in a reward model.However, these reward models (RMs) often lack direct knowledge of why, or under what principles, the preferences annotations were made.In this study, we identify principles that guide RMs to better align with human preferences, and then develop an axiomatic framework to generate a rich variety of preference signals to uphold them.We use these axiomatic signals to train a model for scoring answers to longform questions.Our approach yields a Preference Model with only about 220M parameters that agrees with gold humanannotated preference labels more often than GPT-4.The contributions of this work include: training a standalone preference model that can score human-and LLM-generated answers on the same scale; developing an axiomatic framework for generating training data pairs tailored to certain principles; and showing that a small amount of axiomatic signals can help small models outperform GPT-4 in preference scoring.We intend to release our model. Corby Rosset, Guoqing Zheng, Victor Dibia, Ahmed Awadallah 0001, Paul N. Bennett |
EMNLP | 4 |
| 2023 | Algorithmic Vibe in Information RetrievalabstractWhen information retrieval systems return a ranked list of results in response to a query, they may be choosing from a large set of candidate results that are equally useful and relevant. This means we might be able to identify a difference between rankers A and B, where ranker A systematically prefers a certain type of relevant results. Ranker A may have this systematic difference (different “vibe”) without having systematically better or worse results according to standard information retrieval metrics. We first show that a vibe difference can exist, comparing two publicly available rankers, where the one that is trained on health-related queries will systematically prefer health-related results, even for non-health queries. We define a vibe metric that lets us see the words that a ranker prefers. We investigate the vibe of search engine clicks vs. human labels. We perform an initial study into correcting for vibe differences to make ranker A more like ranker B via changes in negative sampling during training. Ali Montazeralghaem, Nick Craswell, Ryen W. White, Ahmed Awadallah 0001, Byungki Byun |
WWW | 4 |
| 2022 | SummN: A Multi-Stage Summarization Framework for Long Input Dialogues and DocumentsabstractYusen Zhang, Ansong Ni, Ziming Mao, Chen Henry Wu, Chenguang Zhu, Budhaditya Deb, Ahmed Awadallah, Dragomir Radev, Rui Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yusen Zhang 0001, Ansong Ni, Ziming Mao, Chen Henry Wu, Chenguang Zhu 0001, Budhaditya Deb, Ahmed Awadallah 0001, Dragomir R. Radev, Rui Zhang 0037 |
ACL (1) | 7 |
| 2022 | DYLE: Dynamic Latent Extraction for Abstractive Long-Input SummarizationabstractZiming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang, Rui Zhang, Tao Yu, Budhaditya Deb, Chenguang Zhu, Ahmed Awadallah, Dragomir Radev. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Ziming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang 0001, Rui Zhang 0037, Tao Yu 0009, Budhaditya Deb, Chenguang Zhu 0001, Ahmed Awadallah 0001, Dragomir R. Radev |
ACL (1) | 9 |
| 2022 | The Principle of Diversity: Training Stronger Vision Transformers Calls for Reducing All Levels of RedundancyabstractVision transformers (ViTs) have gained increasing popularity as they are commonly believed to own higher mod-eling capacity and representation flexibility, than traditional convolutional networks. However, it is questionable whether such potential has been fully unleashed in prac-tice, as the learned ViTs often suffer from over-smoothening, yielding likely redundant models. Recent works made pre-liminary attempts to identify and alleviate such redundancy, e.g., via regularizing embedding similarity or re-injecting convolution-like structures. However, a “head-to-toe as-sessment” regarding the extent of redundancy in ViTs, and how much we could gain by thoroughly mitigating such, has been absent for this field. This paper, for the first time, systematically studies the ubiquitous existence of re-dundancy at all three levels: patch embedding, attention map, and weight space. In view of them, we advocate a principle of diversity for training ViTs, by presenting cor-responding regularizers that encourage the representation diversity and coverage at each of those levels, that enabling capturing more discriminative information. Extensive ex-periments on ImageNet with a number of ViT backbones validate the effectiveness of our proposals, largely eliminating the observed ViT redundancy and significantly boosting the model generalization. For example, our diversified DeiT obtains 0.70% ~ 1.76% accuracy boosts on ImageNet with highly reduced similarity. Our codes are fully available in https://github.com/VITA-Group/Diverse-ViT. Tianlong Chen 0001, Zhenyu Zhang 0015, Yu Cheng 0001, Ahmed Awadallah 0001, Zhangyang Wang |
CVPR | 4 |
| 2022 | Scalable Learning to Optimize: A Learned Optimizer Can Train Big Models
Xuxi Chen, Tianlong Chen 0001, Yu Cheng 0001, Weizhu Chen, Ahmed Awadallah 0001, Zhangyang Wang |
ECCV (23) | 5 |
| 2022 | DnA: Improving Few-Shot Transfer Learning with Low-Rank Decomposition and Alignment
Ziyu Jiang, Tianlong Chen 0001, Xuxi Chen, Yu Cheng 0001, Luowei Zhou, Lu Yuan 0001, Ahmed Awadallah 0001, Zhangyang Wang |
ECCV (20) | 7 |
| 2022 | Boosting Natural Language Generation from Instructions with Meta-LearningabstractRecent work has shown that language models (LMs) trained with multi-task instructional learning (MTIL) can solve diverse NLP tasks in zero-and few-shot settings with improved performance compared to prompt tuning.MTIL illustrates that LMs can extract and use information about the task from instructions beyond the surface patterns of the inputs and outputs.This suggests that meta-learning may further enhance the utilization of instructions for effective task transfer.In this paper we investigate whether meta-learning applied to MTIL can further improve generalization to unseen tasks in a zero-shot setting.Specifically, we propose to adapt meta-learning to MTIL in three directions: 1) Model Agnostic Meta Learning (MAML), 2) Hyper-Network (HNet) based adaptation to generate task specific parameters conditioned on instructions, and 3) an approach combining HNet and MAML.Through extensive experiments on the large scale Natural Instructions V2 dataset, we show that our proposed approaches significantly improve over strong baselines in zero-shot settings.In particular, meta-learning improves the effectiveness of instructions and is most impactful when the test tasks are strictly zero-shot (i.e.no similar tasks in the training set) and are "hard" for LMs, illustrating the potential of meta-learning for MTIL for out-of-distribution tasks. Budhaditya Deb, Ahmed Awadallah 0001, Guoqing Zheng |
EMNLP | 2 |
| 2022 | Leveraging Locality in Abstractive Text SummarizationabstractNeural attention models have achieved significant improvements on many natural language processing tasks.However, the quadratic memory complexity of the self-attention module with respect to the input length hinders their applications in long text summarization.Instead of designing more efficient attention modules, we approach this problem by investigating if models with a restricted context can have competitive performance compared with the memory-efficient attention models that maintain a global context by treating the input as a single sequence.Our model is applied to individual pages, which contain parts of inputs grouped by the principle of locality, during both the encoding and decoding stages.We empirically investigated three kinds of locality in text summarization at different levels of granularity, ranging from sentences to documents.Our experimental results show that our model has a better performance compared with strong baseline models with efficient attention modules, and our analysis provides further insights into our locality-aware modeling strategy.1 Yixin Liu 0003, Ansong Ni, Linyong Nan, Budhaditya Deb, Chenguang Zhu 0001, Ahmed Awadallah 0001, Dragomir R. Radev |
EMNLP | 6 |
| 2022 | AdaMix: Mixture-of-Adaptations for Parameter-efficient Model TuningabstractYaqing Wang, Sahaj Agarwal, Subhabrata Mukherjee, Xiaodong Liu, Jing Gao, Ahmed Hassan Awadallah, Jianfeng Gao. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Yaqing Wang 0001, Sahaj Agarwal, Subhabrata Mukherjee, Xiaodong Liu 0003, Jing Gao 0004, Ahmed Awadallah 0001, Jianfeng Gao 0001 |
EMNLP | 6 |
| 2022 | Knowledge Infused Decoding
Ruibo Liu, Guoqing Zheng, Radhika Gaonkar, Chongyang Gao, Soroush Vosoughi, Milad Shokouhi, Ahmed Awadallah 0001 |
ICLR | 8 |
| 2022 | WALNUT: A Benchmark on Semi-weakly Supervised Learning for Natural Language UnderstandingabstractGuoqing Zheng, Giannis Karamanolakis, Kai Shu, Ahmed Awadallah. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Guoqing Zheng, Giannis Karamanolakis, Kai Shu, Ahmed Awadallah 0001 |
NAACL-HLT | 4 |
| 2022 | Few-shot Task-agnostic Neural Architecture Search for Distilling Large Language ModelsabstractTraditional knowledge distillation (KD) methods manually design student architectures to compress large models given pre-specified computational cost. This requires several trials to find viable students, and repeating the process with change in computational budget. We use Neural Architecture Search (NAS) to automatically distill several compressed students with variable cost from a large model. Existing NAS methods train a single SuperLM consisting of millions of subnetworks with weight-sharing, resulting in interference between subnetworks of different sizes. Additionally, many of these works are task-specific requiring task labels for SuperLM training. Our framework AutoDistil addresses above challenges with the following steps: (a) Incorporates inductive bias and heuristics to partition Transformer search space into K compact sub-spaces (e.g., K=3 can generate typical student sizes of base, small and tiny); (b) Trains one SuperLM for each sub-space using task-agnostic objective (e.g., self-attention distillation) with weight-sharing of students; (c) Lightweight search for the optimal student without re-training. Task-agnostic training and search allow students to be reused for fine-tuning on any downstream task. Experiments on GLUE benchmark demonstrate AutoDistil to outperform state-of-the-art KD and NAS methods with upto 3x reduction in computational cost and negligible loss in task performance. Code and model checkpoints are available at https://github.com/microsoft/autodistil. Dongkuan Xu, Subhabrata Mukherjee, Xiaodong Liu 0003, Debadeepta Dey, Wenhui Wang 0003, Xiang Zhang 0001, Ahmed Awadallah 0001, Jianfeng Gao 0001 |
NeurIPS | 7 |
| 2021 | Meta Label Correction for Noisy Label LearningabstractLeveraging weak or noisy supervision for building effective machine learning models has long been an important research problem. Its importance has further increased recently due to the growing need for large-scale datasets to train deep learning models. Weak or noisy supervision could originate from multiple sources including non-expert annotators or automatic labeling based on heuristics or user interaction signals. There is an extensive amount of previous work focusing on leveraging noisy labels. Most notably, recent work has shown impressive gains by using a meta-learned instance re-weighting approach where a meta-learning framework is used to assign instance weights to noisy labels. In this paper, we extend this approach via posing the problem as a label correction problem within a meta-learning framework. We view the label correction procedure as a meta-process and propose a new meta-learning based framework termed MLC (Meta Label Correction) for learning with noisy labels. Specifically, a label correction network is adopted as a meta-model to produce corrected labels for noisy labels while the main model is trained to leverage the corrected labels. Both models are jointly trained by solving a bi-level optimization problem. We run extensive experiments with different label noise levels and types on both image recognition and text classification tasks. We compare the re-weighing and correction approaches showing that the correction framing addresses some of the limitations of re-weighting. We also show that the proposed MLC approach outperforms previous methods in both image and language tasks. Guoqing Zheng, Ahmed Awadallah 0001, Susan T. Dumais |
AAAI | 2 |
| 2021 | A Dataset and Baselines for Multilingual Reply SuggestionabstractMozhi Zhang, Wei Wang, Budhaditya Deb, Guoqing Zheng, Milad Shokouhi, Ahmed Hassan Awadallah. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Mozhi Zhang, Wei Wang 0238, Budhaditya Deb, Guoqing Zheng, Milad Shokouhi, Ahmed Awadallah 0001 |
ACL/IJCNLP (1) | 6 |
| 2021 | SCoRe: Pre-Training for Context Representation in Conversational Semantic Parsing
Tao Yu 0009, Rui Zhang 0037, Oleksandr Polozov, Christopher Meek, Ahmed Awadallah 0001 |
ICLR | 5 |
| 2021 | Meta Self-training for Few-shot Neural Sequence LabelingabstractNeural sequence labeling is widely adopted for many Natural Language Processing (NLP) tasks, such as Named Entity Recognition (NER) and slot tagging for dialog systems and semantic parsing. Recent advances with large-scale pre-trained language models have shown remarkable success in these tasks when fine-tuned on large amounts of task-specific labeled data. However, obtaining such large-scale labeled training data is not only costly, but also may not be feasible in many sensitive user applications due to data access and privacy constraints. This is exacerbated for sequence labeling tasks requiring such annotations at token-level. In this work, we develop techniques to address the label scarcity challenge for neural sequence labeling models. Specifically, we propose a meta self-training framework which leverages very few manually annotated labels for training neural sequence models. While self-training serves as an effective mechanism to learn from large amounts of unlabeled data via iterative knowledge exchange -- meta-learning helps in adaptive sample re-weighting to mitigate error propagation from noisy pseudo-labels. Extensive experiments on six benchmark datasets including two for massive multilingual NER and four slot tagging datasets for task-oriented dialog systems demonstrate the effectiveness of our method. With only 10 labeled examples for each class in each task, the proposed method achieves 10% improvement over state-of-the-art methods demonstrating its effectiveness for limited training labels regime. Yaqing Wang 0001, Subhabrata Mukherjee, Haoda Chu, Yuancheng Tu, Jing Gao 0004, Ahmed Awadallah 0001 |
KDD | 7 |
| 2021 | Structure-Grounded Pretraining for Text-to-SQLabstractXiang Deng, Ahmed Hassan Awadallah, Christopher Meek, Oleksandr Polozov, Huan Sun, Matthew Richardson. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Xiang Deng 0001, Ahmed Awadallah 0001, Christopher Meek, Oleksandr Polozov, Huan Sun 0001, Matthew Richardson |
NAACL-HLT | 2 |
| 2021 | NL-EDIT: Correcting Semantic Parse Errors through Natural Language InteractionabstractAhmed Elgohary, Christopher Meek, Matthew Richardson, Adam Fourney, Gonzalo Ramos, Ahmed Hassan Awadallah. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Ahmed Elgohary, Christopher Meek, Matthew Richardson, Adam Fourney, Gonzalo A. Ramos, Ahmed Awadallah 0001 |
NAACL-HLT | 6 |
| 2021 | Self-Training with Weak SupervisionabstractGiannis Karamanolakis, Subhabrata Mukherjee, Guoqing Zheng, Ahmed Hassan Awadallah. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Giannis Karamanolakis, Subhabrata Mukherjee, Guoqing Zheng, Ahmed Awadallah 0001 |
NAACL-HLT | 4 |
| 2021 | MetaXL: Meta Representation Transformation for Low-resource Cross-lingual LearningabstractMengzhou Xia, Guoqing Zheng, Subhabrata Mukherjee, Milad Shokouhi, Graham Neubig, Ahmed Hassan Awadallah. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Mengzhou Xia, Guoqing Zheng, Subhabrata Mukherjee, Milad Shokouhi, Graham Neubig, Ahmed Awadallah 0001 |
NAACL-HLT | 6 |
| 2021 | QMSum: A New Benchmark for Query-based Multi-domain Meeting SummarizationabstractMing Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, Dragomir Radev. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Ming Zhong 0005, Da Yin, Tao Yu 0009, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Awadallah 0001, Asli Celikyilmaz, Yang Liu 0124, Xipeng Qiu, Dragomir R. Radev |
NAACL-HLT | 7 |
| 2021 | Fairness via Representation NeutralizationabstractExisting bias mitigation methods for DNN models primarily work on learning debiased encoders. This process not only requires a lot of instance-level annotations for sensitive attributes, it also does not guarantee that all fairness sensitive information has been removed from the encoder. To address these limitations, we explore the following research question: Can we reduce the discrimination of DNN models by only debiasing the classification head, even with biased representations as inputs? To this end, we propose a new mitigation technique, namely, Representation Neutralization for Fairness (RNF) that achieves fairness by debiasing only the task-specific classification head of DNN models. To this end, we leverage samples with the same ground-truth label but different sensitive attributes, and use their neutralized representations to train the classification head of the DNN model. The key idea of RNF is to discourage the classification head from capturing spurious correlation between fairness sensitive information in encoder representations with specific class labels. To address low-resource settings with no access to sensitive attribute annotations, we leverage a bias-amplified model to generate proxy annotations for sensitive attributes. Experimental results over several benchmark datasets demonstrate our RNF framework to effectively reduce discrimination of DNN models with minimal degradation in task-specific performance. Mengnan Du, Subhabrata Mukherjee, Guanchu Wang, Ruixiang Tang, Ahmed Awadallah 0001, Xia Ben Hu |
NeurIPS | 5 |
| 2020 | Speak to your Parser: Interactive Text-to-SQL with Natural Language FeedbackabstractWe study the task of semantic parse correction with natural language feedback.Given a natural language utterance, most semantic parsing systems pose the problem as one-shot translation where the utterance is mapped to a corresponding logical form.In this paper, we investigate a more interactive scenario where humans can further interact with the system by providing free-form natural language feedback to correct the system when it generates an inaccurate interpretation of an initial utterance.We focus on natural language to SQL systems and construct, SPLASH, a dataset of utterances, incorrect SQL interpretations and the corresponding natural language feedback.We compare various reference models for the correction task and show that incorporating such a rich form of feedback can significantly improve the overall semantic parsing accuracy while retaining the flexibility of natural language interaction.While we estimated human correction accuracy is 81.5%, our best model achieves only 25.1%, which leaves a large gap for improvement in future research.SPLASH is publicly available at https:// aka.ms/Splash_dataset. Exact Match Accuracy (%) BaselineCorrection End-to-End Without Feedback ⇒ Seq2Struct N/A 41.30 ⇒ Re- Ahmed Elgohary, Saghar Hosseini, Ahmed Awadallah 0001 |
ACL | 3 |
| 2020 | XtremeDistil: Multi-stage Distillation for Massive Multilingual ModelsabstractDeep and large pre-trained language models are the state-of-the-art for various natural language processing tasks.However, the huge size of these models could be a deterrent to using them in practice.Some recent works use knowledge distillation to compress these huge models into shallow ones.In this work we study knowledge distillation with a focus on multilingual Named Entity Recognition (NER).In particular, we study several distillation strategies and propose a stage-wise optimization scheme leveraging teacher internal representations, that is agnostic of teacher architecture, and show that it outperforms strategies employed in prior works.Additionally, we investigate the role of several factors like the amount of unlabeled data, annotation resources, model architecture and inference latency to name a few.We show that our approach leads to massive compression of teacher models like mBERT by upto 35x in terms of parameters and 51x in terms of latency for batch inference while retaining 95% of its F 1 -score for NER over 41 languages. Subhabrata Mukherjee, Ahmed Awadallah 0001 |
ACL | 2 |
| 2020 | Smart To-Do: Automatic Generation of To-Do Items from EmailsabstractIntelligent features in email service applications aim to increase productivity by helping people organize their folders, compose their emails and respond to pending tasks.In this work, we explore a new application, Smart-To-Do, that helps users with task management over emails.We introduce a new task and dataset for automatically generating To-Do items from emails where the sender has promised to perform an action.We design a two-stage process leveraging recent advances in neural text generation and sequenceto-sequence learning, obtaining BLEU and ROUGE scores of 0.23 and 0.63 for this task.To the best of our knowledge, this is the first work to address the problem of composing To-Do items from emails. Sudipto Mukherjee 0001, Subhabrata Mukherjee, Marcello Hasegawa, Ahmed Awadallah 0001, Ryen W. White |
ACL | 4 |
| 2020 | Gender Bias in Multilingual Embeddings and Cross-Lingual TransferabstractMultilingual representations embed words from many languages into a single semantic space such that words with similar meanings are close to each other regardless of the language.These embeddings have been widely used in various settings, such as cross-lingual transfer, where a natural language processing (NLP) model trained on one language is deployed to another language.While the crosslingual transfer techniques are powerful, they carry gender bias from the source to target languages.In this paper, we study gender bias in multilingual embeddings and how it affects transfer learning for NLP applications.We create a multilingual dataset for bias analysis and propose several ways for quantifying bias in multilingual representations from both the intrinsic and extrinsic perspectives.Experimental results show that the magnitude of bias in the multilingual representations changes differently when we align the embeddings to different target spaces and that the alignment direction can also have an influence on the bias in transfer learning.We further provide recommendations for using the multilingual word representations for downstream tasks. Jieyu Zhao 0001, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang 0001, Ahmed Awadallah 0001 |
ACL | 5 |
| 2020 | Analyzing Web Search Behavior for Software Engineering TasksabstractWeb search plays an integral role in software engineering (SE) to help with various tasks such as finding documentation, debugging, installation, etc. In this work, we present the first large-scale analysis of web search behavior for SE tasks using the search query logs from Bing, a commercial web search engine. First, we use distant supervision to build a machine learning classifier to extract the SE search queries with an F1 score of 93%. We then perform an analysis on one million search sessions to understand how software engineering related queries and sessions differ from other queries and sessions. Subsequently, we propose a taxonomy of intents to identify the various contexts in which web search is used in software engineering. Lastly, we analyze millions of SE queries to understand the distribution, search metrics and trends across these SE search intents. Our analysis shows that SE related queries form a significant portion of the overall web search traffic. Additionally, we found that there are six major intent categories for which web search is used in software engineering. The techniques and insights can not only help improve existing tools but can also inspire the development of new tools that aid in finding information for SE related tasks. Nikitha Rao, Chetan Bansal, Thomas Zimmermann 0001, Ahmed Awadallah 0001, Nachiappan Nagappan |
IEEE BigData | 4 |
| 2020 | Effects of Past Interactions on User Experience with Recommended DocumentsabstractRecommender systems are commonly used in entertainment, news, e-commerce, and social media. Document recommendation is a new and under-explored application area, in which both re-finding and discovery of documents need to be supported. In this paper we provide an initial exploration of users' experience with recommended documents, with a focus on how prior interactions influence recognition and interest. Through a field study of more than 100 users, we investigate the effects of past interactions with recommended documents on users' recognition of, prior intent to open, and interest in the documents. We examined different presentations of interaction history, and the recency and richness of prior interaction. We found that presentation only influenced recognition time. Our findings also indicate that people are more likely to recognize documents they had accessed recently and to do so more quickly. Similarly, documents that people had interacted with more deeply were also more frequently and quickly recognized. However, people were more interested in older documents or those with which they had less involved interactions. This finding suggests that in addition to helping users quickly access documents they intend to re-find, document recommendation can add value in helping users discover other documents. Our results offer implications for designing document recommendation systems that help users fulfil different needs. Farnaz Jahanbakhsh, Ahmed Awadallah 0001, Susan T. Dumais, Xuhai Xu |
CHIIR | 2 |
| 2020 | An Empirical Study of Software Exceptions in the Field using Search LogsabstractBackground: Software engineers spend a substantial amount of time using Web search to accomplish software engineering tasks. Such search tasks include finding code snippets, API documentation, seeking help with debugging, etc. While debugging a bug or crash, one of the common practices of software engineers is to search for information about the associated error or exception traces on the internet. Aims: In this paper, we analyze query logs from Bing to carry out a large scale study of software exceptions. To the best of our knowledge, this is the first large scale study to analyze how Web search is used to find information about exceptions. Method: We analyzed about 1 million exception related search queries from a random sample of 5 billion web search queries. To extract exceptions from unstructured query text, we built a novel machine learning model. With the model, we extracted exceptions from raw queries and performed popularity, effort, success, query characteristic and web domain analysis. We also performed programming language-specific analysis to give a better view of the exception search behavior. Results: Using the model with an F1-score of 0.82, our study identifies most frequent, most effort-intensive, or less successful exceptions and popularity of community Q&A sites. Conclusion: These techniques can help improve existing methods, documentation and tools for exception analysis and prediction. Further, similar techniques can be applied for APIs, frameworks, etc. Foyzul Hassan, Chetan Bansal, Nachiappan Nagappan, Thomas Zimmermann 0001, Ahmed Awadallah 0001 |
ESEM | 5 |
| 2020 | Uncertainty-aware Self-training for Few-shot Text ClassificationabstractRecent success of pre-trained language models crucially hinges on fine-tuning them on large amounts of labeled data for the downstream task, that are typically expensive to acquire or difficult to access for many applications. We study self-training as one of the earliest semi-supervised learning approaches to reduce the annotation bottleneck by making use of large-scale unlabeled data for the target task. Standard self-training mechanism randomly samples instances from the unlabeled pool to generate pseudo-labels and augment labeled data. We propose an approach to improve self-training by incorporating uncertainty estimates of the underlying neural network leveraging recent advances in Bayesian deep learning. Specifically, we propose (i) acquisition functions to select instances from the unlabeled pool leveraging Monte Carlo (MC) Dropout, and (ii) learning mechanism leveraging model confidence for self-training. As an application, we focus on text classification with five benchmark datasets. We show our methods leveraging only 20-30 labeled samples per class for each task for training and for validation perform within 3% of fully supervised pre-trained language models fine-tuned on thousands of labels with an aggregate accuracy of 91% and improvement of up to 12% over baselines. Subhabrata Mukherjee, Ahmed Awadallah 0001 |
NeurIPS | 2 |
| 2020 | Early Detection of Fake News with Multi-source Weak Social Supervision
Kai Shu, Guoqing Zheng, Yichuan Li 0001, Subhabrata Mukherjee, Ahmed Awadallah 0001, Scott W. Ruston, Huan Liu 0001 |
ECML/PKDD (3) | 5 |
| 2020 | Learning with Weak Supervision for Email Intent DetectionabstractEmail remains one of the most frequently used means of online communication. People spend significant amount of time every day on emails to exchange information, manage tasks and schedule events. Previous work has studied different ways for improving email productivity by prioritizing emails, suggesting automatic replies or identifying intents to recommend appropriate actions. The problem has been mostly posed as a supervised learning problem where models of different complexities were proposed to classify an email message into a predefined taxonomy of intents or classes. The need for labeled data has always been one of the largest bottlenecks in training supervised models. This is especially the case for many real-world tasks, such as email intent classification, where large scale annotated examples are either hard to acquire or unavailable due to privacy or data access constraints. Email users often take actions in response to intents expressed in an email (e.g., setting up a meeting in response to an email with a scheduling request). Such actions can be inferred from user interaction logs. In this paper, we propose to leverage user actions as a source of weak supervision, in addition to a limited set of annotated examples, to detect intents in emails. We develop an end-to-end robust deep neural network model for email intent identification that leverages both clean annotated data and noisy weak supervision along with a self-paced learning mechanism. Extensive experiments on three different intent detection tasks show that our approach can effectively leverage the weakly supervised data to improve intent detection in emails. Kai Shu, Subhabrata Mukherjee, Guoqing Zheng, Ahmed Awadallah 0001, Milad Shokouhi, Susan T. Dumais |
SIGIR | 4 |
| 2020 | Understanding User Behavior For Document RecommendationabstractPersonalized document recommendation systems aim to provide users with a quick shortcut to the documents they may want to access next, usually with an explanation about why the document is recommended. Previous work explored various methods for better recommendations and better explanations in different domains. However, there are few efforts that closely study how users react to the recommended items in a document recommendation scenario. We conducted a large-scale log study of users’ interaction behavior with the explainable recommendation on one of the largest cloud document platforms office.com. Our analysis reveals a number of factors, including display position, file type, authorship, recency of last access, and most importantly, the recommendation explanations, that are associated with whether users will recognize or open the recommended documents. Moreover, we specifically focus on explanations and conduct an online experiment to investigate the influence of different explanations on user behavior. Our analysis indicates that the recommendations help users access their documents significantly faster, but sometimes users miss a recommendation and resort to other more complicated methods to open the documents. Our results suggest opportunities to improve explanations and more generally the design of systems that provide and explain recommendations for documents. Xuhai Xu, Ahmed Awadallah 0001, Susan T. Dumais, Farheen Omar, Bogdan Popp, Robert Rounthwaite, Farnaz Jahanbakhsh |
WWW | 2 |
| 2020 | Special issue on learning from user interactions
Rishabh Mehrotra, Ahmed Awadallah 0001, Emine Yilmaz |
Inf. Retr. J. | 2 |
| 2019 | Adversarial Training for Community Question Answer Selection Based on Multi-Scale MatchingabstractCommunity-based question answering (CQA) websites represent an important source of information. As a result, the problem of matching the most valuable answers to their corresponding questions has become an increasingly popular research topic. We frame this task as a binary (relevant/irrelevant) classification problem, and present an adversarial training framework to alleviate label imbalance issue. We employ a generative model to iteratively sample a subset of challenging negative samples to fool our classification model. Both models are alternatively optimized using REINFORCE algorithm. The proposed method is completely different from previous ones, where negative samples in training set are directly used or uniformly down-sampled. Further, we propose using Multi-scale Matching which explicitly inspects the correlation between words and ngrams of different levels of granularity. We evaluate the proposed method on SemEval 2016 and SemEval 2017 datasets and achieves state-of-the-art or similar performance. Xiao Yang 0004, Madian Khabsa, Miaosen Wang, Wei Wang 0238, Ahmed Awadallah 0001, Daniel Kifer, C. Lee Giles |
AAAI | 5 |
| 2019 | Multi-Source Cross-Lingual Model Transfer: Learning What to ShareabstractModern NLP applications have enjoyed a great boost utilizing neural networks models.Such deep neural models, however, are not applicable to most human languages due to the lack of annotated training data for various NLP tasks.Cross-lingual transfer learning (CLTL) is a viable method for building NLP models for a low-resource target language by leveraging labeled data from other (source) languages.In this work, we focus on the multilingual transfer setting where training data in multiple source languages is leveraged to further boost target language performance.Unlike most existing methods that rely only on language-invariant features for CLTL, our approach coherently utilizes both languageinvariant and language-specific features at instance level.Our model leverages adversarial networks to learn language-invariant features, and mixture-of-experts models to dynamically exploit the similarity between the target language and each individual source language 1 .This enables our model to learn effectively what to share between various languages in the multilingual setup.Moreover, when coupled with unsupervised multilingual embeddings, our model can operate in a zero-resource setting where neither target language training data nor cross-lingual resources are available.Our model achieves significant performance gains over prior art, as shown in an extensive set of experiments over multiple text classification and sequence tagging tasks including a large-scale industry dataset. Xilun Chen 0002, Ahmed Awadallah 0001, Hany Hassan, Wei Wang 0238, Claire Cardie |
ACL (1) | 2 |
| 2019 | Exploring Email Triage: Challenges and OpportunitiesabstractDespite the centrality of email in the daily routines of knowledge workers, fundamental aspects of its usage are still poorly understood. We are particularly interested in understanding one aspect of email management, email triage, the process of going through unhandled email and deciding what to do with them. In this paper we investigate the email triage behavior by presenting interview and survey results that characterize user behavior and needs. The results highlight current challenges and enhance our understanding of how the triage process can be more effectively supported. Bahareh Sarrafzadeh, Ahmed Awadallah 0001, Milad Shokouhi |
CHIIR | 2 |
| 2019 | Learning About Work Tasks to Inform Intelligent Assistant DesignabstractIntelligent assistants can serve many purposes, including entertainment (e.g. playing music), home automation, and task management (e.g. timers, reminders). The role of these assistants is evolving to also support people engaged in work tasks, in workplaces and beyond. To design truly useful intelligent assistants for work, it is important to better understand the work tasks that people are performing. Based on a survey of 401 respondents' daily tasks and activities in a work setting, we present a classification of work-related tasks, and analyze their key characteristics, including the frequency of their self-reported tasks, the environment in which they undertake the tasks, and which, if any, electronic devices are used. We also investigate the cyber, physical, and social aspects of tasks. Finally, we reflect on how intelligent assistants could influence and help people in a work environment to complete their tasks, and synthesize our findings to provide insight on the future of intelligent assistants in support of amplifying personal productivity. Johanne R. Trippas, Damiano Spina, Falk Scholer, Ahmed Awadallah 0001, Peter Bailey, Paul N. Bennett, Ryen W. White, Jonathan Liono, Yongli Ren, Flora D. Salim, Mark Sanderson |
CHIIR | 4 |
| 2019 | Context-Aware Intent Identification in Email ConversationsabstractEmail continues to be one of the most important means of online communication. People spend a significant amount of time sending, reading, searching and responding to email in order to manage tasks, exchange information, etc. In this paper, we study intent identification in workplace email. We use a large scale publicly available email dataset to characterize intents in enterprise email and propose methods for improving intent identification in email conversations. Previous work focused on classifying email messages into broad topical categories or detecting sentences that contain action items or follow certain speech acts. In this work, we focus on sentence-level intent identification and study how incorporating more context (such as the full message body and other metadata) could improve the performance of the intent identification models. We experiment with several models for leveraging context including both classical machine learning and deep learning approaches. We show that modeling the interaction between sentence and context can significantly improve the performance. Wei Wang 0238, Saghar Hosseini, Ahmed Awadallah 0001, Paul N. Bennett, Chris Quirk |
SIGIR | 3 |
| 2019 | Task Completion Detection: A Study in the Context of Intelligent SystemsabstractPeople can record their pending tasks using to-do lists, digital assistants, and other task management software. In doing so, users of these systems face at least two challenges: (1) they must manually mark their tasks as complete, and (2) when systems proactively remind them about their pending tasks, say, via interruptive notifications, they lack information on task completion status. As a result, people may not realize the full benefits of to-do lists (since these lists can contain both completed and pending tasks) and they may be reminded about tasks they have already done (wasting time and causing frustration). In this paper, we present methods to automatically detect task completion. These inferences can be used to deprecate completed tasks and/or suppress notifications for these tasks (or for other purposes, e.g., task prioritization). Using log data from a popular digital assistant, we analyze temporal dynamics in the completion of tasks and train machine-learned models to detect completion with accuracy exceeding 80% using a variety of features (time elapsed since task creation, task content, email, notifications, user history). The findings have implications for the design of intelligent systems to help people manage their tasks. Ryen W. White, Ahmed Awadallah 0001, Robert Sim |
SIGIR | 2 |
| 2019 | Task Intelligence Workshop @ WSDM 2019abstractThe task intelligence workshop at the 2019 ACM Web Search and Data Mining (WSDM) conference comprised a mixture of research paper presentations, reports from data challenge participants, invited keynote(s) on broad topics related to tasks, and a workshop-wide discussion about task intelligence and its implications for system development. Ahmed Awadallah 0001, Cathal Gurrin, Mark Sanderson, Ryen W. White |
WSDM | 1 |
| 2019 | Characterizing and Predicting Email Deferral BehaviorabstractEmail triage involves going throughunhandled emails and deciding what to do with them. This familiar process can become increasingly challenging as the number of unhandled email grows. During a triage session, users commonlydefer handling emails that they cannot immediately deal with to later. These deferred emails, are often related to tasks that are postponed until the user has more time or the right information to deal with them. In this paper, through qualitative interviews and a large-scale log analysis, we study when and whatenterprise email users tend to defer. We found that users are more likely to defer emails when handling them involves replying, reading carefully, or clicking on links and attachments. We also learned that the decision to defer emails depends on many factors such as user's workload and the importance of the sender. Our qualitative results suggested that deferring is very common, and our quantitative log analysis confirms that 12% of triage sessions and 16% of daily active users had at least one deferred email on weekdays. We also discuss severaldeferral strategies such as marking emails as unread and flagging that are reported by our interviewees, and illustrate how such patterns can be also observed in user logs. Inspired by the characteristics of deferred emails and contextual factors involved in deciding if an email should be deferred, we train a classifier for predicting whether a recently triaged email is actually deferred. Our experimental results suggests that deferral can be classified with modest effectiveness. Overall, our work provides novel insights about how users handle their emails and how deferral can be modeled. Bahareh Sarrafzadeh, Ahmed Awadallah 0001, Christopher H. Lin, Milad Shokouhi, Susan T. Dumais |
WSDM | 2 |
| 2019 | Task Duration EstimationabstractEstimating how long a task will take to complete (i.e., the task duration) is important for many applications, including calendaring and project management. Population-scale calendar data contains distributional information about time allocated by individuals for tasks that may be useful to build computational models for task duration estimation. This study analyzes anonymized large-scale calendar appointment data from hundreds of thousands of individuals and millions of tasks to understand expected task durations and the longitudinal evolution in these durations. Machine-learned models are trained using the appointment data to estimate task duration. Study findings show that task attributes, including content (anonymized appointment subjects), context, and history, are correlated with time allocated for tasks. We also show that machine-learned models can be trained to estimate task duration, with multiclass classification accuracies of almost 80%. The findings have implications for understanding time estimation in populations, and in the design of support in digital assistants and calendaring applications to find time for tasks and to help people, especially those who are new to a task, block sufficient time for task completion. Ryen W. White, Ahmed Awadallah 0001 |
WSDM | 2 |
| 2019 | Exploring User Behavior in Email Re-Finding TasksabstractEmail continues to be one of the most commonly used forms of online communication. As inboxes grow larger, users rely more heavily on email search to effectively find what they are looking for. However, previous studies on email have been exclusive to enterprises with access to large user logs, or limited to small-scale qualitative surveys and analyses on limited public datasets such as Enron1 and Avocado2. In this work, we propose a novel framework that allows for experimentation with real email data. In particular, our approach provides a realistic way of simulating email re-finding tasks in a crowdsourcing environment using the workers' personal email data. We use our approach to experiment with various ranking functions and quality degradation to measure how users behave under different conditions, and conduct analysis across various email types and attributes. Our results show that user behavior can be significantly impacted as a result of the quality of the search ranker, but only when differences in quality are very pronounced. Our analysis confirms that time-based ranking begins to fail as email age increases, suggesting that hybrid approaches may help bridge the gap between relevance-based rankers and the traditional time-based ranking approach. Finally, we also found that users typically reformulate search queries by either entirely re-writing the query, or simply appending terms to the query, which may have implications for email query suggestion facilities. Joel Mackenzie, Kshitiz Gupta, Fang Qiao, Ahmed Awadallah 0001, Milad Shokouhi |
WWW | 4 |
| 2018 | The Lifetime of Email Messages: A Large-Scale Analysis of Email RevisitationabstractEmail continues to be one of the most important means of online communication, leading to a number of challenges related to information overload and email management. To better understand email management practices in detail, we examine the distribution of visits to emails over time. During their lifetime, emails may be visited one or more times, and with each visit different actions may be taken. Emails that are revisited over time are especially interesting because they represent an opportunity to improve email management and search. In this paper, we present a large-scale log analysis of email revisitation, the activities that people perform on revisited email messages (e.g. responding to, organizing or deleting messages, and opening attachments), and the strategies they use to go back to these emails. We find that most emails have a short lifetime, with more than 33% having a lifetime of less than 5 minutes. We also find that deleting is the most common action taken on messages visited once, and that responding and organizing are more common for messages visited more than once. We complement the log analysis with a survey to understand the motivation behind revisits and the types of emails that are revisited. The survey results show that 73% of the visits are to find information (e.g. a link or document, instructions to perform a task, or answers to questions), while 20% of revisits are to respond to the email. Our findings have implications for designing email clients and intelligent agents that support both short- and long-term revisitation patterns. Tarfah Alrashed, Ahmed Awadallah 0001, Susan T. Dumais |
CHIIR | 2 |
| 2018 | Natural Language Interfaces with Fine-Grained User Interaction: A Case Study on Web APIsabstractThe rapidly increasing ubiquity of computing puts a great demand on next-generation human-machine interfaces. Natural language interfaces, exemplified by virtual assistants like Apple Siri and Microsoft Cortana, are widely believed to be a promising direction. However, current natural language interfaces provide users with little help in case of incorrect interpretation of user commands. We hypothesize that the support of fine-grained user interaction can greatly improve the usability of natural language interfaces. In the specific setting of natural language interface to web APIs, we conduct a systematic study to verify our hypothesis. To facilitate this study, we propose a novel modular sequence-to-sequence model to create interactive natural language interfaces. By decomposing the complex prediction process of a typical sequence-to-sequence model into small, highly-specialized prediction units called modules, it becomes straightforward to explain the model prediction to the user, and solicit user feedback to correct possible prediction errors at a fine-grained level. We test our hypothesis by comparing an interactive natural language interface with its non-interactive version through both simulation and human subject experiments with real-world APIs. We show that with the interactive natural language interface, users can achieve a higher success rate and a lower task completion time, which lead to greatly improved user satisfaction. Yu Su 0001, Ahmed Awadallah 0001, Miaosen Wang, Ryen W. White |
SIGIR | 2 |
| 2018 | Characterizing and Supporting Question Answering in Human-to-Human CommunicationabstractEmail continues to be one of the most important means of online communication. People spend a significant amount of time sending, reading, searching and responding to email in order to manage tasks, exchange information, etc. In this paper, we focus on information exchange over enterprise email in the form of questions and answers. We study a large scale publicly available email dataset to characterize information exchange via questions and answers in enterprise email. We augment our analysis with a survey to gain insights on the types of questions exchanged, when and how do people get back to them and whether this behavior is adequately supported by existing email management and search functionality. We leverage this understanding to define the task of extracting question/answer pairs from threaded email conversations. We propose a neural network based approach that matches the question to the answer considering comparisons at different levels of granularity. We also show that we can improve the performance by leveraging external data of question and answer pairs. We test our approach using a manually labeled email data collected using a crowd-sourcing annotation study. Our findings have implications for designing email clients and intelligent agents that support question answering and information lookup in email. Xiao Yang 0004, Ahmed Awadallah 0001, Madian Khabsa, Wei Wang 0238, Miaosen Wang |
SIGIR | 2 |
| 2018 | LearnIR: WSDM 2018 Workshop on Learning from User InteractionsabstractWhile users interact with online services(e.g. search engines, recommender systems, conversational agents), they leave behind fine grained traces of interaction patterns. The ability to understand user behavior, record and interpret user interaction signals, gauge user satisfaction and incorporate user feedback gives online systems a vast treasure trove of insights for improvement and experimentation. More generally, the ability to learn from user interactions promises pathways for solving a number of problems and improving user engagement and satisfaction. Rishabh Mehrotra, Ahmed Awadallah 0001, Emine Yilmaz |
WSDM | 2 |
| 2017 | Self-Es: The Role of Emails-to-Self in Personal Information ManagementabstractEmail has been central to online communication for the past two decades. Through constant use, new information flows are being defined around users' interactions with emails. Alongside traditional messages, the email inbox is an always-available repository of to-do lists, reminders, files and notes. In this paper, we investigate the use of self-addressed emails (self-Es) as an information management tool, by analysing both: (i) responses to a survey about email use; and (ii) a collection of user donated self-addressed emails. Our results show that sending self-Es is a frequent behaviour among the users we questioned. In addition, we find that to-dos and reminders are the most popular type of information contained in emails-to-self. Our findings have direct implications for the development of systems that support novel interactions with information flows centred around email. Horatiu S. Bota, Paul N. Bennett, Ahmed Awadallah 0001, Susan T. Dumais |
CHIIR | 3 |
| 2017 | Beyond Success Rate: Utility as a Search Quality Metric for Online ExperimentsabstractUser satisfaction metrics are an integral part of search engine development as they help system developers to understand and evaluate the quality of the user experience. Research to date has mostly focused on predicting success or frustration as a proxy for satisfaction. However, users' search experience is more complex than merely being either successful or not. As such, using success rate as a measure of satisfaction can be limiting. In this work, we propose the use of utility as a measure of searcher satisfaction. This concept represents the fulfillment a user receives from con-suming a service and explains how users aim to gain optimal overall satisfaction. Our utility metrics measure the user satisfac-tion by aggregating all their interaction with the search engine. These interactions are represented as a timeline of actions and their dwelltimes, where each action is classified as having a posi-tive or negative effect on the user. We examine sessions mined from Bing logs, with multi-point scale assessment of searcher satisfaction and show that utility is a better proxy for satisfaction compared to success. Leveraging that data, we design metrics of searcher satisfaction that assess the overall utility accumulated by a user during her search session. We use real user traffic from millions of users in an A/B setting to compare utility metrics to success rate metrics. We show that utility is a better metric for evaluating searcher satisfaction with the search engine, and a more sensitive and accurate metric when compared to predicting success. These metrics are currently adopted as the top-level met-ric for evaluating the thousands of A/B experiments that are run on Bing each year. Widad Machmouchi, Ahmed Awadallah 0001, Imed Zitouni, Georg Buscher |
CIKM | 2 |
| 2017 | Deep Sequential Models for Task Satisfaction PredictionabstractDetecting and understanding implicit signals of user satisfaction are essential for experimentation aimed at predicting searcher satisfaction. As retrieval systems have advanced, search tasks have steadily emerged as accurate units not only to capture searcher's goals but also in understanding how well a system is able to help the user achieve that goal. However, a major portion of existing work on modeling searcher satisfaction has focused on query level satisfaction. The few existing approaches for task satisfaction prediction have narrowly focused on simple tasks aimed at solving atomic information needs. Rishabh Mehrotra, Ahmed Awadallah 0001, Milad Shokouhi, Emine Yilmaz, Imed Zitouni, Ahmed El Kholy, Madian Khabsa |
CIKM | 2 |
| 2017 | Building Natural Language Interfaces to Web APIsabstractAs the Web evolves towards a service-oriented architecture, application program interfaces (APIs) are becoming an increasingly important way to provide access to data, services, and devices. We study the problem of natural language interface to APIs (NL2APIs), with a focus on web APIs for web services. Such NL2APIs have many potential benefits, for example, facilitating the integration of web services into virtual assistants. Yu Su 0001, Ahmed Awadallah 0001, Madian Khabsa, Patrick Pantel, Michael Gamon, Mark J. Encarnación |
CIKM | 2 |
| 2017 | Characterizing and Predicting Enterprise Email Reply BehaviorabstractEmail is still among the most popular online activities. People spend a significant amount of time sending, reading and responding to email in order to communicate with others, manage tasks and archive personal information. Most previous research on email is based on either relatively small data samples from user surveys and interviews, or on consumer email accounts such as those from Yahoo! Mail or Gmail. Much less has been published on how people interact with enterprise email even though it contains less automatically generated commercial email and involves more organizational behavior than is evident in personal accounts. In this paper, we extend previous work on predicting email reply behavior by looking at enterprise settings and considering more than dyadic communications. We characterize the influence of various factors such as email content and metadata, historical interaction features and temporal features on email reply behavior. We also develop models to predict whether a recipient will reply to an email and how long it will take to do so. Experiments with the publicly-available Avocado email collection show that our methods outperform all baselines with large gains. We also analyze the importance of different features on reply behavior predictions. Our findings provide new insights about how people interact with enterprise email and have implications for the design of the next generation of email clients. Liu Yang 0005, Susan T. Dumais, Paul N. Bennett, Ahmed Awadallah 0001 |
SIGIR | 4 |
| 2017 | User Interaction Sequences for Search Satisfaction PredictionabstractDetecting and understanding implicit measures of user satisfaction are essential for meaningful experimentation aimed at enhancing web search quality. While most existing studies on satisfaction prediction rely on users' click activity and query reformulation behavior, often such signals are not available for all search sessions and as a result, not useful in predicting satisfaction. On the other hand, user interaction data (such as mouse cursor movement) is far richer than just click data and can provide useful signals for predicting user satisfaction. In this work, we focus on considering holistic view of user interaction with the search engine result page (SERP) and construct detailed universal interaction sequences of their activity. We propose novel ways of leveraging the universal interaction sequences to automatically extract informative, interpretable subsequences. In addition to extracting frequent, discriminatory and interleaved subsequences, we propose a Hawkes process model to incorporate temporal aspects of user interaction. Through extensive experimentation we show that encoding the extracted subsequences as features enables us to achieve significant improvements in predicting user satisfaction. We additionally present an analysis of the correlation between various subsequences and user satisfaction. Finally, we demonstrate the usefulness of the proposed approach in covering abandonment cases. Our findings provide a valuable tool for fine-grained analysis of user interaction behavior for metric development. Rishabh Mehrotra, Imed Zitouni, Ahmed Awadallah 0001, Ahmed El Kholy, Madian Khabsa |
SIGIR | 3 |
| 2016 | Understanding User Satisfaction with Intelligent AssistantsabstractVoice-controlled intelligent personal assistants, such as Cortana, Google Now, Siri and Alexa, are increasingly becoming a part of users' daily lives, especially on mobile devices. They introduce a significant change in information access, not only by introducing voice control and touch gestures but also by enabling dialogues where the context is preserved. This raises the need for evaluation of their effectiveness in assisting users with their tasks. However, in order to understand which type of user interactions reflect different degrees of user satisfaction we need explicit judgements. In this paper, we describe a user study that was designed to measure user satisfaction over a range of typical scenarios of use: controlling a device, web search, and structured search dialogue. Using this data, we study how user satisfaction varied with different usage scenarios and what signals can be used for modeling satisfaction in the different scenarios. We find that the notion of satisfaction varies across different scenarios, and show that, in some scenarios (e.g. making a phone call), task completion is very important while for others (e.g. planning a night out), the amount of effort spent is key. We also study how the nature and complexity of the task at hand affects user satisfaction, and find that preserving the conversation context is essential and that overall task-level satisfaction cannot be reduced to query-level satisfaction alone. Finally, we shed light on the relative effectiveness and usefulness of voice-controlled intelligent agents, explaining their increasing popularity and uptake relative to the traditional query-response interaction. Julia Kiseleva, Kyle Williams 0001, Jiepu Jiang, Ahmed Awadallah 0001, Aidan C. Crook, Imed Zitouni, Tasos Anastasakos |
CHIIR | 4 |
| 2016 | Learning to Account for Good Abandonment in Search Success MetricsabstractAbandonment in web search has been widely used as a proxy to measure user satisfaction. Initially it was considered a signal of dissatisfaction, however with search engines moving towards providing answer-like results, a new category of abandonment was introduced and referred to as Good Abandonment. Predicting good abandonment is a hard problem and it was the subject of several previous studies. All those studies have focused, though, on predicting good abandonment in offline settings using manually labeled data. Thus, it remained a challenge how to have an online metric that accounts for good abandonment. In this work we describe how a search success metric can be augmented to account for good abandonment sessions using a machine learned metric that depends on user's viewport information. We use real user traffic from millions of users to evaluate the proposed metric in an A/B experiment. We show that taking good abandonment into consideration has a significant effect on the overall performance of the online metric. Madian Khabsa, Aidan C. Crook, Ahmed Awadallah 0001, Imed Zitouni, Tasos Anastasakos, Kyle Williams 0001 |
CIKM | 3 |
| 2016 | Activity Modeling in EmailabstractAshequl Qadir, Michael Gamon, Patrick Pantel, Ahmed Hassan Awadallah. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Ashequl Qadir, Michael Gamon, Patrick Pantel, Ahmed Awadallah 0001 |
HLT-NAACL | 4 |
| 2016 | Predicting User Satisfaction with Intelligent AssistantsabstractThere is a rapid growth in the use of voice-controlled intelligent personal assistants on mobile devices, such as Microsoft's Cortana, Google Now, and Apple's Siri. They significantly change the way users interact with search systems, not only because of the voice control use and touch gestures, but also due to the dialogue-style nature of the interactions and their ability to preserve context across different queries. Predicting success and failure of such search dialogues is a new problem, and an important one for evaluating and further improving intelligent assistants. While clicks in web search have been extensively used to infer user satisfaction, their significance in search dialogues is lower due to the partial replacement of clicks with voice control, direct and voice answers, and touch gestures. Julia Kiseleva, Kyle Williams 0001, Ahmed Awadallah 0001, Aidan C. Crook, Imed Zitouni, Tasos Anastasakos |
SIGIR | 3 |
| 2016 | Is This Your Final Answer?: Evaluating the Effect of Answers on Good Abandonment in Mobile SearchabstractAnswers on mobile search result pages have become a common way to attempt to satisfy users without them needing to click on search results. Many different types of answers exist, such as weather, flight and currency answers. Understanding the effect that these different answer types have on mobile user behavior and how they contribute to satisfaction is important for search engine evaluation. We study these two aspects by analyzing the logs of a commercial search engine and through a user study. Our results show that user click, abandonment and engagement behavior differs depending on the answer types present on a page. Furthermore, we find that satisfaction rates differ in the presence of different answer types with simple answer types, such as time zone answers, leading to more satisfaction than more complex answers, such as news answers. Our findings have implications for the study and application of user satisfaction for search systems. Kyle Williams 0001, Julia Kiseleva, Aidan C. Crook, Imed Zitouni, Ahmed Awadallah 0001, Madian Khabsa |
SIGIR | 5 |
| 2016 | Detecting Good Abandonment in Mobile SearchabstractWeb search queries for which there are no clicks are referred to as abandoned queries and are usually considered as leading to user dissatisfaction. However, there are many cases where a user may not click on any search result page (SERP) but still be satisfied. This scenario is referred to as good abandonment and presents a challenge for most approaches measuring search satisfaction, which are usually based on clicks and dwell time. The problem is exacerbated further on mobile devices where search providers try to increase the likelihood of users being satisfied directly by the SERP. This paper proposes a solution to this problem using gesture interactions, such as reading times and touch actions, as signals for differentiating between good and bad abandonment. These signals go beyond clicks and characterize user behavior in cases where clicks are not needed to achieve satisfaction. We study different good abandonment scenarios and investigate the different elements on a SERP that may lead to good abandonment. We also present an analysis of the correlation between user gesture features and satisfaction. Finally, we use this analysis to build models to automatically identify good abandonment in mobile search achieving an accuracy of 75%, which is significantly better than considering query and session signals alone. Our findings have implications for the study and application of user satisfaction in search systems. Kyle Williams 0001, Julia Kiseleva, Aidan C. Crook, Imed Zitouni, Ahmed Awadallah 0001, Madian Khabsa |
WWW | 5 |
| 2015 | Characterizing and Predicting Voice Query ReformulationabstractVoice interactions are becoming more prevalent as the usage of voice search and intelligent assistants gains more popularity. Users frequently reformulate their requests in hope of getting better results either because the system was unable to recognize what they said or because it was able to recognize it but was unable to return the desired response. Query reformulation has been extensively studied in the context of text input. Many of the characteristics studied in the context of text query reformulation are potentially useful for voice query reformulation. However, voice query reformulation has its unique characteristics in terms of the reasons that lead users to reformulating their queries and how they reformulate them. In this paper, we study the problem of voice query reformulation. We perform a large scale human annotation study to collect thousands of labeled instances of voice reformulation and non-reformulation query pairs. We use this data to compare and contrast characteristics of reformulation and non-reformulation queries over a large a number of dimensions. We then train classifiers to distinguish between reformulation and non-reformulation query pairs and to predict the rationale behind reformulation. We demonstrate through experiments with the human labeled data that our classifiers achieve good performance in both tasks. Ahmed Awadallah 0001, Ranjitha Gurunath Kulkarni, Umut Ozertem, Rosie Jones |
CIKM | 1 |
| 2015 | Struggling and Success in Web SearchabstractWeb searchers sometimes struggle to find relevant information. Struggling leads to frustrating and dissatisfying search experiences, even if searchers ultimately meet their search objectives. Better understanding of search tasks where people struggle is important in improving search systems. We address this important issue using a mixed methods study using large-scale logs, crowd-sourced labeling, and predictive modeling. We analyze anonymized search logs from the Microsoft Bing Web search engine to characterize aspects of struggling searches and better explain the relationship between struggling and search success. To broaden our understanding of the struggling process beyond the behavioral signals in log data, we develop and utilize a crowd-sourced labeling methodology. We collect third-party judgments about why searchers appear to struggle and, if appropriate, where in the search task it became clear to the judges that searches would succeed (i.e., the pivotal query). We use our findings to propose ways in which systems can help searchers reduce struggling. Key components of such support are algorithms that accurately predict the nature of future actions and their anticipated impact on search outcomes. Our findings have implications for the design of search systems that help searchers struggle less and succeed more. Daan Odijk, Ryen W. White, Ahmed Awadallah 0001, Susan T. Dumais |
CIKM | 3 |
| 2015 | Personalizing Search on Shared DevicesabstractSearch personalization tailors the search experience to individual searchers. To do this, search engines construct interest models comprising signals from observed behavior associated with ma-chines, often via Web browser cookies or other user identifiers. However, shared device usage is common, meaning that the activities of multiple searchers may be interwoven in the interest models generated. Recent research on activity attribution has led to methods to automatically disentangle the histories of multiple searchers and correctly ascribe newly-observed search activity to the correct per-son. Building on this, we introduce attribution-based personalization (ABP), a procedure that extends traditional personalization to target individual searchers on shared devices. Activity attribution may improve personalization, but its benefits are not yet fully understood. We present an oracle study (with perfect knowledge of which searchers perform each action on each machine) to under-stand the effectiveness of ABP in predicting searchers' future interests. We utilize a large Web search log dataset containing both per-son identifiers and machine identifiers to quantify the gain in personalization performance from ABP, identify the circumstances under which ABP is most effective, and develop a classifier to determine when to apply it that yields sizable gains in personalization performance. ABP allows search providers to personalize experiences for individuals rather than targeting all users of a device collectively. Ryen W. White, Ahmed Awadallah 0001 |
SIGIR | 2 |
| 2015 | Understanding and Predicting Graded Search SatisfactionabstractUnderstanding and estimating satisfaction with search engines is an important aspect of evaluating retrieval performance. Research to date has modeled and predicted search satisfaction on a binary scale, i.e., the searchers are either satisfied or dissatisfied with their search outcome. However, users' search experience is a complex construct and there are different degrees of satisfaction. As such, binary classification of satisfaction may be limiting. To the best of our knowledge, we are the first to study the problem of understanding and predicting graded (multi-level) search satisfaction. We ex-amine sessions mined from search engine logs, where searcher satisfaction was also assessed on multi-point scale by human annotators. Leveraging these search log data, we observe rich and non-monotonous changes in search behavior in sessions with different degrees of satisfaction. The findings suggest that we should predict finer-grained satisfaction levels. To address this issue, we model search satisfaction using features indicating search outcome, search effort, and changes in both outcome and effort during a session. We show that our approach can predict subtle changes in search satisfaction more accurately than state-of-the-art methods, affording greater insight into search satisfaction. The strong performance of our models has implications for search providers seeking to accu-rately measure satisfaction with their services. Jiepu Jiang, Ahmed Awadallah 0001, Ryen W. White |
WSDM | 2 |
| 2015 | Automatic Online Evaluation of Intelligent AssistantsabstractVoice-activated intelligent assistants, such as Siri, Google Now, and Cortana, are prevalent on mobile devices. However, it is challenging to evaluate them due to the varied and evolving number of tasks supported, e.g., voice command, web search, and chat. Since each task may have its own procedure and a unique form of correct answers, it is expensive to evaluate each task individually. This paper is the first attempt to solve this challenge. We develop consistent and automatic approaches that can evaluate different tasks in voice-activated intelligent assistants. We use implicit feedback from users to predict whether users are satisfied with the intelligent assistant as well as its components, i.e., speech recognition and intent classification. Using this approach, we can potentially evaluate and compare different tasks within and across intelligent assistants ac-cording to the predicted user satisfaction rates. Our approach is characterized by an automatic scheme of categorizing user-system interaction into task-independent dialog actions, e.g., the user is commanding, selecting, or confirming an action. We use the action sequence in a session to predict user satisfaction and the quality of speech recognition and intent classification. We also incorporate other features to further improve our approach, including features derived from previous work on web search satisfaction prediction, and those utilizing acoustic characteristics of voice requests. We evaluate our approach using data collected from a user study. Results show our approach can accurately identify satisfactory and unsatisfactory sessions. Jiepu Jiang, Ahmed Awadallah 0001, Rosie Jones, Umut Ozertem, Imed Zitouni, Ranjitha Gurunath Kulkarni, Omar Zia Khan |
WWW | 2 |
| 2014 | Supporting Complex Search TasksabstractWe present methods to automatically identify and recommend sub-tasks to help people explore and accomplish complex search tasks. Although Web searchers often exhibit directed search behaviors such as navigating to a particular Website or locating a particular item of information, many search scenarios involve more complex tasks such as learning about a new topic or planning a vacation. These tasks often involve multiple search queries and can span multiple sessions. Current search systems do not provide adequate support for tackling these tasks. Instead, they place most of the burden on the searcher for discovering which aspects of the task they should explore. Particularly challenging is the case when a searcher lacks the task knowledge necessary to decide which step to tackle next. In this paper, we propose methods to automatically mine search logs for tasks and build an association graph connecting multiple tasks together. We then leverage the task graph to assist new searchers in exploring new search topics or tackling multi-step search tasks. We demonstrate through experiments with human participants that we can discover related and interesting tasks to assist with complex search scenarios. Ahmed Awadallah 0001, Ryen W. White, Patrick Pantel, Susan T. Dumais, Yi-Min Wang |
CIKM | 1 |
| 2014 | Machine-Assisted Search Preference EvaluationabstractInformation Retrieval systems are traditionally evaluated using the relevance of web pages to individual queries. Other work on IR evaluation has focused on exploring the use of preference judgments over two search result lists. Unlike traditional query-document evaluation, collecting preference judgments over two search result-lists takes the context of documents, and hence takes the interaction between search results, into consideration. Moreover, preference judgments have been shown to produce more accurate results compared to absolute judgment. On the other hand result list preference judgments have very high annotation cost. In this work, we investigate how machine learned models can assist human judges in order to collect reliable result list preference judgments at large scale with lower judgment-cost. We build novel models that can predict user preference automatically. We investigate the effect of different features on the prediction quality. We focus on predicting preferences with high confidence and show that these models can be effectively used to assist human judges resulting in significant reduction in annotation cost. Ahmed Awadallah 0001, Imed Zitouni |
CIKM | 1 |
| 2014 | Comparing client and server dwell time estimates for click-level satisfaction predictionabstractClick dwell time is the amount of time that a user spends on a clicked search result. Many previous studies have shown that click dwell time is strongly correlated with result-level satisfaction and document relevance. Accurate estimates of dwell time are therefore important for applications such as search satisfaction prediction and result ranking. However, dwell time can be estimated in different ways according to the information available about the search process. For example, a result reached for the query [Garfield] may involve 145s of "server-side" dwell time (observable to the search engine) and 40s of "client-side" dwell time (observable from the browser). Since search engines can only observe server-side actions (i.e., activity on the search engine result page), server-side dwell times are estimated by measuring the time between a search result click and the next search event (click or query). Conversely, more detailed information about page dwell times can be obtained via client-side methods such as Web browser toolbars. The client-side information enables the estimation of more accurate dwell times by measuring the amount of time that a user spends on pages of interest (either the landing page, or pages on the full navigation trail). In this paper, we define three different dwell times, i.e., server-side, client-side, and trail dwell time, and examine their effectiveness for predicting click satisfaction. For this, we collect toolbar and search engine logs from real users, and provide an analysis of dwell times for improving prediction performance. Moreover, we show further improvements in predicting click-level satisfaction by combining dwell times with other query features (e.g., query clarity). Ahmed Awadallah 0001, Ryen W. White, Imed Zitouni |
SIGIR | 2 |
| 2014 | Enhancing personalization via search activity attributionabstractOnline services rely on machine identifiers to tailor services such as personalized search and advertising to individual users. The assumption made is that each identifier comprises the behavior of a single person. However, shared machine usage is common, and in these cases, the activities of multiple users may be generated under a single identifier, creating a potentially noisy signal for applications such as search personalization. We propose enhancing Web search personalization with methods that can disambiguate among different users of a machine, thus connecting the current query with the appropriate search history. Using logs containing both person and machine identifiers, and logs from a popular commercial search engine, we learn models that accurately assign observed search behaviors to each of different users. This information is then used to augment existing personalization methods that are currently based only on machine identifiers. We show that this new capability to infer users can be used to improve the performance of existing personalization methods. The early findings of our research are promising and have implications for search personalization. Adish Singla, Ryen W. White, Ahmed Awadallah 0001, Eric Horvitz |
SIGIR | 3 |
| 2014 | Context-aware web search abandonment predictionabstractWeb search queries without hyperlink clicks are often referred to as abandoned queries. Understanding the reasons for abandonment is crucial for search engines in evaluating their performance. Abandonment can be categorized as good or bad depending on whether user information needs are satisfied by result page content. Previous research has sought to understand abandonment rationales via user surveys, or has developed models to predict those rationales using behavioral patterns. However, these models ignore important contextual factors such as the relationship between the abandoned query and prior abandonment instances. We propose more advanced methods for modeling and predicting abandonment rationales using contextual information from user search sessions by analyzing search engine logs, and discover dependencies between abandoned queries and user behaviors. We leverage these dependency signals to build a sequential classifier using a structured learning framework designed to handle such signals. Our experimental results show that our approach is 22% more accurate than the state-of-the-art abandonment-rationale classifier. Going beyond prediction, we leverage the prediction results to significantly improve relevance using instances of predicted good and bad abandonment. Yang Song 0008, Ryen W. White, Ahmed Awadallah 0001 |
SIGIR | 4 |
| 2014 | Modeling action-level satisfaction for search task satisfaction predictionabstractSearch satisfaction is a property of a user's search process. Understanding it is critical for search providers to evaluate the performance and improve the effectiveness of search engines. Existing methods model search satisfaction holistically at the search-task level, ignoring important dependencies between action-level satisfaction and overall task satisfaction. We hypothesize that searchers' latent action-level satisfaction (i.e., whether they believe they were satisfied with the results of a query or click) influences their observed search behaviors and contributes to overall search satisfaction. We conjecture that by modeling search satisfaction at the action level, we can build more complete and more accurate predictors of search-task satisfaction. To do this, we develop a latent structural learning method, whereby rich structured features and dependency relations unique to search satisfaction prediction are explored. Using in-situ search satisfaction judgments provided by searchers, we show that there is significant value in modeling action-level satisfaction in search-task satisfaction prediction. In addition, experimental results on large-scale logs from Bing.com demonstrate clear benefit from using inferred action satisfaction labels for other applications such as document relevance estimation and query suggestion. Hongning Wang, Yang Song 0008, Ming-Wei Chang, Xiaodong He 0001, Ahmed Awadallah 0001, Ryen W. White |
SIGIR | 5 |
| 2014 | Struggling or exploring?: disambiguating long search sessionsabstractWeb searchers often exhibit directed search behaviors such as navigating to a particular Website. However, in many circumstances they exhibit different behaviors that involve issuing many queries and visiting many results. In such cases, it is not clear whether the user's rationale is to intentionally explore the results or whether they are struggling to find the information they seek. Being able to disambiguate between these types of long search sessions is important for search engines both in performing retrospective analysis to understand search success, and in developing real-time support to assist searchers. The difficulty of this challenge is amplified since many of the characteristics of exploration (e.g., multiple queries, long duration) are also observed in sessions where people are struggling. In this paper, we analyze struggling and exploring behavior in Web search using log data from a commercial search engine. We first compare and contrast search behaviors along a number dimensions, including query dynamics during the session. We then build classifiers that can accurately distinguish between exploring and struggling sessions using behavioral and topical features. Finally, we show that by considering the struggling/exploring prediction we can more accurately predict search satisfaction. Ahmed Awadallah 0001, Ryen W. White, Susan T. Dumais, Yi-Min Wang |
WSDM | 1 |
| 2014 | Modeling dwell time to predict click-level satisfactionabstractClicks on search results are the most widely used behavioral signals for predicting search satisfaction. Even though clicks are correlated with satisfaction, they can also be noisy. Previous work has shown that clicks are affected by position bias, caption bias, and other factors. A popular heuristic for reducing this noise is to only consider clicks with long dwell time, usually equaling or exceeding 30 seconds. The rationale is that the more time a searcher spends on a page, the more likely they are to be satisfied with its contents. However, having a single threshold value assumes that users need a fixed amount of time to be satisfied with any result click, irrespective of the page chosen. In reality, clicked pages can differ significantly. Pages have different topics, readability levels, content lengths, etc. All of these factors may affect the amount of time spent by the user on the page. In this paper, we study the effect of different page characteristics on the time needed to achieve search satisfaction. We show that the topic of the page, its length and its readability level are critical in determining the amount of dwell time needed to predict whether any click is associated with satisfaction. We propose a method to model and provide a better understanding of click dwell time. We estimate click dwell time distributions for SAT (satisfied) or DSAT (dissatisfied) clicks for different click segments and use them to derive features to train a click-level satisfaction model. We compare the proposed model to baseline methods that use dwell time and other search performance predictors as features, and demonstrate that the proposed model achieves significant improvements. Ahmed Awadallah 0001, Ryen W. White, Imed Zitouni |
WSDM | 2 |
| 2014 | Contextual and dimensional relevance judgments for reusable SERP-level evaluationabstractDocument-level relevance judgments are a major component in the calculation of effectiveness metrics. Collecting high-quality judgments is therefore a critical step in information retrieval evaluation. However, the nature of and the assumptions underlying relevance judgment collection have not received much attention. In particular, relevance judgments are typically collected for each document in isolation, although users read each document in the context of other documents. In this work, we aim to investigate the nature of relevance judgment collection. We collect relevance labels in both isolated and conditional setting, and ask for judgments in various dimensions of relevance as well as overall relevance. Then we compare the relevance metrics based on various types of judgments with other metrics of quality such as user preference. Our analyses illuminate how these settings for judgment collection affect the quality and the characteristics of the judgments. We also find that the metrics based on conditional judgments show higher correlation with user preference than isolated judgments. Peter B. Golbus, Imed Zitouni, Jin Young Kim 0005, Ahmed Awadallah 0001, Fernando Diaz 0001 |
WWW | 4 |
| 2014 | From devices to people: attribution of search activity in multi-user settingsabstractOnline services rely on unique identifiers of machines to tailor offerings to their users. An implicit assumption is made that each machine identifier maps to an individual. However, shared ma-chines are common, leading to interwoven search histories and noisy signals for applications such as personalized search and ad-vertising. We present methods for attributing search activity to individual searchers. Using ground truth data for a sample of almost four million U.S. Web searchers-containing both machine identifiers and person identifiers-we show that over half of the machine identifiers comprise the queries of multiple people. We characterize variations in features of topic, time, and other aspects such as the complexity of the information sought per the number of searchers on a machine, and show significant differences in all measures. Based on these insights, we develop models to accurately estimate when multiple people contribute to the logs ascribed to a single machine identifier. We also develop models to cluster search behavior on a machine, allowing us to attribute historical data accurately and automatically assign new search activity to the correct searcher. The findings have implications for the design of applications such as personalized search and advertising that rely heavily on machine identifiers to custom-tailor their services. Ryen W. White, Ahmed Awadallah 0001, Adish Singla, Eric Horvitz |
WWW | 2 |
| 2014 | A Random Walk-Based Model for Identifying Semantic OrientationabstractAutomatically identifying the sentiment polarity of words is a very important task that has been used as the essential building block of many natural language processing systems such as text classification, text filtering, product review analysis, survey response analysis, and on-line discussion mining. We propose a method for identifying the sentiment polarity of words that applies a Markov random walk model to a large word relatedness graph, and produces a polarity estimate for any given word. The model can accurately and quickly assign a polarity sign and magnitude to any word. It can be used both in a semi-supervised setting where a training set of labeled words is used, and in a weakly supervised setting where only a handful of seed words is used to define the two polarity classes. The method is experimentally tested using a gold standard set of positive and negative words from the General Inquirer lexicon. We also show how our method can be used for three-way classification which identifies neutral words in addition to positive and negative words. Our experiments show that the proposed method outperforms the state-of-the-art methods in the semi-supervised setting and is comparable to the best reported values in the weakly supervised setting. In addition, the proposed method is faster and does not need a large corpus. We also present extensions of our methods for identifying the polarity of foreign words and out-of-vocabulary words. Ahmed Awadallah 0001, Amjad Abu-Jbara, Wanchen Lu, Dragomir R. Radev |
Comput. Linguistics | 1 |
| 2014 | Content Bias in Online Health SearchabstractSearch engines help people answer consequential questions. Biases in retrieved and indexed content (e.g., skew toward erroneous outcomes that represent deviations from reality), coupled with searchers' biases in how they examine and interpret search results, can lead people to incorrect answers. In this article, we seek to better understand biases in search and retrieval, and in particular those affecting the accuracy of content in search results, including the search engine index, features used for ranking, and the formulation of search queries. Focusing on the important domain of online health search, this research broadens previous work on biases in search to examine the role of search systems in contributing to biases. To assess bias, we focus on questions about medical interventions and employ reliable ground truth data from authoritative medical sources. In the course of our study, we utilize large-scale log analysis using data from a popular Web search engine, deep probes of result lists on that search engine, and crowdsourced human judgments of search result captions and landing pages. Our findings reveal bias in results, amplifying searchers' existing biases that appear evident in their search activity. We also highlight significant bias in indexed content and show that specific ranking signals and specific query terms support bias. Both of these can degrade result accuracy and increase skewness in search results. Our analysis has implications for bias mitigation strategies in online search systems, and we offer recommendations for search providers based on our findings. Ryen W. White, Ahmed Awadallah 0001 |
ACM Trans. Web | 2 |
| 2013 | Beyond clicks: query reformulation as a predictor of search satisfactionabstractTo understand whether a user is satisfied with the current search results, implicit behavior is a useful data source, with clicks being the best-known implicit signal. However, it is possible for a non-clicking user to be satisfied and a clicking user to be dissatisfied. Here we study additional implicit signals based on the relationship between the user's current query and the next query, such as their textual similarity and the inter-query time. Using a large unlabeled dataset, a labeled dataset of queries and a labeled dataset of user tasks, we analyze the relationship between these signals. We identify an easily-implemented rule that indicates dissatisfaction: that a similar query issued within a time interval that is short enough (such as five minutes) implies dissatisfaction. By incorporating additional query-based features in the model, we show that a query-based model (with no click information) can indicate satisfaction more accurately than click-based models. The best model uses both query and click features. In addition, by comparing query sequences in successful tasks and unsuccessful tasks, we observe that search success is an incremental process for successful tasks with multiple queries. Ahmed Awadallah 0001, Nick Craswell, Bill Ramsey |
CIKM | 1 |
| 2013 | Personalized models of search satisfactionabstractSearch engines need to model user satisfaction to improve their services. Since it is not practical to request feedback on searchers' perceptions and search outcomes directly from users, search engines must estimate satisfaction from behavioral signals such as query refinement, result clicks, and dwell times. This analysis of behavior in the aggregate leads to the development of global metrics such as satisfied result clickthrough (typically operationalized as result-page clicks with dwell time exceeding a particular threshold) that are then applied to all searchers' behavior to estimate satisfac-tion levels. However, satisfaction is a personal belief and how users behave when they are satisfied can also differ. In this paper we verify that searcher behavior when satisfied and dissatisfied is indeed different among individual searchers along a number of dimensions. As a result, we introduce and evaluate learned models of satisfaction for individual searchers and searcher cohorts. Through experimentation via logs from a large commercial Web search engine, we show that our proposed models can predict search satisfaction more accurately than a global baseline that applies the same satisfaction model across all users. Our findings have implications for the study and application of user satisfaction in search systems. Ahmed Awadallah 0001, Ryen W. White |
CIKM | 1 |
| 2013 | Toward self-correcting search engines: using underperforming queries to improve searchabstractSearch engines receive queries with a broad range of different search intents. However, they do not perform equally well for all queries. Understanding where search engines perform poorly is critical for improving their performance. In this paper, we present a method for automatically identifying poorly-performing query groups where a search engine may not meet searcher needs. This allows us to create coherent query clusters that help system design-ers generate actionable insights about necessary changes and helps learning-to-rank algorithms better learn relevance signals via spe-cialized rankers. The result is a framework capable of estimating dissatisfaction from Web search logs and learning to improve per-formance for dissatisfied queries. Through experimentation, we show that our method yields good quality groups that align with established retrieval performance metrics. We also show that we can significantly improve retrieval effectiveness via specialized rankers, and that coherent grouping of underperforming queries generated by our method is important in improving each group. Ahmed Awadallah 0001, Ryen W. White, Yi-Min Wang |
SIGIR | 1 |
| 2013 | Playing by the rules: mining query associations to predict search performanceabstractUnderstanding the characteristics of queries where a search engine is failing is important for improving engine performance. Previous work largely relies on user-interaction features (e.g., clickthrough statistics) to identify such underperforming queries. However, relying on interaction behavior means that searchers need to become dissatisfied and need to exhibit that in their search behavior, by which point it may be too late to help them. In this paper, we propose a method to generate underperforming query identification rules instantly using topical and lexical attributes. The method first generates query attributes using sources such as topics, concepts (entities), and keywords in queries. Then, association rules are learned by exploiting the FP-growth algorithm and decision trees using underperforming query examples. We develop a query classification model capable of accurately estimating dissatisfaction using the generated rules, and demonstrate significant performance gains over state-of-the-art query performance prediction models. Ahmed Awadallah 0001, Ryen W. White, Yi-Min Wang |
WSDM | 2 |
| 2013 | Enhancing personalized search by mining and modeling task behaviorabstractPersonalized search systems tailor search results to the current user intent using historic search interactions. This relies on being able to find pertinent information in that user's search history, which can be challenging for unseen queries and for new search scenarios. Building richer models of users' current and historic search tasks can help improve the likelihood of finding relevant content and enhance the relevance and coverage of personalization methods. The task-based approach can be applied to the current user's search history, or as we focus on here, all users' search histories as so-called "groupization" (a variant of personalization whereby other users' profiles can be used to personalize the search experience). We describe a method whereby we mine historic search-engine logs to find other users performing similar tasks to the current user and leverage their on-task behavior to identify Web pages to promote in the current ranking. We investigate the effectiveness of this approach versus query-based matching and finding related historic activity from the current user (i.e., group versus individual). As part of our studies we also explore the use of the on-task behavior of particular user cohorts, such as people who are expert in the topic currently being searched, rather than all other users. Our approach yields promising gains in retrieval performance, and has direct implications for improving personalization in search systems. Ryen W. White, Ahmed Awadallah 0001, Xiaodong He 0001, Yang Song 0008, Hongning Wang |
WWW | 3 |
| 2012 | Task tours: helping users tackle complex search tasksabstractComplex search tasks such as planning a vacation often comprise multiple queries and may span a number of search sessions. When engaged in such tasks, users may require holistic support in determining the required task activities. Unfortunately, current search engines do not offer such support to their users. In this paper, we propose methods to automatically generate task tours comprising a starting task and a set of relevant related tasks, some or all of which may be necessary to satisfy a user's information needs. Applications of the tours include helping users understand the required steps to complete a task, finding URLs related to the active task, and alerting users to activities they may have missed. We demonstrate through experimentation with human judges and large-scale search logs that our tours are of good quality and can benefit a significant fraction of search engine users. Ahmed Awadallah 0001, Ryen W. White |
CIKM | 1 |
| 2012 | Detecting Subgroups in Online Discussions by Modeling Positive and Negative Relations among Participants
Ahmed Awadallah 0001, Amjad Abu-Jbara, Dragomir R. Radev |
EMNLP-CoNLL | 1 |
| 2012 | AttitudeMiner: Mining Attitude from Online Discussions
Amjad Abu-Jbara, Ahmed Awadallah 0001, Dragomir R. Radev |
HLT-NAACL | 2 |
| 2011 | A task level metric for measuring web search satisfaction and its application on improving relevance estimationabstractUnderstanding the behavior of satisfied and unsatisfied Web search users is very important for improving users search experience. Collecting labeled data that characterizes search behavior is a very challenging problem. Most of the previous work used a limited amount of data collected in lab studies or annotated by judges lacking information about the actual intent. In this work, we performed a large scale user study where we collected explicit judgments of user satisfaction with the entire search task. Results were analyzed using sequence models that incorporate user behavior to predict whether the user ended up being satisfied with a search or not. We test our metric on millions of queries collected from real Web search traffic and show empirically that user behavior models trained using explicit judgments of user satisfaction outperform several other search quality metrics. The proposed model can also be used to optimize different search engine components. We propose a method that uses task level success prediction to provide a better interpretation of clickthrough data. Clickthough data has been widely used to improve relevance estimation. We use our user satisfaction model to distinguish between clicks that lead to satisfaction and clicks that do not. We show that adding new features derived from this metric allowed us to improve the estimation of document relevance. Ahmed Awadallah 0001, Yang Song 0008, Li-wei He |
CIKM | 1 |
| 2010 | Identifying Text Polarity Using Random Walks
Ahmed Awadallah 0001, Dragomir R. Radev |
ACL | 1 |
| 2010 | What's with the Attitude? Identifying Sentences with Attitude in Online Discussions
Ahmed Awadallah 0001, Vahed Qazvinian, Dragomir R. Radev |
EMNLP | 1 |
| 2010 | Beyond DCG: user behavior as a predictor of a successful searchabstractWeb search engines are traditionally evaluated in terms of the relevance of web pages to individual queries. However, relevance of web pages does not tell the complete picture, since an individual query may represent only a piece of the user's information need and users may have different information needs underlying the same queries. In this work, we address the problem of predicting user search goal success by modeling user behavior. We show empirically that user behavior alone can give an accurate picture of the success of the user's web search goals, without considering the relevance of the documents displayed. In fact, our experiments show that models using user behavior are more predictive of goal success than those using document relevance. We build novel sequence models incorporating time distributions for this task and our experiments show that the sequence and time distribution models are more accurate than static models based on user behavior, or predictions based on document relevance. Ahmed Awadallah 0001, Rosie Jones, Kristina Lisa Klinkner |
WSDM | 1 |
| 2009 | A case study of using geographic cues to predict query news intentabstractGeographic information retrieval encompasses important tasks including finding the location of a user, and locations relevant to their search queries. Web-based search engines receive queries from numerous users located in very different parts of the world. A typical way for people to find news is through a general web search engine, which makes it important for search engines to recognize queries with news intent. An important question for geographic information retrieval is how we can benefit from geographic cues to predict the intent of users. This work presents a case study of an application using geographic features to improve the quality of an important web search task, involving predicting which queries have news intent and hence are likely to receive clicks on news search results. Our case study suggests that information derived from geographic features can help the task. The information we consider includes cues derived from the location of the user, from the IP address, the location relevant to the query, automatically extracted from the query string, and the relation between the two locations. We build a classifier that uses geographical cues to predict whether a query will result in a news click or not. We compare our classifier to a strong baseline that use non-geographic click-based features and we show that our classifier outperforms the baseline for geographic queries. Ahmed Awadallah 0001, Rosie Jones, Fernando Diaz 0001 |
GIS | 1 |
| 2009 | Content Based Recommendation and Summarization in the Blogosphere
Ahmed Awadallah 0001, Dragomir R. Radev, Junghoo Cho, Amruta Joshi |
ICWSM | 1 |
| 2009 | Using Citations to Generate surveys of Scientific Paradigms
Saif M. Mohammad, Bonnie J. Dorr, Melissa Egan, Ahmed Awadallah 0001, Pradeep Muthukrishnan, Vahed Qazvinian, Dragomir R. Radev, David M. Zajic |
HLT-NAACL | 4 |
| 2008 | Tracking the Dynamic Evolution of Participants Salience in a Discussion
Ahmed Awadallah 0001, Anthony Fader, Michael H. Crespin, Kevin M. Quinn, Burt L. Monroe, Michael P. Colaresi, Dragomir R. Radev |
COLING | 1 |
| 2008 | Language Independent Text Correction using Finite State Automata
Ahmed Awadallah 0001, Sara Noeman, Hany Hassan |
IJCNLP | 1 |
| 2006 | Unsupervised Information Extraction Approach Using Graph Mutual Reinforcement
Hany Hassan, Ahmed Awadallah 0001, Ossama Emam |
EMNLP | 2 |