Sameer Singh 0001

dblp:13/3568-1 · DBLP profile ↗
← Back
90ranked-venue papers
11as first author
39since 2021 · last 2026
0000-0003-0621-6323ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 82 · 10 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
abstract
Preethi Seshadri, Samuel Cahyawijaya, Ayomide Odumakinde, Sameer Singh, Seraphina Goldfarb-Tarrant. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Preethi Seshadri, Samuel Cahyawijaya, Ayomide Odumakinde, Sameer Singh 0001, Seraphina Goldfarb-Tarrant
ACL (1)4
2025 Nudging: Inference-time Alignment of LLMs via Guided Decoding
abstract
Large language models (LLMs) require alignment to effectively and safely follow user instructions. This process necessitates training an aligned version for every base model, resulting in significant computational overhead. In this work, we propose NUDGING, a simple, training-free algorithm that aligns any base model at inference time using a small aligned model. NUDGING is motivated by recent findings that alignment primarily alters the model’s behavior on a small subset of stylistic tokens (e.g., discourse markers). We find that base models are significantly more uncertain when generating these tokens. Building on this insight, NUDGING employs a small aligned model to generate nudging tokens to guide the base model’s output during decoding when the base model’s uncertainty is high, with only a minor additional inference overhead. We evaluate NUDGING across 3 model families on a diverse range of open-instruction tasks. Without any training, nudging a large base model with a 7×-14× smaller aligned model achieves zero-shot performance comparable to, and sometimes surpassing, that of large aligned models. By operating at the token level, NUDGING enables off-the-shelf collaboration between model families. For instance, nudging Gemma-2-27b with Llama-27b-chat outperforms Llama-2-70b-chat on various tasks. Overall, our work offers a modular and cost-efficient solution to LLM alignment. Our code and demo are available at: https://fywalter.github.io/nudging/.
Yasaman Razeghi, Sameer Singh 0001
ACL (1)3
2025 TurtleBench: A Visual Programming Benchmark in Turtle Geometry
abstract
Sina Rismanchian, Yasaman Razeghi, Sameer Singh, Shayan Doroudi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sina Rismanchian, Yasaman Razeghi, Sameer Singh 0001, Shayan Doroudi
NAACL (Long Papers)3
2024 Perceptions of Linguistic Uncertainty by Language Models and Humans
abstract
Uncertainty expressions such as "probably" or "highly unlikely" are pervasive in human language.While prior work has established that there is population-level agreement in terms of how humans quantitatively interpret these expressions, there has been little inquiry into the abilities of language models in the same context.In this paper, we investigate how language models map linguistic expressions of uncertainty to numerical responses.Our approach assesses whether language models can employ theory of mind in this setting: understanding the uncertainty of another agent about a particular statement, independently of the model's own certainty about that statement.We find that 7 out of 10 models are able to map uncertainty expressions to probabilistic responses in a human-like manner.However, we observe systematically different behavior depending on whether a statement is actually true or false.This sensitivity indicates that language models are substantially more susceptible to bias based on their prior knowledge (as compared to humans).These findings raise important questions and have broad implications for human-AI and AI-AI communication.
Catarina G. Belém, Markelle Rösti, Mark Steyvers, Sameer Singh 0001, Padhraic Smyth
EMNLP4
2024 Are Models Biased on Text without Gender-related Language?
abstract
Gender bias research has been pivotal in revealing undesirable behaviors in large language models, exposing serious gender stereotypes associated with occupations, and emotions. A key observation in prior work is that models reinforce stereotypes as a consequence of the gendered correlations that are present in the training data. In this paper, we focus on bias where the effect from training data is unclear, and instead address the question: *Do language models still exhibit gender bias in non-stereotypical settings?* To do so, we introduce **UnStereoEval (USE)**, a novel framework tailored for investigating gender bias in stereotype-free scenarios. USE defines a sentence-level score based on pretraining data statistics to determine if the sentence contain minimal word-gender associations. To systematically benchmark the fairness of popular language models in stereotype-free scenarios, we utilize USE to automatically generate benchmarks without any gender-related language. By leveraging USE's sentence-level score, we also repurpose prior gender bias benchmarks (Winobias and Winogender) for non-stereotypical evaluation. Surprisingly, we find low fairness across all 28 tested models. Concretely, models demonstrate fair behavior in only 9%-41% of stereotype-free sentences, suggesting that bias does not solely stem from the presence of gender-related words. These results raise important questions about where underlying model biases come from and highlight the need for more systematic and comprehensive bias evaluation. We release the full dataset and code at [ucinlp.github.io/unstereo-eval](https://ucinlp.github.io/unstereo-eval).
Catarina G. Belém, Preethi Seshadri, Yasaman Razeghi, Sameer Singh 0001
ICLR4
2024 What's In My Big Data?
abstract
Large text corpora are the backbone of language models. However, we have a limited understanding of the content of these corpora, including general statistics, quality, social factors, and inclusion of evaluation data (contamination). In this work, we propose What's In My Big Data? (WIMBD), a platform and a set of sixteen analyses that allow us to reveal and compare the contents of large text corpora. WIMBD builds on two basic capabilities---count and search---*at scale*, which allows us to analyze more than 35 terabytes on a standard compute node. We apply WIMBD to ten different corpora used to train popular language models, including *C4*, *The Pile*, and *RedPajama*. Our analysis uncovers several surprising and previously undocumented findings about these corpora, including the high prevalence of duplicate, synthetic, and low-quality content, personally identifiable information, toxic language, and benchmark contamination. For instance, we find that about 50% of the documents in *RedPajama* and *LAION-2B-en* are duplicates. In addition, several datasets used for benchmarking models trained on such corpora are contaminated with respect to important benchmarks, including the Winograd Schema Challenge and parts of GLUE and SuperGLUE. We open-source WIMBD's code and artifacts to provide a standard set of evaluations for new text-based corpora and to encourage more analyses and transparency around them.
Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Pete Walsh 0001, Dirk Groeneveld, Luca Soldaini, Sameer Singh 0001, Hannaneh Hajishirzi, Noah A. Smith, Jesse Dodge
ICLR10
2024 Skill Set Optimization: Reinforcing Language Model Behavior via Transferable Skills
abstract
Large language models (LLMs) have recently been used for sequential decision making in interactive environments. However, leveraging environment reward signals for continual LLM actor improvement is not straightforward. We propose Skill Set Optimization (SSO) for improving LLM actor performance through constructing and refining sets of transferable skills. SSO constructs skills by extracting common subtrajectories with high rewards and generating subgoals and instructions to represent each skill. These skills are provided to the LLM actor in-context to reinforce behaviors with high rewards. Then, SSO further refines the skill set by pruning skills that do not continue to result in high rewards. We evaluate our method in the classic videogame NetHack and the text environment ScienceWorld to demonstrate SSO's ability to optimize a set of skills and perform in-context policy improvement. SSO outperforms baselines by 40% in our custom NetHack task and outperforms the previous state-of-the-art in ScienceWorld by 35%.
Kolby Nottingham, Bodhisattwa Prasad Majumder, Bhavana Dalvi, Sameer Singh 0001, Peter Clark, Roy Fox
ICML4
2024 MisgenderMender: A Community-Informed Approach to Interventions for Misgendering
abstract
Tamanna Hossain, Sunipa Dev, Sameer Singh. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Tamanna Hossain, Sunipa Dev, Sameer Singh 0001
NAACL-HLT3
2024 The Bias Amplification Paradox in Text-to-Image Generation
abstract
Preethi Seshadri, Sameer Singh, Yanai Elazar. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Preethi Seshadri, Sameer Singh 0001, Yanai Elazar
NAACL-HLT2
2024 Benchmark Data Repositories for Better Benchmarking
abstract
In machine learning research, it is common to evaluate algorithms via their performance on standard benchmark datasets. While a growing body of work establishes guidelines for---and levies criticisms at---data and benchmarking practices in machine learning, comparatively less attention has been paid to the data repositories where these datasets are stored, documented, and shared. In this paper, we analyze the landscape of these benchmark data repositories and the role they can play in improving benchmarking. This role includes addressing issues with both datasets themselves (e.g., representational harms, construct validity) and the manner in which evaluation is carried out using such datasets (e.g., overemphasis on a few datasets and metrics, lack of reproducibility). To this end, we identify and discuss a set of considerations surrounding the design and use of benchmark data repositories, with a focus on improving benchmarking practices in machine learning.
Rachel Longjohn, Markelle Rösti, Sameer Singh 0001, Padhraic Smyth
NeurIPS3
2023 Maestro: A Gamified Platform for Teaching AI Robustness
abstract
Although the prevention of AI vulnerabilities is critical to preserve the safety and privacy of users and businesses, educational tools for robust AI are still underdeveloped worldwide. We present the design, implementation, and assessment of Maestro. Maestro is an effective open-source game-based platform that contributes to the advancement of robust AI education. Maestro provides "goal-based scenarios" where college students are exposed to challenging life-inspired assignments in a "competitive programming" environment. We assessed Maestro's influence on students' engagement, motivation, and learning success in robust AI. This work also provides insights into the design features of online learning tools that promote active learning opportunities in the robust AI domain. We analyzed the reflection responses (measured with Likert scales) of 147 undergraduate students using Maestro in two quarterly college courses in AI. According to the results, students who felt the acquisition of new skills in robust AI tended to appreciate highly Maestro and scored highly on material consolidation, curiosity, and maestry in robust AI. Moreover, the leaderboard, our key gamification element in Maestro, has effectively contributed to students' engagement and learning. Results also indicate that Maestro can be effectively adapted to any course length and depth without losing its educational quality.
Margarita Geleta, Jiacen Xu 0001, Manikanta Loya, Sameer Singh 0001, Zhou Li 0001, Sergio Gago Masagué
AAAI5
2023 Factual and Informative Review Generation for Explainable Recommendation
abstract
Recent models can generate fluent and grammatical synthetic reviews while accurately predicting user ratings. The generated reviews, expressing users' estimated opinions towards related products, are often viewed as natural language ‘rationales’ for the jointly predicted rating. However, previous studies found that existing models often generate repetitive, universally applicable, and generic explanations, resulting in uninformative rationales. Further, our analysis shows that previous models' generated content often contain factual hallucinations. These issues call for novel solutions that could generate both informative and factually grounded explanations. Inspired by recent success in using retrieved content in addition to parametric knowledge for generation, we propose to augment the generator with a personalized retriever, where the retriever's output serves as external knowledge for enhancing the generator. Experiments on Yelp, TripAdvisor, and Amazon Movie Reviews dataset show our model could generate explanations that more reliably entail existing reviews, are more diverse, and are rated more informative by human evaluators.
Zhouhang Xie, Sameer Singh 0001, Julian J. McAuley, Bodhisattwa Prasad Majumder
AAAI2
2023 To Adapt or to Annotate: Challenges and Interventions for Domain Adaptation in Open-Domain Question Answering
abstract
Recent advances in open-domain question answering (ODQA) have demonstrated impressive accuracy on general-purpose domains like Wikipedia.While some work has been investigating how well ODQA models perform when tested for out-of-domain (OOD) generalization, these studies have been conducted only under conservative shifts in data distribution and typically focus on a single component (i.e., retriever or reader) rather than an end-to-end system.This work proposes a more realistic endto-end domain shift evaluation setting covering five diverse domains.We not only find that endto-end models fail to generalize but that high retrieval scores often still yield poor answer prediction accuracy.To address these failures, we investigate several interventions, in the form of data augmentations, for improving model adaption and use our evaluation set to elucidate the relationship between the efficacy of an intervention scheme and the particular type of dataset shifts we consider.We propose a generalizability test that estimates the type of shift in a target dataset without training a model in the target domain and that the type of shift is predictive of which data augmentation schemes will be effective for domain adaption.Overall, we find that these interventions increase end-to-end performance by up to ∼24 points.* *This work was done while authors were at Google.Average F1 over all target datasets Average F1 over target datasets with specific shifts
Dheeru Dua, Emma Strubell, Sameer Singh 0001, Patrick Verga
ACL (1)3
2023 MISGENDERED: Limits of Large Language Models in Understanding Pronouns
abstract
Content Warning: This paper contains examples of misgendering and erasure that could be offensive and potentially triggering.Gender bias in language technologies has been widely studied, but research has mostly been restricted to a binary paradigm of gender.It is essential also to consider non-binary gender identities, as excluding them can cause further harm to an already marginalized group.In this paper, we comprehensively evaluate popular language models for their ability to correctly use English gender-neutral pronouns (e.g., singular they, them) and neo-pronouns (e.g., ze, xe, thon) that are used by individuals whose gender identity is not represented by binary pronouns.We introduce MISGENDERED, a framework for evaluating large language models' ability to correctly use preferred pronouns, consisting of (i) instances declaring an individual's pronoun, followed by a sentence with a missing pronoun, and (ii) an experimental setup for evaluating masked and auto-regressive language models using a unified method.When prompted outof-the-box, language models perform poorly at correctly predicting neo-pronouns (averaging 7.6% accuracy) and gender-neutral pronouns (averaging 31.0%accuracy).This inability to generalize results from a lack of representation of non-binary pronouns in training data and memorized associations.Few-shot adaptation with explicit examples in the prompt improves the performance but plateaus at only 45.4% for neo-pronouns.We release the full dataset, code, and demo at https://tamannahossainkay. github.io/misgendered/.
Tamanna Hossain, Sunipa Dev, Sameer Singh 0001
ACL (1)3
2023 Do Embodied Agents Dream of Pixelated Sheep: Embodied Decision Making using Language Guided World Modelling
abstract
Reinforcement learning (RL) agents typically learn tabula rasa, without prior knowledge of the world. However, if initialized with knowledge of high-level subgoals and transitions between subgoals, RL agents could utilize this Abstract World Model (AWM) for planning and exploration. We propose using few-shot large language models (LLMs) to hypothesize an AWM, that will be verified through world experience, to improve sample efficiency of RL agents. Our DECKARD agent applies LLM-guided exploration to item crafting in Minecraft in two phases: (1) the Dream phase where the agent uses an LLM to decompose a task into a sequence of subgoals, the hypothesized AWM; and (2) the Wake phase where the agent learns a modular policy for each subgoal and verifies or corrects the hypothesized AWM. Our method of hypothesizing an AWM with LLMs and then verifying the AWM based on agent experience not only increases sample efficiency over contemporary methods by an order of magnitude but is also robust to and corrects errors in the LLM, successfully blending noisy internet-scale information from LLMs with knowledge grounded in environment dynamics.
Kolby Nottingham, Prithviraj Ammanabrolu, Alane Suhr, Yejin Choi 0001, Hannaneh Hajishirzi, Sameer Singh 0001, Roy Fox
ICML6
2023 Post Hoc Explanations of Language Models Can Improve Language Models
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities in performing complex tasks. Moreover, recent research has shown that incorporating human-annotated rationales (e.g., Chain-of-Thought prompting) during in-context learning can significantly enhance the performance of these models, particularly on tasks that require reasoning capabilities. However, incorporating such rationales poses challenges in terms of scalability as this requires a high degree of human involvement. In this work, we present a novel framework, Amplifying Model Performance by Leveraging In-Context Learning with Post Hoc Explanations (AMPLIFY), which addresses the aforementioned challenges by automating the process of rationale generation. To this end, we leverage post hoc explanation methods which output attribution scores (explanations) capturing the influence of each of the input features on model predictions. More specifically, we construct automated natural language rationales that embed insights from post hoc explanations to provide corrective signals to LLMs. Extensive experimentation with real-world datasets demonstrates that our framework, AMPLIFY, leads to prediction accuracy improvements of about 10-25% over a wide range of tasks, including those where prior approaches which rely on human-annotated rationales such as Chain-of-Thought prompting fall short. Our work makes one of the first attempts at highlighting the potential of post hoc explanations as valuable tools for enhancing the effectiveness of LLMs. Furthermore, we conduct additional empirical analyses and ablation studies to demonstrate the impact of each of the components of AMPLIFY, which, in turn, lead to critical insights for refining in context learning.
Satyapriya Krishna, Jiaqi W. Ma, Dylan Slack, Asma Ghandeharioun, Sameer Singh 0001, Himabindu Lakkaraju
NeurIPS5
2023 Design Factors of Maestro: A Serious Game for Robust AI Education
abstract
Training tools targeting robust AI are still in their infancy. We present Maestro, an effective open-source game-based platform for robust AI training in higher education, which includes counter- measures and prevention of AI vulnerabilities. Maestro provides goal-based scenarios (GBSs) where students are exposed to challenging life-inspired assignments in a competitive programming environment. The assessment of Maestro showed that its leader-board, a key gamification element, has been crucial for effective student learning. Students who felt the acquisition of new skills in robust AI tended to appreciate highly Maestro and scored highly on material consolidation, curiosity and maestry in robust AI.
Margarita Geleta, Jiacen Xu 0001, Manikanta Loya, Sameer Singh 0001, Zhou Li 0001, Sergio Gago Masagué
SIGCSE (2)5
2023 Evaluating the generalisability of neural rumour verification models
abstract
Research on automated social media rumour verification, the task of identifying the veracity of questionable information circulating on social media, has yielded neural models achieving high performance, with accuracy scores that often exceed 90%. However, none of these studies focus on the real-world generalisability of the proposed approaches, that is whether the models perform well on datasets other than those on which they were initially trained and tested. In this work we aim to fill this gap by assessing the generalisability of top performing neural rumour verification models covering a range of different architectures from the perspectives of both topic and temporal robustness. For a more complete evaluation of generalisability, we collect and release COVID-RV, a novel dataset of Twitter conversations revolving around COVID-19 rumours. Unlike other existing COVID-19 datasets, our COVID-RV contains conversations around rumours that follow the format of prominent rumour verification benchmarks, while being different from them in terms of topic and time scale, thus allowing better assessment of the temporal robustness of the models. We evaluate model performance on COVID-RV and three popular rumour verification datasets to understand limitations and advantages of different model architectures, training datasets and evaluation scenarios. We find a dramatic drop in performance when testing models on a different dataset from that used for training. Further, we evaluate the ability of models to generalise in a few-shot learning setup, as well as when word embeddings are updated with the vocabulary of a new, unseen rumour. Drawing upon our experiments we discuss challenges and make recommendations for future research directions in addressing this important problem.
Elena Kochkina, Tamanna Hossain, Robert L. Logan IV, Miguel Arana-Catania, Rob Procter, Arkaitz Zubiaga, Sameer Singh 0001, Yulan He 0001, Maria Liakata
Inf. Process. Manag.7
2022 PYLON: A PyTorch Framework for Learning with Constraints
abstract
Deep learning excels at learning task information from large amounts of data, but struggles with learning from declarative high-level knowledge that can be more succinctly expressed directly. In this work, we introduce PYLON, a neuro-symbolic training framework that builds on PyTorch to augment procedurally trained models with declaratively specified knowledge. PYLON lets users programmatically specify constraints as Python functions and compiles them into a differentiable loss, thus training predictive models that fit the data whilst satisfying the specified constraints. PYLON includes both exact as well as approximate compilers to efficiently compute the loss, employing fuzzy logic, sampling methods, and circuits, ensuring scalability even to complex models and constraints. Crucially, a guiding principle in designing PYLON is the ease with which any existing deep learning codebase can be extended to learn from constraints in a few lines code: a function that expresses the constraint, and a single line to compile it into a loss. Our demo comprises of models in NLP, computer vision, logical games, and knowledge graphs that can be interactively trained using constraints as supervision.
Kareem Ahmed, Tao Li 0039, Thy Ton, Quan Guo, Kai-Wei Chang 0001, Parisa Kordjamshidi, Vivek Srikumar, Guy Van den Broeck, Sameer Singh 0001
AAAI9
2022 ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension
abstract
Sanjay Subramanian, William Merrill, Trevor Darrell, Matt Gardner, Sameer Singh, Anna Rohrbach. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Sanjay Subramanian, William Merrill, Trevor Darrell, Matt Gardner 0001, Sameer Singh 0001, Anna Rohrbach
ACL (1)5
2022 Successive Prompting for Decomposing Complex Questions
abstract
Answering complex questions that require making latent decisions is a challenging task, especially when limited supervision is available.Recent works leverage the capabilities of large language models (LMs) to perform complex question answering in a few-shot setting by demonstrating how to output intermediate rationalizations while solving the complex question in a single pass.We introduce "Successive Prompting", where we iteratively break down a complex task into a simple task, solve it, and then repeat the process until we get the final solution.Successive prompting decouples the supervision for decomposing complex questions from the supervision for answering simple questions, allowing us to (1) have multiple opportunities to query in-context examples at each reasoning step (2) learn question decomposition separately from question answering, including using synthetic data, and (3) use bespoke (fine-tuned) components for reasoning steps where a large LM does not perform well.The intermediate supervision is typically manually written, which can be expensive to collect.We introduce a way to generate a synthetic dataset which can be used to bootstrap a model's ability to decompose and answer intermediate questions.Our best model (with successive prompting) achieves an improvement of ∼5% absolute F1 on a few-shot version of the DROP dataset when compared with a stateof-the-art model with the same supervision.
Dheeru Dua, Shivanshu Gupta, Sameer Singh 0001, Matt Gardner 0001
EMNLP3
2022 Continued Pretraining for Better Zero- and Few-Shot Promptability
abstract
Recently introduced language model prompting methods can achieve high accuracy in zeroand few-shot settings while requiring few to no learned task-specific parameters.Nevertheless, these methods still often trail behind full model finetuning.In this work, we investigate if a dedicated continued pretraining stage could improve "promptability", i.e., zero-shot performance with natural language prompts or few-shot performance with prompt tuning.We reveal settings where existing continued pretraining methods lack promptability.We also identify current methodological gaps, which we fill with thorough large-scale experiments.We demonstrate that a simple recipe, continued pretraining that incorporates a trainable prompt during multi-task learning, leads to improved promptability in both zero-and fewshot settings compared to existing methods, up to 31% relative.On the other hand, we find that continued pretraining using MAML-style metalearning, a method that directly optimizes fewshot promptability, yields subpar performance.We validate our findings with two prompt tuning methods, and, based on our results, we provide concrete recommendations to optimize promptability for different use cases.
Zhaofeng Wu, Robert L. Logan IV, Pete Walsh 0001, Akshita Bhagia, Dirk Groeneveld, Sameer Singh 0001, Iz Beltagy
EMNLP6
2022 Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts
abstract
Daniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Daniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson 0001, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh 0001, Yejin Choi 0001
NAACL-HLT10
2022 FRUIT: Faithfully Reflecting Updated Information in Text
abstract
Robert Iv, Alexandre Passos, Sameer Singh, Ming-Wei Chang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Robert L. Logan IV, Alexandre Tachard Passos, Sameer Singh 0001, Ming-Wei Chang
NAACL-HLT3
2022 BottleFit: Learning Compressed Representations in Deep Neural Networks for Effective and Efficient Split Computing
abstract
Although mission-critical applications require the use of deep neural networks (DNNs), their continuous execution at mobile devices results in a significant increase in energy consumption. While edge offloading can decrease energy consumption, erratic patterns in channel quality, network and edge server load can lead to severe disruption of the system’s key operations. An alternative approach, called split computing, generates compressed representations within the model (called "bottlenecks"), to reduce bandwidth usage and energy consumption. Prior work has proposed approaches that introduce additional layers, to the detriment of energy consumption and latency. For this reason, we propose a new framework called BottleFit, which, in addition to targeted DNN architecture modifications, includes a novel training strategy to achieve high accuracy even with strong compression rates. We apply BottleFit on cutting-edge DNN models in image classification, and show that BottleFit achieves 77.1% data compression with up to 0.6% accuracy loss on ImageNet dataset, while state of the art such as SPINN loses up to 6% in accuracy. We experimentally measure the power consumption and latency of an image classification application running on an NVIDIA Jetson Nano board (GPU-based) and a Raspberry PI board (GPU-less). We show that BottleFit decreases power consumption and latency respectively by up to 49% and 89% with respect to (w.r.t.) local computing and by 37% and 55% w.r.t. edge offloading. We also compare BottleFit with state-of-the-art autoencoders-based approaches, and show that (i) BottleFit reduces power consumption and execution time respectively by up to 54% and 44% on the Jetson and 40% and 62% on Raspberry PI; (ii) the size of the head model executed on the mobile device is 83 times smaller. We publish the code repository for reproducibility of the results in this study.
Yoshitomo Matsubara, Davide Callegaro, Sameer Singh 0001, Marco Levorato, Francesco Restuccia 0001
WoWMoM3
2021 Improved Consistency Regularization for GANs
abstract
Recent work has increased the performance of Generative Adversarial Networks (GANs) by enforcing a consistency cost on the discriminator. We improve on this technique in several ways. We first show that consistency regularization can introduce artifacts into the GAN samples and explain how to fix this issue. We then propose several modifications to the consistency regularization procedure designed to improve its performance. We carry out extensive experiments quantifying the benefit of our improvements. For unconditional image synthesis on CIFAR-10 and CelebA, our modifications yield the best known FID scores on various GAN architectures. For conditional image synthesis on CIFAR-10, we improve the state-of-the-art FID score from 11.48 to 9.21. Finally, on ImageNet-2012, we apply our technique to the original BigGAN model and improve the FID from 6.66 to 5.38, which is the best score at that model size.
Zhengli Zhao, Sameer Singh 0001, Honglak Lee, Augustus Odena, Han Zhang 0010
AAAI2
2021 Evaluating Entity Disambiguation and the Role of Popularity in Retrieval-Based NLP
abstract
Anthony Chen, Pallavi Gudipati, Shayne Longpre, Xiao Ling, Sameer Singh. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Anthony Chen, Pallavi Gudipati, Shayne Longpre, Sameer Singh 0001
ACL/IJCNLP (1)5
2021 Benchmarking Scalable Methods for Streaming Cross Document Entity Coreference
abstract
Robert L Logan IV, Andrew McCallum, Sameer Singh, Dan Bikel. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Robert L. Logan IV, Andrew McCallum, Sameer Singh 0001, Dan Bikel
ACL/IJCNLP (1)3
2021 Competency Problems: On Finding and Removing Artifacts in Language Data
abstract
Much recent work in NLP has documented dataset artifacts, bias, and spurious correlations between input features and output labels.However, how to tell which features have "spurious" instead of legitimate correlations is typically left unspecified.In this work we argue that for complex language understanding tasks, all simple feature correlations are spurious, and we formalize this notion into a class of problems which we call competency problems.For example, the word "amazing" on its own should not give information about a sentiment label independent of the context in which it appears, which could include negation, metaphor, sarcasm, etc.We theoretically analyze the difficulty of creating data for competency problems when human bias is taken into account, showing that realistic datasets will increasingly deviate from competency problems as dataset size increases.This analysis gives us a simple statistical test for dataset artifacts, which we use to show more subtle biases than were described in prior work, including demonstrating that models are inappropriately affected by these less extreme biases.Our theoretical treatment of this problem also allows us to analyze proposed solutions, such as making local edits to dataset instances, and to give recommendations for future data collection and model design efforts that target competency problems.
Matt Gardner 0001, William Merrill, Jesse Dodge, Matthew E. Peters, Alexis Ross, Sameer Singh 0001, Noah A. Smith
EMNLP (1)6
2021 Learning with Instance Bundles for Reading Comprehension
abstract
When training most modern reading comprehension models, all the questions associated with a context are treated as being independent from each other.However, closely related questions and their corresponding answers are not independent, and leveraging these relationships could provide a strong supervision signal to a model.Drawing on ideas from contrastive estimation, we introduce several new supervision losses that compare question-answer scores across multiple related instances.Specifically, we normalize these scores across various neighborhoods of closely contrasting questions and/or answers, adding a cross entropy loss term in addition to traditional maximum likelihood estimation.Our techniques require bundles of related question-answer pairs, which we either mine from within existing data or create using automated heuristics.We empirically demonstrate the effectiveness of training with instance bundles on two datasets-HotpotQA and ROPES-showing up to 9% absolute gains in accuracy.
Dheeru Dua, Pradeep Dasigi, Sameer Singh 0001, Matt Gardner 0001
EMNLP (1)3
2021 Generative Context Pair Selection for Multi-hop Question Answering
abstract
Dheeru Dua, Cicero Nogueira dos Santos, Patrick Ng, Ben Athiwaratkun, Bing Xiang, Matt Gardner, Sameer Singh. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Dheeru Dua, Cícero Nogueira dos Santos, Patrick Ng, Ben Athiwaratkun, Bing Xiang, Matt Gardner 0001, Sameer Singh 0001
EMNLP (1)7
2021 Paired Examples as Indirect Supervision in Latent Decision Models
abstract
Compositional, structured models are appealing because they explicitly decompose problems and provide interpretable intermediate outputs that give confidence that the model is not simply latching onto data artifacts.Learning these models is challenging, however, because end-task supervision only provides a weak indirect signal on what values the latent decisions should take.This often results in the model failing to learn to perform the intermediate tasks correctly.In this work, we introduce a way to leverage paired examples that provide stronger cues for learning latent decisions.When two related training examples share internal substructure, we add an additional training objective to encourage consistency between their latent decisions.Such an objective does not require external supervision for the values of the latent output, or even the end task, yet provides an additional training signal to that provided by individual training examples themselves.We apply our method to improve compositional question answering using neural module networks on the DROP dataset.We explore three ways to acquire paired questions in DROP: (a) discovering naturally occurring paired examples within the dataset, (b) constructing paired examples using templates, and (c) generating paired examples using a question generation model.We empirically demonstrate that our proposed approach improves both in-and outof-distribution generalization and leads to correct latent decision predictions.
Nitish Gupta, Sameer Singh 0001, Matt Gardner 0001, Dan Roth 0001
EMNLP (1)2
2021 Entity-Based Knowledge Conflicts in Question Answering
abstract
Knowledge-dependent tasks typically use two sources of knowledge: parametric, learned at training time, and contextual, given as a passage at inference time.To understand how models use these sources together, we formalize the problem of knowledge conflicts, where the contextual information contradicts the learned information.Analyzing the behaviour of popular models, we measure their over-reliance on memorized information (the cause of hallucinations), and uncover important factors that exacerbate this behaviour.Lastly, we propose a simple method to mitigate over-reliance on parametric knowledge which minimizes hallucination and improves out-of-distribution generalization by 4% -7%.Our findings demonstrate the importance for practitioners to evaluate model tendency to hallucinate rather than read, and show that our mitigation strategy encourages generalization to evolving information (i.e., time-dependent queries).To encourage these practices, we have released our framework for generating knowledge conflicts.1
Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Christopher DuBois, Sameer Singh 0001
EMNLP (1)6
2021 Calibrate Before Use: Improving Few-shot Performance of Language Models
abstract
GPT-3 can perform numerous tasks when provided a natural language prompt that contains a few training examples. We show that this type of few-shot learning can be unstable: the choice of prompt format, training examples, and even the order of the examples can cause accuracy to vary from near chance to near state-of-the-art. We demonstrate that this instability arises from the bias of language models towards predicting certain answers, e.g., those that are placed near the end of the prompt or are common in the pre-training data. To mitigate this, we first estimate the model’s bias towards each answer by asking for its prediction when given a training prompt and a content-free test input such as "N/A". We then fit calibration parameters that cause the prediction for this input to be uniform across answers. On a diverse set of tasks, this contextual calibration procedure substantially improves GPT-3 and GPT-2’s accuracy (up to 30.0% absolute) across different choices of the prompt, while also making learning considerably more stable.
Eric Wallace, Shi Feng 0005, Daniel Klein 0001, Sameer Singh 0001
ICML5
2021 Beyond Accuracy: Behavioral Testing of NLP Models with Checklist (Extended Abstract)
abstract
Although measuring held-out accuracy has been the primary approach to evaluate generalization, it often overestimates the performance of NLP models, while alternative approaches for evaluating models either focus on individual tasks or on specific behaviors. Inspired by principles of behavioral testing in software engineering, we introduce CheckList, a task-agnostic methodology for testing NLP models. CheckList includes a matrix of general linguistic capabilities and test types that facilitate comprehensive test ideation, as well as a software tool to generate a large and diverse number of test cases quickly. We illustrate the utility of CheckList with tests for three tasks, identifying critical failures in both commercial and state-of-art models. In a user study, a team responsible for a commercial sentiment analysis model found new and actionable bugs in an extensively tested model. In another user study, NLP practitioners with CheckList created twice as many tests, and found almost three times as many bugs as users without it.
Marco Túlio Ribeiro, Sherry Tongshuang Wu, Carlos Guestrin, Sameer Singh 0001
IJCAI4
2021 An Empirical Comparison of Instance Attribution Methods for NLP
abstract
Pouya Pezeshkpour, Sarthak Jain, Byron Wallace, Sameer Singh. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Pouya Pezeshkpour, Byron C. Wallace, Sameer Singh 0001
NAACL-HLT4
2021 Concealed Data Poisoning Attacks on NLP Models
abstract
Adversarial attacks alter NLP model predictions by perturbing test-time inputs.However, it is much less understood whether, and how, predictions can be manipulated with small, concealed changes to the training data.In this work, we develop a new data poisoning attack that allows an adversary to control model predictions whenever a desired trigger phrase is present in the input.For instance, we insert 50 poison examples into a sentiment model's training set that causes the model to frequently predict Positive whenever the input contains "James Bond".Crucially, we craft these poison examples using a gradient-based procedure so that they do not mention the trigger phrase.We also apply our poison attack to language modeling ("Apple iPhone" triggers negative generations) and machine translation ("iced coffee" mistranslated as "hot coffee").We conclude by proposing three defenses that can mitigate our attack at some cost in prediction accuracy or extra human annotation.
Eric Wallace, Tony Z. Zhao, Shi Feng 0005, Sameer Singh 0001
NAACL-HLT4
2021 Counterfactual Explanations Can Be Manipulated
abstract
Counterfactual explanations are emerging as an attractive option for providing recourse to individuals adversely impacted by algorithmic decisions. As they are deployed in critical applications (e.g. law enforcement, financial lending), it becomes important to ensure that we clearly understand the vulnerabilties of these methods and find ways to address them. However, there is little understanding of the vulnerabilities and shortcomings of counterfactual explanations. In this work, we introduce the first framework that describes the vulnerabilities of counterfactual explanations and shows how they can be manipulated. More specifically, we show counterfactual explanations may converge to drastically different counterfactuals under a small perturbation indicating they are not robust. Leveraging this insight, we introduce a novel objective to train seemingly fair models where counterfactual explanations find much lower cost recourse under a slight perturbation. We describe how these models can unfairly provide low-cost recourse for specific subgroups in the data while appearing fair to auditors. We perform experiments on loan and violent crime prediction data sets where certain subgroups achieve up to 20x lower cost recourse under the perturbation. These results raise concerns regarding the dependability of current counterfactual explanation techniques, which we hope will inspire investigations in robust counterfactual explanations.
Dylan Slack, Anna Hilgard, Himabindu Lakkaraju, Sameer Singh 0001
NeurIPS4
2021 Reliable Post hoc Explanations: Modeling Uncertainty in Explainability
abstract
As black box explanations are increasingly being employed to establish model credibility in high stakes settings, it is important to ensure that these explanations are accurate and reliable. However, prior work demonstrates that explanations generated by state-of-the-art techniques are inconsistent, unstable, and provide very little insight into their correctness and reliability. In addition, these methods are also computationally inefficient, and require significant hyper-parameter tuning. In this paper, we address the aforementioned challenges by developing a novel Bayesian framework for generating local explanations along with their associated uncertainty. We instantiate this framework to obtain Bayesian versions of LIME and KernelSHAP which output credible intervals for the feature importances, capturing the associated uncertainty. The resulting explanations not only enable us to make concrete inferences about their quality (e.g., there is a 95% chance that the feature importance lies within the given range), but are also highly consistent and stable. We carry out a detailed theoretical analysis that leverages the aforementioned uncertainty to estimate how many perturbations to sample, and how to sample for faster convergence. This work makes the first attempt at addressing several critical issues with popular explanation methods in one shot, thereby generating consistent, stable, and reliable explanations with guarantees in a computationally efficient manner. Experimental evaluation with multiple real world datasets and user studies demonstrate that the efficacy of the proposed framework.
Dylan Slack, Anna Hilgard, Sameer Singh 0001, Himabindu Lakkaraju
NeurIPS3
2020 Minecraft as a Platform for Project-Based Learning in AI
abstract
Undergraduate courses that focus on open-ended, project-based learning teach students how to define concrete goals, transfer conceptual understanding of algorithms to code, and evaluate/analyze/present their solution. However, AI, along with machine learning, is getting increasingly varied in terms of both the approaches and applications, making it challenging to design project courses that span a sufficiently wide spectrum of AI. For these reasons, existing AI project courses are restricted to a narrow set of approaches (e.g. only reinforcement learning) or applications (e.g. only computer vision).In this paper, we propose to use Minecraft as the platform for teaching AI via project-based learning. Minecraft is an open-world sandbox game with elements of exploration, resource gathering, crafting, construction, and combat, and is supported by the Malmo library that provides a programmatic interface to the player observations and actions at various levels of granularity. In Minecraft, students can design projects to use approaches like search-based AI, reinforcement learning, supervised learning, and constraint satisfaction, on data types like text, audio, images, and tabular data. We describe our experience with an open-ended, undergraduate AI projects course using Minecraft that includes 82 different projects, covering themes that ranged from navigation, instruction following, object detection, combat, and music/image generation.
Sameer Singh 0001
AAAI1
2020 Benefits of Intermediate Annotations in Reading Comprehension
abstract
Complex, compositional reading comprehension datasets require performing latent sequential decisions that are learned via supervision from the final answer.A large combinatorial space of possible decision paths that result in the same answer, compounded by the lack of intermediate supervision to help choose the right path, makes the learning particularly hard for this task.In this work, we study the benefits of collecting intermediate reasoning supervision along with the answer during data collection.We find that these intermediate annotations can provide two-fold benefits.First, we observe that for any collection budget, spending a fraction of it on intermediate annotations results in improved model performance, for two complex compositional datasets: DROP and Quoref.Second, these annotations encourage the model to learn the correct latent reasoning steps, helping combat some of the biases introduced during the data collection process.
Dheeru Dua, Sameer Singh 0001, Matt Gardner 0001
ACL2
2020 Dynamic Sampling Strategies for Multi-Task Reading Comprehension
abstract
Building general reading comprehension systems, capable of solving multiple datasets at the same time, is a recent aspirational goal in the research community.Prior work has focused on model architectures or generalization to held out datasets, and largely passed over the particulars of the multi-task learning set up.We show that a simple dynamic sampling strategy, selecting instances for training proportional to the multi-task model's current performance on a dataset relative to its singletask performance, gives substantive gains over prior multi-task sampling strategies, mitigating the catastrophic forgetting that is common in multi-task learning.We also demonstrate that allowing instances of different tasks to be interleaved as much as possible between each epoch and batch has a clear benefit in multitask performance over forcing task homogeneity at the epoch or batch level.Our final model shows greatly increased performance over the best model on ORB, a recently-released multitask reading comprehension benchmark.
Ananth Gottumukkala, Dheeru Dua, Sameer Singh 0001, Matt Gardner 0001
ACL3
2020 On Importance Sampling-Based Evaluation of Latent Language Models
abstract
Language models that use additional latent structures (e.g., syntax trees, coreference chains, and knowledge graph links) provide several advantages over traditional language models.However, likelihood-based evaluation of these models is often intractable as it requires marginalizing over the latent space.Existing methods avoid this issue by using importance sampling.Although this approach has asymptotic guarantees, analysis is rarely conducted on the effect of decisions such as sample size, granularity of sample aggregation, and the proposal distribution on the reported estimates.In this paper, we measure the effect these factors have on perplexity estimates for three different latent language models.In addition, we elucidate subtle differences in how importance sampling is applied, which can have substantial effects on the final estimates, as well as provide theoretical results that reinforce the validity of importance sampling for evaluating latent language models.
Robert L. Logan IV, Matt Gardner 0001, Sameer Singh 0001
ACL3
2020 Beyond Accuracy: Behavioral Testing of NLP Models with CheckList
abstract
Although measuring held-out accuracy has been the primary approach to evaluate generalization, it often overestimates the performance of NLP models, while alternative approaches for evaluating models either focus on individual tasks or on specific behaviors.Inspired by principles of behavioral testing in software engineering, we introduce CheckList, a taskagnostic methodology for testing NLP models.CheckList includes a matrix of general linguistic capabilities and test types that facilitate comprehensive test ideation, as well as a software tool to generate a large and diverse number of test cases quickly.We illustrate the utility of CheckList with tests for three tasks, identifying critical failures in both commercial and state-of-art models.In a user study, a team responsible for a commercial sentiment analysis model found new and actionable bugs in an extensively tested model.In another user study, NLP practitioners with CheckList created twice as many tests, and found almost three times as many bugs as users without it.
Marco Túlio Ribeiro, Sherry Tongshuang Wu, Carlos Guestrin, Sameer Singh 0001
ACL4
2020 Obtaining Faithful Interpretations from Compositional Neural Networks
abstract
Neural module networks (NMNs) are a popular approach for modeling compositionality: they achieve high accuracy when applied to problems in language and vision, while reflecting the compositional structure of the problem in the network architecture.However, prior work implicitly assumed that the structure of the network modules, describing the abstract reasoning process, provides a faithful explanation of the model's reasoning; that is, that all modules perform their intended behaviour.In this work, we propose and conduct a systematic evaluation of the intermediate outputs of NMNs on NLVR2 and DROP, two datasets which require composing multiple reasoning steps.We find that the intermediate outputs differ from the expected output, illustrating that the network structure does not provide a faithful explanation of model behaviour.To remedy that, we train the model with auxiliary supervision and propose particular choices for module architecture that yield much better faithfulness, at a minimal cost to accuracy.
Sanjay Subramanian, Ben Bogin, Nitish Gupta, Tomer Wolfson, Sameer Singh 0001, Jonathan Berant, Matt Gardner 0001
ACL5
2020 Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods
abstract
As machine learning black boxes are increasingly being deployed in domains such as healthcare and criminal justice, there is growing emphasis on building tools and techniques for explaining these black boxes in an interpretable manner. Such explanations are being leveraged by domain experts to diagnose systematic errors and underlying biases of black boxes. In this paper, we demonstrate that post hoc explanations techniques that rely on input perturbations, such as LIME and SHAP, are not reliable. Specifically, we propose a novel scaffolding technique that effectively hides the biases of any given classifier by allowing an adversarial entity to craft an arbitrary desired explanation. Our approach can be used to scaffold any biased classifier in such a way that its predictions on the input data distribution still remain biased, but the post hoc explanations of the scaffolded classifier look innocuous. Using extensive evaluation with multiple real world datasets (including COMPAS), we demonstrate how extremely biased (racist) classifiers crafted by our framework can easily fool popular explanation techniques such as LIME and SHAP into generating innocuous explanations which do not reflect the underlying biases.
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh 0001, Himabindu Lakkaraju
AIES4
2020 MOCHA: A Dataset for Training and Evaluating Generative Reading Comprehension Metrics
abstract
Posing reading comprehension as a generation problem provides a great deal of flexibility, allowing for open-ended questions with few restrictions on possible answers.However, progress is impeded by existing generation metrics, which rely on token overlap and are agnostic to the nuances of reading comprehension.To address this, we introduce a benchmark for training and evaluating generative reading comprehension metrics: MOdeling Correctness with Human Annotations.MOCHA contains 40K human judgement scores on model outputs from 6 diverse question answering datasets and an additional set of minimal pairs for evaluation.Using MOCHA, we train a Learned Evaluation metric for Reading Comprehension, LERC, to mimic human judgement scores.LERC outperforms baseline metrics by 10 to 36 absolute Pearson points on held-out annotations.When we evaluate robustness on minimal pairs, LERC achieves 80% accuracy, outperforming baselines by 14 to 26 absolute percentage points while leaving significant room for improvement.MOCHA presents a challenging problem for developing accurate and robust generative reading comprehension metrics. 1
Anthony Chen, Gabriel Stanovsky, Sameer Singh 0001, Matt Gardner 0001
EMNLP (1)3
2020 AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
abstract
The remarkable success of pretrained language models has motivated the study of what kinds of knowledge these models learn during pretraining.Reformulating tasks as fillin-the-blanks problems (e.g., cloze tests) is a natural approach for gauging such knowledge, however, its usage is limited by the manual effort and guesswork required to write suitable prompts.To address this, we develop AUTOPROMPT, an automated method to create prompts for a diverse set of tasks, based on a gradient-guided search.Using AUTO-PROMPT, we show that masked language models (MLMs) have an inherent capability to perform sentiment analysis and natural language inference without additional parameters or finetuning, sometimes achieving performance on par with recent state-of-the-art supervised models.We also show that our prompts elicit more accurate factual knowledge from MLMs than the manually created prompts on the LAMA benchmark, and that MLMs can be used as relation extractors more effectively than supervised relation extraction models.These results demonstrate that automatically generated prompts are a viable parameter-free alternative to existing probing methods, and as pretrained LMs become more sophisticated and capable, potentially a replacement for finetuning.
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, Sameer Singh 0001
EMNLP (1)5
2020 Neural Module Networks for Reasoning over Text
Nitish Gupta, Dan Roth 0001, Sameer Singh 0001, Matt Gardner 0001
ICLR4
2020 Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature Attribution
Nikaash Puri, Sukriti Verma, Dhruv Kayastha, Shripad V. Deshmukh, Balaji Krishnamurthy, Sameer Singh 0001
ICLR7
2020 Building a Better Lie Detector with BERT: The Difference Between Truth and Lies
abstract
Detecting lies or deceptive statements in text is a valuable skill. This is partly because the patterns that underlie deceptive text are not known. The aim of this work is to identify patterns that characterize deceptive text. A key step in this approach is to train a classifier based on the BERT (Bidirectional Encoder Representations from Transformers) network. BERT beats the state of the art in deception classification accuracy on the Ott Deceptive Opinion Spam corpus. The results of our ablation study indicate that certain components of the input, such as some parts of speech, are more informative to the classifier than others. Further part-of-speech analysis in "swing" sentences that are considered important to BERT's classification indicates that deceptive text is more formulaic and less varied than truthful text. We expanded our classifier into a new Generative Adversarial Network based on BERT to create exemplars of deceptive and truthful text that further showed the differences between truth and deception, reinforcing the underlying similarity of deceptive text in terms of part-of-speech makeup.
Dan Barsever, Sameer Singh 0001, Emre Neftci
IJCNN2
2019 Barack's Wife Hillary: Using Knowledge Graphs for Fact-Aware Language Modeling
abstract
Modeling human language requires the ability to not only generate fluent text but also encode factual knowledge.However, traditional language models are only capable of remembering facts seen at training time, and often have difficulty recalling them.To address this, we introduce the knowledge graph language model (KGLM), a neural language model with mechanisms for selecting and copying facts from a knowledge graph that are relevant to the context.These mechanisms enable the model to render information it has never seen before, as well as generate out-of-vocabulary tokens.We also introduce the Linked WikiText-2 dataset, 1 a corpus of annotated text aligned to the Wikidata knowledge graph whose contents (roughly) match the popular WikiText-2 benchmark (Merity et al., 2017).In experiments, we demonstrate that the KGLM achieves significantly better performance than a strong baseline language model.We additionally compare different language models' ability to complete sentences requiring factual knowledge, and show that the KGLM outperforms even very large language models in generating facts.
Robert L. Logan IV, Nelson F. Liu, Matthew E. Peters, Matt Gardner 0001, Sameer Singh 0001
ACL (1)5
2019 Compositional Questions Do Not Necessitate Multi-hop Reasoning
abstract
Multi-hop reading comprehension (RC) questions are challenging because they require reading and reasoning over multiple paragraphs.We argue that it can be difficult to construct large multi-hop RC datasets.For example, even highly compositional questions can be answered with a single hop if they target specific entity types, or the facts needed to answer them are redundant.Our analysis is centered on HOTPOTQA, where we show that single-hop reasoning can solve much more of the dataset than previously thought.We introduce a single-hop BERT-based RC model that achieves 67 F1-comparable to state-of-theart multi-hop models.We also design an evaluation setting where humans are not shown all of the necessary paragraphs for the intended multi-hop reasoning but can still answer over 80% of questions.Together with detailed error analysis, these results suggest there should be an increasing focus on the role of evidence in multi-hop reasoning and possibly even a shift towards information retrieval style evaluations with large and diverse evidence collections.
Sewon Min, Eric Wallace, Sameer Singh 0001, Matt Gardner 0001, Hannaneh Hajishirzi, Luke Zettlemoyer
ACL (1)3
2019 Are Red Roses Red? Evaluating Consistency of Question-Answering Models
abstract
Although current evaluation of questionanswering systems treats predictions in isolation, we need to consider the relationship between predictions to measure true understanding.A model should be penalized for answering "no" to "Is the rose red?" if it answers "red" to "What color is the rose?".We propose a method to automatically extract such implications for instances from two QA datasets, VQA and SQuAD, which we then use to evaluate the consistency of models.Human evaluation shows these generated implications are well formed and valid.Consistency evaluation provides crucial insights into gaps in existing models, and retraining with implicationaugmented data improves consistency on both synthetic and human-generated implications.
Marco Túlio Ribeiro, Carlos Guestrin, Sameer Singh 0001
ACL (1)3
2019 Knowledge Enhanced Contextual Word Representations
abstract
Matthew E. Peters, Mark Neumann, Robert Logan, Roy Schwartz, Vidur Joshi, Sameer Singh, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Matthew E. Peters, Mark Neumann, Robert L. Logan IV, Roy Schwartz 0001, Vidur Joshi, Sameer Singh 0001, Noah A. Smith
EMNLP/IJCNLP (1)6
2019 Universal Adversarial Triggers for Attacking and Analyzing NLP
abstract
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, Sameer Singh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Eric Wallace, Shi Feng 0005, Nikhil Kandpal, Matt Gardner 0001, Sameer Singh 0001
EMNLP/IJCNLP (1)5
2019 Do NLP Models Know Numbers? Probing Numeracy in Embeddings
abstract
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh, Matt Gardner. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh 0001, Matt Gardner 0001
EMNLP/IJCNLP (1)4
2019 Detecting conversation topics in primary care office visits from transcripts of patient-provider interactions
abstract
OBJECTIVE: Amid electronic health records, laboratory tests, and other technology, office-based patient and provider communication is still the heart of primary medical care. Patients typically present multiple complaints, requiring physicians to decide how to balance competing demands. How this time is allocated has implications for patient satisfaction, payments, and quality of care. We investigate the effectiveness of machine learning methods for automated annotation of medical topics in patient-provider dialog transcripts. MATERIALS AND METHODS: We used dialog transcripts from 279 primary care visits to predict talk-turn topic labels. Different machine learning models were trained to operate on single or multiple local talk-turns (logistic classifiers, support vector machines, gated recurrent units) as well as sequential models that integrate information across talk-turn sequences (conditional random fields, hidden Markov models, and hierarchical gated recurrent units). RESULTS: Evaluation was performed using cross-validation to measure 1) classification accuracy for talk-turns and 2) precision, recall, and F1 scores at the visit level. Experimental results showed that sequential models had higher classification accuracy at the talk-turn level and higher precision at the visit level. Independent models had higher recall scores at the visit level compared with sequential models. CONCLUSIONS: Incorporating sequential information across talk-turns improves the accuracy of topic prediction in patient-provider dialog by smoothing out noisy information from talk-turns. Although the results are promising, more advanced prediction techniques and larger labeled datasets will likely be required to achieve prediction performance appropriate for real-world clinical applications.
Dimitrios Kotzias, Patty Kuo, Robert L. Logan IV, Kritzia Merced, Sameer Singh 0001, Michael Tanana, Efi Karra Taniskidou, Jennifer Elston-Lafata, David C. Atkins, Ming Tai-Seale, Zac E. Imel, Padhraic Smyth
J. Am. Medical Informatics Assoc.6
2018 Anchors: High-Precision Model-Agnostic Explanations
abstract
We introduce a novel model-agnostic system that explains the behavior of complex models with high-precision rules called anchors, representing local, "sufficient" conditions for predictions. We propose an algorithm to efficiently compute these explanations for any black-box model with high-probability guarantees. We demonstrate the flexibility of anchors by explaining a myriad of different models for different domains and tasks. In a user study, we show that anchors enable users to predict how a model would behave on unseen instances with less effort and higher precision, as compared to existing linear explanations or no explanations.
Marco Túlio Ribeiro, Sameer Singh 0001, Carlos Guestrin
AAAI2
2018 Semantically Equivalent Adversarial Rules for Debugging NLP models
abstract
Complex machine learning models for NLP are often brittle, making different predictions for input instances that are extremely similar semantically. To automatically detect this behavior for individual instances, we present semantically equivalent adversaries (SEAs) – semantic-preserving perturbations that induce changes in the model’s predictions. We generalize these adversaries into semantically equivalent adversarial rules (SEARs) – simple, universal replacement rules that induce adversaries on many instances. We demonstrate the usefulness and flexibility of SEAs and SEARs by detecting bugs in black-box state-of-the-art models for three domains: machine comprehension, visual question-answering, and sentiment analysis. Via user studies, we demonstrate that we generate high-quality local adversaries for more instances than humans, and that SEARs induce four times as many mistakes as the bugs discovered by human experts. SEARs are also actionable: retraining models using data augmentation significantly reduces bugs, while maintaining accuracy.
Marco Túlio Ribeiro, Sameer Singh 0001, Carlos Guestrin
ACL (1)2
2018 Embedding Multimodal Relational Data for Knowledge Base Completion
abstract
Representing entities and relations in an embedding space is a well-studied approach for machine learning on relational data.Existing approaches, however, primarily focus on simple link structure between a finite set of entities, ignoring the variety of data types that are often used in knowledge bases, such as text, images, and numerical values.In this paper, we propose multimodal knowledge base embeddings (MKBE) that use different neural encoders for this variety of observed data, and combine them with existing relational models to learn embeddings of the entities and multimodal data.Further, using these learned embedings and different neural decoders, we introduce a novel multimodal imputation model to generate missing multimodal values, like text and images, from information in the knowledge base.We enrich existing relational datasets to create two novel benchmarks that contain additional information such as textual descriptions and images of the original entities.We demonstrate that our models utilize this additional information effectively to provide more accurate link prediction, achieving state-of-the-art results with a considerable gap of 5-7% over existing methods.Further, we evaluate the quality of our generated multimodal values via a user study.We have release the datasets and the opensource implementation of our models at https: //github.com/pouyapez/mkbe.
Pouya Pezeshkpour, Sameer Singh 0001
EMNLP3
2018 Interpretation of Natural Language Rules in Conversational Machine Reading
abstract
Marzieh Saeidi, Max Bartolo, Patrick Lewis, Sameer Singh, Tim Rocktäschel, Mike Sheldon, Guillaume Bouchard, Sebastian Riedel. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Marzieh Saeidi, Max Bartolo, Patrick S. H. Lewis, Sameer Singh 0001, Tim Rocktäschel, Mike Sheldon, Guillaume Bouchard, Sebastian Riedel 0001
EMNLP4
2018 Combining Symbolic Expressions and Black-box Function Evaluations in Neural Programs
Forough Arabshahi, Sameer Singh 0001, Anima Anandkumar
ICLR (Poster)2
2018 Generating Natural Adversarial Examples
Zhengli Zhao, Dheeru Dua, Sameer Singh 0001
ICLR (Poster)3
2018 Mining Knowledge Graphs From Text
abstract
Knowledge graphs have become an increasingly crucial component in machine intelligence systems, powering ubiquitous digital assistants and inspiring several large scale academic projects across the globe. Our tutorial explains why knowledge graphs are important, how knowledge graphs are constructed, and where new research opportunities exist for improving the state-of-the-art. In this tutorial, we cover the many sophisticated approaches that complete and correct knowledge graphs. We organize this exploration into two main classes of models. The first include probabilistic logical frameworks that use graphical models, random walks, or statistical rule mining to construct knowledge graphs. The second class of models includes latent space models such as matrix and tensor factorization and neural networks. We conclude the tutorial with a critical comparison of techniques and results. We will offer practical advice for novices to identify common empirical challenges and concrete data sets for initial experimentation. Finally, we will highlight promising areas of current and future work.
Jay Pujara, Sameer Singh 0001
WSDM2
2018 A Framework of Rapid Regional Tsunami Damage Recognition From Post-event TerraSAR-X Imagery Using Deep Neural Networks
abstract
Near real-time building damage mapping is an indispensable prerequisite for governments to make decisions for disaster relief. With high-resolution synthetic aperture radar (SAR) systems, such as TerraSAR-X, the provision of such products in a fast and effective way becomes possible. In this letter, a deep learning-based framework for rapid regional tsunami damage recognition using post-event SAR imagery is proposed. To perform such a rapid damage mapping, a series of tile-based image split analysis is employed to generate the data set. Next, a selection algorithm with the SqueezeNet network is developed to swiftly distinguish between built-up (BU) and nonbuilt-up regions. Finally, a recognition algorithm with a modified wide residual network is developed to classify the BU regions into wash away, collapsed, and slightly damaged regions. Experiments performed on the TerraSAR-X data from the 2011 Tohoku earthquake and tsunami in Japan show a BU region extraction accuracy of 80.4% and a damage-level recognition accuracy of 74.8%, respectively. Our framework takes around 2 h to train on a new region, and only several minutes for prediction.
Yanbing Bai, Chang Gao 0001, Sameer Singh 0001, Magaly Koch, Bruno Adriano, Erick Mas, Shunichi Koshimura
IEEE Geosci. Remote. Sens. Lett.3
2017 Entity Linking via Joint Encoding of Types, Descriptions, and Context
abstract
For accurate entity linking, we need to capture various information aspects of an entity, such as its description in a KB, contexts in which it is mentioned, and structured knowledge.Additionally, a linking system should work on texts from different domains without requiring domain-specific training data or hand-engineered features.In this work we present a neural, modular entity linking system that learns a unified dense representation for each entity using multiple sources of information, such as its description, contexts around its mentions, and its fine-grained types.We show that the resulting entity linking system is effective at combining these sources, and performs competitively, sometimes out-performing current state-of-theart systems across datasets, without requiring any domain-specific training data or hand-engineered features.We also show that our model can effectively "embed" entities that are new to the KB, and is able to link its mentions accurately.
Nitish Gupta, Sameer Singh 0001, Dan Roth 0001
EMNLP2
2016 Creating Interactive and Visual Educational Resources for AI
abstract
Teaching artificial intelligence is effective if the experience is a visual and interactive one, with educational materials that utilize combinations of various content types such as text, math, and code into an integrated experience. Unfortunately, easy-to-use tools for creating such pedagogical resources are not available to the educators, resulting in most courses being taught using a disconnected set of static materials, which is not only ineffective for learning AI, but further, requires repeated and redundant effort for the instructor. In this paper, we introduce Moro, a software tool for easily creating and presenting AI-friendly teaching materials. Moro notebooks integrate content of different types (text, math, code, images), allow real-time interactions via modifiable and executable code blocks, and are viewable in browsers both as long-form pages and as presentations. Creating notebooks is easy and intuitive; the creation tool is also in-browser, is WYSIWYG for quick iterations of editing, and supports a variety of shortcuts and customizations for efficiency. We present three deployed case studies of Moro that widely differ from each other, demonstrating its utility in a variety of scenarios such as in-class teaching and conference tutorials.
Sameer Singh 0001, Sebastian Riedel 0001
AAAI1
2016 Connotation Frames: A Data-Driven Investigation
abstract
Through a particular choice of a predicate (e.g., "x violated y"), a writer can subtly connote a range of implied sentiment and presupposed facts about the entities x and y: (1) writer's perspective: projecting x as an "antagonist" and y as a "victim", (2) entities' perspective: y probably dislikes x, (3) effect: something bad happened to y, (4) value: y is something valuable, and (5) mental state: y is distressed by the event.We introduce connotation frames as a representation formalism to organize these rich dimensions of connotation using typed relations.First, we investigate the feasibility of obtaining connotative labels through crowdsourcing experiments.We then present models for predicting the connotation frames of verb predicates based on their distributional word representations and the interplay between different types of connotative relations.Empirical results confirm that connotation frames can be induced from various data sources that reflect how language is used in context.We conclude with analytical results that show the potential use of connotation frames for analyzing subtle biases in online news media.
Hannah Rashkin, Sameer Singh 0001, Yejin Choi 0001
ACL (1)2
2016 Better call Saul: Flexible Programming for Learning and Inference in NLP
abstract
We present a novel way for designing complex joint inference and learning models using Saul (Kordjamshidi et al., 2015), a recently-introduced declarative learning-based programming language (DeLBP). We enrich Saul with components that are necessary for a broad range of learning based Natural Language Processing tasks at various levels of granularity. We illustrate these advances using three different, well-known NLP problems, and show how these generic learning and inference modules can directly exploit Saul’s graph-based data representation. These properties allow the programmer to easily switch between different model formulations and configurations, and consider various kinds of dependencies and correlations among variables of interest with minimal programming effort. We argue that Saul provides an extremely useful paradigm both for the design of advanced NLP systems and for supporting advanced research in NLP.
Parisa Kordjamshidi, Daniel Khashabi, Christos Christodoulopoulos 0001, Bhargav Mangipudi, Sameer Singh 0001, Dan Roth 0001
COLING5
2016 "Why Should I Trust You?": Explaining the Predictions of Any Classifier
abstract
Despite widespread adoption, machine learning models remain mostly black boxes. Understanding the reasons behind predictions is, however, quite important in assessing trust, which is fundamental if one plans to take action based on a prediction, or when choosing whether to deploy a new model. Such understanding also provides insights into the model, which can be used to transform an untrustworthy model or prediction into a trustworthy one.
Marco Túlio Ribeiro, Sameer Singh 0001, Carlos Guestrin
KDD2
2015 Efficient Second-Order Gradient Boosting for Conditional Random Fields
abstract
Conditional random fields (CRFs) are an important class of models for accurate structured prediction, but effective design of the feature functions is a major challenge when applying CRF models to real world data. Gradient boosting, which is used to automatically induce and select feature functions, is a natural candidate solution to the problem. However, it is non-trivial to derive gradient boosting algorithms for CRFs due to the dense Hessian matrices introduced by variable dependencies. Existing approaches thus use only first-order information when optimizing likelihood, and hence face convergence issues. We incorporate second-order information by deriving a Markov Chain mixing rate bound to quantify the dependencies, and introduce a gradient boosting algorithm that iteratively optimizes an adaptive upper bound of the objective function. The resulting algorithm induces and selects features for CRFs via functional space optimization, with provable convergence guarantees. Experimental results on three real world datasets demonstrate that the mixing rate based upper bound is effective for learning CRFs with non-linear potentials.
Tianqi Chen 0001, Sameer Singh 0001, Ben Taskar, Carlos Guestrin
AISTATS2
2015 Injecting Logical Background Knowledge into Embeddings for Relation Extraction
abstract
Matrix factorization approaches to relation extraction provide several attractive features: they support distant supervision, handle open schemas, and leverage unlabeled data.Unfortunately, these methods share a shortcoming with all other distantly supervised approaches: they cannot learn to extract target relations without existing data in the knowledge base, and likewise, these models are inaccurate for relations with sparse data.Rule-based extractors, on the other hand, can be easily extended to novel relations and improved for existing but inaccurate relations, through first-order formulae that capture auxiliary domain knowledge.However, usually a large set of such formulae is necessary to achieve generalization.In this paper, we introduce a paradigm for learning low-dimensional embeddings of entity-pairs and relations that combine the advantages of matrix factorization with first-order logic domain knowledge.We introduce simple approaches for estimating such embeddings, as well as a novel training algorithm to jointly optimize over factual and first-order logic information.Our results show that this method is able to learn accurate extractors with little or no distant supervision alignments, while at the same time generalizing to textual patterns that do not appear in the formulae.
Tim Rocktäschel, Sameer Singh 0001, Sebastian Riedel 0001
HLT-NAACL2
2015 WOLFE: An NLP-friendly Declarative Machine Learning Stack
abstract
Sameer Singh, Tim Rocktäschel, Luke Hewitt, Jason Naradowsky, Sebastian Riedel. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. 2015.
Sameer Singh 0001, Tim Rocktäschel, Luke Hewitt, Jason Naradowsky, Sebastian Riedel 0001
HLT-NAACL1
2015 Design Challenges for Entity Linking
abstract
Recent research on entity linking (EL) has introduced a plethora of promising techniques, ranging from deep neural networks to joint inference. But despite numerous papers there is surprisingly little understanding of the state of the art in EL. We attack this confusion by analyzing differences between several versions of the EL problem and presenting a simple yet effective, modular, unsupervised system, called Vinculum, for entity linking. We conduct an extensive evaluation on nine data sets, comparing Vinculum with two state-of-the-art systems, and elucidate key aspects of the system that include mention extraction, candidate generation, entity type prediction, entity coreference, and coherence.
Sameer Singh 0001, Daniel S. Weld
Trans. Assoc. Comput. Linguistics2
2013 Automated probabilistic modeling for relational data
abstract
Probabilistic graphical model representations of relational data provide a number of desired features, such as inference of missing values, detection of errors, visualization of data, and probabilistic answers to relational queries. However, adoption has been slow due to the high level of expertise expected both in probability and in the domain from the user. Instead of requiring a domain expert to specify the probabilistic dependencies of the data, we present an approach that uses the relational DB schema to automatically construct a Bayesian graphical model for a database. This resulting model contains customized distributions for the attributes, latent variables that cluster the records, and factors that reflect and represent the foreign key links, whilst allowing efficient inference. Experiments demonstrate the accuracy of the model and scalability of inference on synthetic and real-world data.
Sameer Singh 0001, Thore Graepel
CIKM1
2013 AKBC 2013: third workshop on automated knowledge base construction
abstract
The AKBC 2013 workshop aims to be a venue of excellence and vision in the area of knowledge base construction. This year's workshop will feature keynotes by ten leading researchers in the field, including from Google, Microsoft, Stanford, and CMU. The submissions focus on visionary ideas instead of on experimental evaluation. Nineteen accepted papers will be presented as posters, with nine exceptional papers also highlighted as spotlight talks. Thereby, the workshop aims provides a vivid forum of discussion about the field of automated knowledge base construction.
Fabian M. Suchanek, Sebastian Riedel 0001, Sameer Singh 0001, Partha P. Talukdar
CIKM3
2013 Dynamic Knowledge-Base Alignment for Coreference Resolution
Jiaping Zheng, Luke Vilnis, Sameer Singh 0001, Jinho D. Choi, Andrew McCallum
CoNLL3
2012 A Discriminative Hierarchical Model for Fast Coreference at Large Scale
Michael L. Wick, Sameer Singh 0001, Andrew McCallum
ACL (1)2
2012 Monte Carlo MCMC: Efficient Inference by Approximate Sampling
Sameer Singh 0001, Michael L. Wick, Andrew McCallum
EMNLP-CoNLL1
2011 Large-Scale Cross-Document Coreference Using Distributed Inference and Hierarchical Models
Sameer Singh 0001, Amarnag Subramanya, Fernando Pereira 0003, Andrew McCallum
ACL1
2010 Minimally-Supervised Extraction of Entities from Text Advertisements
Sameer Singh 0001, Dustin Hillard, Chris Leggetter
HLT-NAACL1
2010 Constraint-Driven Rank-Based Learning for Information Extraction
Sameer Singh 0001, Limin Yao, Sebastian Riedel 0001, Andrew McCallum
HLT-NAACL1
2009 FACTORIE: Probabilistic Programming via Imperatively Defined Factor Graphs
abstract
Discriminatively trained undirected graphical models have had wide empirical success, and there has been increasing interest in toolkits that ease their application to complex relational data. The power in relational models is in their repeated structure and tied parameters; at issue is how to define these structures in a powerful and flexible way. Rather than using a declarative language, such as SQL or first-order logic, we advocate using an imperative language to express various aspects of model structure, inference, and learning. By combining the traditional, declarative, statistical semantics of factor graphs with imperative definitions of their construction and operation, we allow the user to mix declarative and procedural domain knowledge, and also gain significant efficiencies. We have implemented such imperatively defined factor graphs in a system we call Factorie, a software library for an object-oriented, strongly-typed, functional language. In experimental comparisons to Markov Logic Networks on joint segmentation and coreference, we find our approach to be 3-15 times faster while reducing error by 20-25%-achieving a new state of the art.
Andrew McCallum, Karl Schultz, Sameer Singh 0001
NIPS3
2009 Training Factor Graphs with Reinforcement Learning for Efficient MAP Inference
abstract
Large, relational factor graphs with structure defined by first-order logic or other languages give rise to notoriously difficult inference problems. Because unrolling the structure necessary to represent distributions over all hypotheses has exponential blow-up, solutions are often derived from MCMC. However, because of limitations in the design and parameterization of the jump function, these sampling-based methods suffer from local minima|the system must transition through lower-scoring configurations before arriving at a better MAP solution. This paper presents a new method of explicitly selecting fruitful downward jumps by leveraging reinforcement learning (RL). Rather than setting parameters to maximize the likelihood of the training data, parameters of the factor graph are treated as a log-linear function approximator and learned with temporal difference (TD); MAP inference is performed by executing the resulting policy on held out test data. Our method allows efficient gradient updates since only factors in the neighborhood of variables affected by an action need to be computed|we bypass the need to compute marginals entirely. Our method provides dramatic empirical success, producing new state-of-the-art results on a complex joint model of ontology alignment, with a 48\% reduction in error over state-of-the-art in that domain.
Michael L. Wick, Khashayar Rohanimanesh, Sameer Singh 0001, Andrew McCallum
NIPS3
2009 Bi-directional Joint Inference for Entity Resolution and Segmentation Using Imperatively-Defined Factor Graphs
Sameer Singh 0001, Karl Schultz, Andrew McCallum
ECML/PKDD (2)1
2009 Parallel Large Scale Feature Selection for Logistic Regression
abstract
In this paper we examine the problem of efficient feature evaluation for logistic regression on very large data sets. We present a new forward feature selection heuristic that ranks features by their estimated effect on the resulting model's performance. An approximate optimization, based on backfitting, provides a fast and accurate estimate of each new feature's coefficient in the logistic regression model. Further, the algorithm is highly scalable by parallelizing simultaneously over both features and records, allowing us to quickly evaluate billions of potential features even for very large data sets.
Sameer Singh 0001, Jeremy Kubica, Scott Larsen, Daria Sorokina
SDM1
2007 Fine-grain analysis of common coupling and its application to a Linux case study
Dror G. Feitelson, Tokunbo O. S. Adeshiyan, Daniel Balasubramanian, Yoav Etsion, Gabor Madl, Esteban Osses, Sameer Singh 0001, Karlkim Suwanmongkol, Minhui Xie, Stephen R. Schach
J. Syst. Softw.7
2007 Common coupling and pointer variables, with application to a Linux case study
Stephen R. Schach, Tokunbo O. S. Adeshiyan, Daniel Balasubramanian, Gabor Madl, Esteban Osses, Sameer Singh 0001, Karlkim Suwanmongkol, Minhui Xie, Dror G. Feitelson
Softw. Qual. J.6
2006 Transfer of Learning for Complex Task Domains: a Demonstration using Multiple Robots
abstract
This paper demonstrates a learning mechanism for complex tasks. Such tasks may be inherently expensive to learn in terms of training time and/or cost of obtaining each training pattern. Learning simple, safe tasks and extending them to more complex tasks can cause faster convergence to the solution. This method has been formalized and demonstrated on a simulated multiple robot (multi-robot) scenario. The objective is to effectively search out and destroy stationary hostile agents present in an unknown urban terrain map. Using the presented method, the robots learn how to effectively map the area, and then improve their learning modules for the complex task. The robots are simple behavioral agents with minimal communication
Sameer Singh 0001, Julie A. Adams
ICRA1