EDBT 2026 Demo / reviewers in the wild / expert
Hal Daumé III
dblp:77/2856
· DBLP profile ↗
167ranked-venue papers
24as first author
46since 2021 · last 2026
0000-0002-3760-345XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 148 · 23 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 since 2021Human-computer interaction and ubiquitous computing · 10 · 9 since 2021Databases, data management, data science and information retrieval · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Systems, architecture and hardware · 2Security and privacy · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal StyleabstractConnor Baumler, Calvin Bao, Huy Nghiem, Xinchen Yang, Marine Carpuat, Hal Daumé Iii. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Connor Baumler, Calvin Bao, Huy Nghiem, Xinchen Yang, Marine Carpuat, Hal Daumé III |
ACL (1) | 6 |
| 2026 | SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language ModelsabstractWARNING: This paper contains examples of offensive materials.To address toxic content on social media, we introduce SMARTER, a data-efficient 2-stage framework for explainable content moderation using Large Language Models (LLMs).In Stage 1, we leverage LLMs' own outputs to generate synthetic explanations for correct and incorrect labels, enabling preference optimization with minimal supervision.In Stage 2, we refine explanation quality through cross-model training, allowing weaker models to align with stronger ones.Experiments on 3 classification tasks (HateXplain, Latent Hate, Implicit Hate) show SMARTER achieves up to 13% macro-F1 improvement over few-shot baselines using only 6-57% of training data.Our framework offers a scalable strategy for low-data settings by harnessing LLMs' selfimprovement for explainable moderation. Huy Nghiem, Advik Sachdeva, Hal Daumé III |
ACL (1) | 3 |
| 2026 | When Stereotypes GTG: The Impact of Predictive Text Suggestions on Gender Bias in Human-AI Co-WritingabstractAI-based systems such as language models have been shown to replicate and even amplify social biases reflected in their training data. Among other questionable behaviors, this can lead to AI-generated text–and text suggestions–that contain normatively inappropriate stereotypical associations. Little is known, however, about how this behavior impacts the writing produced by people using these systems. We address this gap by measuring how much impact stereotypes or anti-stereotypes in English single-word LM predictive text suggestions have on the stories that people write using those tools in a co-writing scenario. We find that (n = 414), LM suggestions that challenge stereotypes sometimes lead to a significantly increased rate of anti-stereotypical co-written stories. However, despite this increased rate of anti-stereotypical stories, pro-stereotypical narratives still dominated the co-written stories, demonstrating that technical debiasing is only a partially effective strategy to alleviate harms from human-AI collaboration. Connor Baumler, Hal Daumé III |
CHI | 2 |
| 2026 | Surveilling Suitability: How AI Hiring Interviews Impact Job Seekers with DisabilitiesabstractAI hiring interviews, asynchronous video recording platforms that use AI to assess candidate suitability, are increasingly used by employers to streamline hiring processes. These platforms often promise to standardize assessments and mitigate subjective biases in hiring decisions. Yet, little is known about how these technologies are perceived and experienced by people with disabilities, a group historically underrepresented in the workforce and particularly vulnerable to injustices perpetuated by technology. To address this gap, we conducted focus groups and semi-structured interviews with 19 people with disabilities. We found that people with disabilities perceive and experience discrimination by AI hiring interviews that: 1) center normative characteristics, 2) exacerbate information asymmetries, 3) undermine autonomy, and 4) intrude on privacy. We use the analytical frame of surveillance to interrogate the role of AI in reconfiguring social relations between job seekers and employers. We discuss implications of our work for design and policy. Vaishnav Kameswaran, Valentina Hong, Jazmin Clark, Hal Daumé III, Katie Shilton |
CHI | 5 |
| 2026 | Say It My Way: Exploring Control in Conversational Visual Question Answering with Blind UsersabstractPrompting and steering techniques are well established in general-purpose generative AI, yet assistive visual question answering (VQA) tools for blind users still follow rigid interaction patterns with limited opportunities for customization. User control can be helpful when system responses are misaligned with their goals and contexts, a gap that becomes especially consequential for blind users that may rely on these systems for access. We invite 11 blind users to customize their interactions with a real-world conversational VQA system. Drawing on 418 interactions, reflections, and post-study interviews, we analyze prompting-based techniques participants adopted, including those introduced in the study and those developed independently in real-world settings. VQA interactions were often lengthy: participants averaged 3 turns, sometimes up to 21, with input text typically tenfold shorter than the responses they heard. Built on state-of-the-art LLMs, the system lacked verbosity controls, was limited in estimating distance in space and time, relied on inaccessible image framing, and offered little to no camera guidance. We discuss how customization techniques such as prompt engineering can help participants work around these limitations. Alongside a new publicly available dataset, we offer insights for interaction design at both query and system levels. Farnaz Zamiri Zeraati, Yang Trista Cao, Yuehan Qiao, Hal Daumé III, Hernisa Kacorri |
CHI | 4 |
| 2026 | Care, Wisdom, and Civics: Value Sensitive Design of Large Language Model Support for Online ModerationabstractHuman–computer interaction (HCI) and natural language processing (NLP) research increasingly explore large language model (LLM) support for online content moderation tasks. This study conducts a value sensitive design process with volunteer moderators of two heavily moderated subreddits (history Q&A, and legal advice). Through an empirical investigation using iterative interviews and a conceptual investigation centered in virtue ethics, we find moderators center values of care, wisdom, and civics in their work. A technical investigation then matches these values to the known capabilities and limitations of LLMs. We find current LLMs potentially well-suited to supporting care in managing sensitive content and wisdom in bridging content to context. However, many aspects of civic community-building are challenging to support with today’s models. Our study provides guidelines for designing LLM support for moderation tools and demonstrates value sensitive design methods to connect work practices, values, and the possibilities and limits of automation. Lovely-Frances Domingo, Sarah A. Gilbert, Yang (Trista) Cao, Hal Daumé III, Michelle L. Mazurek, Katie Shilton |
ACM Trans. Comput. Hum. Interact. | 4 |
| 2025 | Exploring Collaboration to Center the Deaf Community in Sign Language AIabstractSign language processing holds great promise for advancing societal inclusivity, yet it often excludes meaningful participation from the Deaf community, raising ethical and practical concerns about the applicability of AI solutions to their needs. This paper addresses these gaps through two interrelated studies. First, surveys identify differences in priorities and expectations between machine learning (ML) practitioners and Deaf American Sign Language (ASL) signers. Second, paired co-design sessions bring ML and ASL experts together to generate guiding questions that support practices for aligning AI development with community goals. Our findings reveal critical points of friction that reflect deeper systemic and epistemic barriers to effective collaboration. By synthesizing unique and shared insights from both groups, we provide empirically grounded resources to guide collaborative frameworks that promote the agency and expertise of the Deaf community. This research paves actionable pathways toward equitable, community-centered advancements in AI. Rie Kamikubo, Abraham Glasser, Alex Lu 0002, Hal Daumé III, Hernisa Kacorri, Danielle Bragg |
ASSETS | 4 |
| 2025 | An Interdisciplinary Approach to Human-Centered Machine TranslationabstractMarine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Fred Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé Iii, Kevin Duh, Ge Gao, Alvin C Grissom II, Marzena Karpinska, Elaine C Khoong, William D. Lewis, Andre Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Frédéric Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé III, Kevin Duh, Ge Gao 0001, Alvin Grissom II, Marzena Karpinska, Elaine C. Khoong, William D. Lewis, André F. T. Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon |
EMNLP | 8 |
| 2025 | 'Rich Dad, Poor Lad': How do Large Language Models Contextualize Socioeconomic Factors in College Admission ?abstractLarge Language Models (LLMs) are increasingly involved in high-stakes domains, yet how they reason about socially-sensitive decisions still remains underexplored.We present a largescale audit of LLMs' treatment of socioeconomic status (SES) in college admissions decisions using a novel dual-process framework inspired by cognitive science.Leveraging a synthetic dataset of 30,000 applicant profiles 1 grounded in real-world correlations, we prompt 4 open-source LLMs (Qwen 2, Mistral v0.3, Gemma 2, Llama 3.1) under 2 modes: a fast, decision-only setup (System 1) and a slower, explanation-based setup (System 2).Results from 5 million prompts reveals that LLMs consistently favor low-SES applicants-even when controlling for academic performance-and that System 2 amplifies this tendency by explicitly invoking SES as compensatory justification, highlighting both their potential and volatility as decision-makers.We then propose DPAF, a dual-process audit framework to probe LLMs' reasoning behaviors in sensitive applications. Huy Nghiem, Phuong-Anh Nguyen-Le, John Prindle, Rachel Rudinger, Hal Daumé III |
EMNLP | 5 |
| 2025 | A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text ExplanationsabstractFaithful free-text explanations are important to ensure transparency in high-stakes AI decisionmaking contexts, but they are challenging to generate by language models and assess by humans.In this paper, we present a measure for Prediction-EXplanation (PEX) consistency, by extending the concept of weight of evidence.This measure quantifies how much a free-text explanation supports or opposes a prediction, serving as an important aspect of explanation faithfulness.Our analysis reveals that more than 62% explanations generated by large language models lack this consistency.We show that applying direct preference optimization improves the consistency of generated explanations across three model families, with improvement ranging from 43.1% to 292.3%.Furthermore, we demonstrate that optimizing this consistency measure can improve explanation faithfulness by up to 9.7%. 1 Lingjun Zhao, Hal Daumé III |
EMNLP | 2 |
| 2025 | Which Demographic Features Are Relevant for Individual Fairness Evaluation of U.S. Recidivism Risk Assessment Tools?abstractDespite its constitutional relevance, the technical “individual fairness” criterion has not been operationalized in U.S. state or federal statutes/regulations. We conduct a human subjects experiment to address this gap, evaluating which demographic features are relevant for individual fairness evaluation of recidivism risk assessment (RRA) tools. Our analyses conclude that the individual similarity function should consider age and sex, but it should ignore race. Tin Trung Nguyen, Jiannan Xu, Phuong-Anh Nguyen-Le, Jonathan Lazar, Donald Braman, Hal Daumé III, Zubin Jelveh |
ICAIL | 6 |
| 2025 | Natural Language Inference Improves Compositionality in Vision-Language ModelsabstractCompositional reasoning in Vision-Language Models (VLMs) remains challenging as these models often struggle to relate objects, attributes, and spatial relationships. Recent methods aim to address these limitations by relying on the semantics of the textual description, using Large Language Models (LLMs) to break them down into subsets of questions and answers. However, these methods primarily operate on the surface level, failing to incorporate deeper lexical understanding while introducing incorrect assumptions generated by the LLM. In response to these issues, we present Caption Expansion with Contradictions and Entailments (CECE), a principled approach that leverages Natural Language Inference (NLI) to generate entailments and contradictions from a given premise. CECE produces lexically diverse sentences while maintaining their core meaning. Through extensive experiments, we show that CECE enhances interpretability and reduces overreliance on biased or superficial features. By balancing CECE along the original premise, we achieve significant improvements over previous methods without requiring additional fine-tuning, producing state-of-the-art results on benchmarks that score agreement with human judgments for image-text alignment, and achieving an increase in performance on Winoground of $+19.2\%$ (group score) and $+12.9\%$ on EqBen (group score) over the best prior work (finetuned with targeted data). Paola Cascante-Bonilla, Yang Trista Cao, Hal Daumé III, Rachel Rudinger |
ICLR | 4 |
| 2025 | TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic PoliciesabstractAlthough large vision-language-action (VLA) models pretrained on extensive robot datasets offer promising generalist policies for robotic learning, they still struggle with spatial-temporal dynamics in interactive robotics, making them less effective in handling complex tasks, such as manipulation. In this work, we introduce visual trace prompting, a simple yet effective approach to facilitate VLA models’ spatial-temporal awareness for action prediction by encoding state-action trajectories visually. We develop a new TraceVLA model by finetuning
OpenVLA on our own collected dataset of 150K robot manipulation trajectories using visual trace prompting. Evaluations of TraceVLA across 137 configurations in SimplerEnv and 4 tasks on a physical WidowX robot demonstrate state-of-the-art performance, outperforming OpenVLA by 10% on SimplerEnv and 3.5x on real-robot tasks and exhibiting robust generalization across diverse embodiments and scenarios. To further validate the effectiveness and generality of our method, we present a compact VLA model based on 4B Phi-3-Vision, pretrained on the Open-X-Embodiment and finetuned on our dataset, rivals the 7B OpenVLA baseline while significantly improving inference efficiency. Ruijie Zheng, Yongyuan Liang, Shuaiyi Huang, Jianfeng Gao 0001, Hal Daumé III, Andrey Kolobov, Furong Huang |
ICLR | 5 |
| 2025 | Language Models Predict Empathy Gaps Between Social In-groups and Out-groupsabstractYu Hou, Hal Daumé Iii, Rachel Rudinger. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hal Daumé III, Rachel Rudinger |
NAACL (Long Papers) | 2 |
| 2025 | My LLM might Mimic AAE - But When Should It?abstractSandra Camille Sandoval, Christabel Acquaye, Kwesi Adu Cobbina, Mohammad Nayeem Teli, Hal Daumé Iii. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sandra Sandoval, Christabel Acquaye, Kwesi A. Cobbina, Mohammad Nayeem Teli, Hal Daumé III |
NAACL (Long Papers) | 5 |
| 2025 | Causal Differentiating Concepts: Interpreting LM Behavior via Causal Representation LearningabstractLanguage model activations entangle concepts that mediate their behavior, making it difficult to interpret these factors, which has implications for generalizability and robustness. We introduce an approach for disentangling these concepts without supervision. Existing methods for concept discovery often rely on external labels, contrastive prompts, or known causal structures, which limits their scalability and biases them toward predefined, easily annotatable features. In contrast, we propose a new unsupervised algorithm that identifies causal differentiating concepts—interpretable latent directions in LM activations that must be changed to elicit a different model behavior. These concepts are discovered using a constrained contrastive learning objective, guided by the insight that eliciting a target behavior requires only sparse changes to the underlying concepts. We formalize this notion and show that, under a particular assumption about the sparsity of these causal differentiating concepts, our method learns disentangled representations that align with human-interpretable factors influencing LM decisions. We empirically show the ability of our method to recover ground-truth causal factors in synthetic and semi-synthetic settings. Additionally, we illustrate the utility of our method through a case study on refusal behavior in language models. Our approach offers a scalable and interpretable lens into the internal workings of LMs, providing a principled foundation for interpreting language model behavior. Navita Goyal, Hal Daumé III, Alexandre Drouin, Dhanya Sridhar |
NeurIPS | 2 |
| 2024 | Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric MethodabstractExtensive efforts in automated approaches for content moderation have been focused on developing models to identify toxic, offensive, and hateful content with the aim of lightening the load for moderators.Yet, it remains uncertain whether improvements on those tasks have truly addressed moderators' needs in accomplishing their work.In this paper, we surface gaps between past research efforts that have aimed to provide automation for aspects of content moderation and the needs of volunteer content moderators, regarding identifying violations of various moderation rules.To do so, we conduct a model review on Hugging Face to reveal the availability of models to cover various moderation rules and guidelines from three exemplar forums.We further put state-of-the-art LLMs to the test, evaluating how well these models perform in flagging violations of platform rules from one particular forum.Finally, we conduct a user survey study with volunteer moderators to gain insight into their perspectives on useful moderation models.Overall, we observe a nontrivial gap, as missing developed models and LLMs exhibit moderate to low performance on a significant portion of the rules.Moderators' reports provide guides for future work on developing moderation assistant models. Yang Trista Cao, Lovely-Frances Domingo, Sarah A. Gilbert, Michelle L. Mazurek, Katie Shilton, Hal Daumé III |
EMNLP | 6 |
| 2024 | Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRAabstractRecent advancements of large language models (LLMs) have led to claims of AI surpassing humans in natural language processing (NLP) tasks such as textual understanding and reasoning.This work investigates these assertions by introducing CAIMIRA, a novel framework rooted in item response theory (IRT) that enables quantitative assessment and comparison of problem-solving abilities in questionanswering (QA) agents.Through analysis of over 300,000 responses from ~70 AI systems and 155 humans across thousands of quiz questions, CAIMIRA uncovers distinct proficiency patterns in knowledge domains and reasoning skills.Humans outperform AI systems in knowledge-grounded abductive and conceptual reasoning, while state-of-the-art LLMs like GPT-4-TURBO and LLAMA-3-70B demonstrate superior performance on targeted information retrieval and fact-based reasoning, particularly when information gaps are well-defined and addressable through pattern matching or data retrieval.These findings identify key areas for future QA tasks and model development, highlighting the critical need for questions that not only challenge higher-order reasoning and scientific thinking, but also demand nuanced linguistic and cross-contextual application. Maharshi Gor, Hal Daumé III, Tianyi Zhou 0001, Jordan L. Boyd-Graber |
EMNLP | 2 |
| 2024 | "You Gotta be a Doctor, Lin" : An Investigation of Name-Based Bias of Large Language Models in Employment RecommendationsabstractSocial science research has shown that candidates with names indicative of certain races or genders often face discrimination in employment practices.Similarly, Large Language Models (LLMs) have demonstrated racial and gender biases in various applications.In this study, we utilize GPT-3.5-Turbo and Llama 3-70B-Instruct to simulate hiring decisions and salary recommendations for candidates with 320 first names that strongly signal their race and gender, across over 750,000 prompts.Our empirical results indicate a preference among these models for hiring candidates with White female-sounding names over other demographic groups across 40 occupations.Additionally, even among candidates with identical qualifications, salary recommendations vary by as much as 5% between different subgroups.A comparison with real-world labor data reveals inconsistent alignment with U.S. labor market characteristics, underscoring the necessity of risk investigation of LLM-powered systems. Huy Nghiem, John Prindle, Jieyu Zhao 0001, Hal Daumé III |
EMNLP | 4 |
| 2024 | ASL STEM Wiki: Dataset and Benchmark for Interpreting STEM ArticlesabstractKayo Yin, Chinmay Singh, Fyodor O Minakov, Vanessa Milan, Hal Daumé Iii, Cyril Zhang, Alex Xijie Lu, Danielle Bragg. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Kayo Yin, Chinmay Singh, Fyodor O. Minakov, Vanessa Milan, Hal Daumé III, Cyril Zhang, Alex Lu 0002, Danielle Bragg |
EMNLP | 5 |
| 2024 | Successfully Guiding Humans with Imperfect Instructions by Highlighting Potential Errors and Suggesting CorrectionsabstractLanguage models will inevitably err in situations with which they are unfamiliar.However, by effectively communicating uncertainties, they can still guide humans toward making sound decisions in those contexts.We demonstrate this idea by developing HEAR, a system that can successfully guide humans in simulated residential environments despite generating potentially inaccurate instructions.Diverging from systems that provide users with only the instructions they generate, HEAR warns users of potential errors in its instructions and suggests corrections.This rich uncertainty information effectively prevents misguidance and reduces the search space for users.Evaluation with 80 users shows that HEAR achieves a 13% increase in success rate and a 29% reduction in final location error distance compared to only presenting instructions to users.Interestingly, we find that offering users possibilities to explore, HEAR motivates them to make more attempts at the task, ultimately leading to a higher success rate.To our best knowledge, this work is the first to show the practical benefits of uncertainty communication in a long-horizon sequential decision-making problem. 1 Lingjun Zhao, Hal Daumé III |
EMNLP | 3 |
| 2024 | DrM: Mastering Visual Reinforcement Learning through Dormant Ratio MinimizationabstractVisual reinforcement learning (RL) has shown promise in continuous control tasks.
Despite its progress, current algorithms are still unsatisfactory in virtually every aspect of the performance such as sample efficiency, asymptotic performance, and their robustness to the choice of random seeds.
In this paper, we identify a major shortcoming in existing visual RL methods that is the agents often exhibit sustained inactivity during early training, thereby limiting their ability to explore effectively.
Expanding upon this crucial observation, we additionally unveil a significant correlation between the agents' inclination towards motorically inactive exploration and the absence of neuronal activity within their policy networks.
To quantify this inactivity, we adopt dormant ratio as a metric to measure inactivity in the RL agent's network.
Empirically, we also recognize that the dormant ratio can act as a standalone indicator of an agent's activity level, regardless of the received reward signals.
Leveraging the aforementioned insights, we introduce DrM, a method that uses three core mechanisms to guide agents' exploration-exploitation trade-offs by actively minimizing the dormant ratio.
Experiments demonstrate that DrM achieves significant improvements in sample efficiency and asymptotic performance with no broken seeds (76 seeds in total) across three continuous control benchmark environments, including DeepMind Control Suite, MetaWorld, and Adroit.
Most importantly, DrM is the first model-free algorithm that consistently solves tasks in both the Dog and Manipulator domains from the DeepMind Control Suite as well as three dexterous hand manipulation tasks without demonstrations in Adroit, all based on pixel observations. Guowei Xu 0001, Ruijie Zheng, Yongyuan Liang, Zhecheng Yuan, Tianying Ji, Yu Luo 0021, Xiaoyu Liu 0003, Pu Hua, Shuzhen Li, Yanjie Ze, Hal Daumé III, Furong Huang, Huazhe Xu |
ICLR | 13 |
| 2024 | PRISE: LLM-Style Sequence Compression for Learning Temporal Action Abstractions in ControlabstractTemporal action abstractions, along with belief state representations, are a powerful knowledge sharing mechanism for sequential decision making. In this work, we propose a novel view that treats inducing temporal action abstractions as a sequence compression problem. To do so, we bring a subtle but critical component of LLM training pipelines -- input tokenization via byte pair encoding (BPE) -- to bear on the seemingly distant task of learning skills of variable time span in continuous control domains. We introduce an approach called Primitive Sequence Encoding (PRISE) that combines continuous action quantization with BPE to learn powerful action abstractions. We empirically show that high-level skills discovered by PRISE from a multitask set of robotic manipulation demonstrations significantly boost the learning performance of behavior cloning on downstream tasks. Ruijie Zheng, Ching-An Cheng, Hal Daumé III, Furong Huang, Andrey Kolobov |
ICML | 3 |
| 2024 | Premier-TACO is a Few-Shot Policy Learner: Pretraining Multitask Representation via Temporal Action-Driven Contrastive LossabstractWe present Premier-TACO, a multitask feature representation learning approach designed to improve few-shot policy learning efficiency in sequential decision-making tasks. Premier-TACO leverages a subset of multitask offline datasets for pretraining a general feature representation, which captures critical environmental dynamics and is fine-tuned using minimal expert demonstrations. It advances the temporal action contrastive learning (TACO) objective, known for state-of-the-art results in visual control tasks, by incorporating a novel negative example sampling strategy. This strategy is crucial in significantly boosting TACO’s computational efficiency, making large-scale multitask offline pretraining feasible. Our extensive empirical evaluation in a diverse set of continuous control benchmarks including Deepmind Control Suite, MetaWorld, and LIBERO demonstrate Premier-TACO’s effective- ness in pretraining visual representations, significantly enhancing few-shot imitation learning of novel tasks. Ruijie Zheng, Yongyuan Liang, Hal Daumé III, Huazhe Xu, John Langford 0001, Praveen Palanisamy, Kalyan Shankar Basu, Furong Huang |
ICML | 5 |
| 2024 | The Impact of Explanations on Fairness in Human-AI Decision-Making: Protected vs Proxy FeaturesabstractAI systems have been known to amplify biases in real-world data. Explanations may help human-AI teams address these biases for fairer decision-making. Typically, explanations focus on salient input features. If a model is biased against some protected group, explanations may include features that demonstrate this bias, but when biases are realized through proxy features, the relationship between this proxy feature and the protected one may be less clear to a human. In this work, we study the effect of the presence of protected and proxy features on participants’ perception of model fairness and their ability to improve demographic parity over an AI alone. Further, we examine how different treatments—explanations, model bias disclosure and proxy correlation disclosure—affect fairness perception and parity. We find that explanations help people detect direct but not indirect biases. Additionally, regardless of bias type, explanations tend to increase agreement with model biases. Disclosures can help mitigate this effect for indirect biases, improving both unfairness recognition and decision-making fairness. We hope that our findings can help guide further research into advancing explanations in support of fair human-AI decision-making. Navita Goyal, Connor Baumler, Tin Nguyen 0005, Hal Daumé III |
IUI | 4 |
| 2024 | Large Language Models Help Humans Verify Truthfulness - Except When They Are Convincingly WrongabstractChenglei Si, Navita Goyal, Tongshuang Wu, Chen Zhao, Shi Feng, Hal Daumé Iii, Jordan Boyd-Graber. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chenglei Si, Navita Goyal, Sherry Tongshuang Wu, Chen Zhao 0013, Shi Feng 0005, Hal Daumé III, Jordan L. Boyd-Graber |
NAACL-HLT | 6 |
| 2024 | Seamful XAI: Operationalizing Seamful Design in Explainable AIabstractMistakes in AI systems are inevitable, arising from both technical limitations and sociotechnical gaps. While black-boxing AI systems can make the user experience seamless, hiding the seams risks disempowering users to mitigate fallouts from AI mistakes. Instead of hiding these AI imperfections, can we leverage them to help the user? While Explainable AI (XAI) has predominantly tackled algorithmic opaqueness, we propose that seamful design can foster AI explainability by revealing and leveraging sociotechnical and infrastructural mismatches. We introduce the concept of Seamful XAI by (1) conceptually transferring "seams" to the AI context and (2) developing a design process that helps stakeholders anticipate and design with seams. We explore this process with 43 AI practitioners and real end-users, using a scenario-based co-design activity informed by real-world use cases. We found that the Seamful XAI design process helped users foresee AI harms, identify underlying reasons (seams), locate them in the AI's lifecycle, learn how to leverage seamful information to improve XAI and user agency. We share empirical insights, implications, and reflections on how this process can help practitioners anticipate and craft seams in AI, how seamfulness can improve explainability, empower end-users, and facilitate Responsible AI. Upol Ehsan, Qingzi Vera Liao, Samir Passi, Mark O. Riedl, Hal Daumé III |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2023 | FairPrism: Evaluating Fairness-Related Harms in Text GenerationabstractEve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett, Hal Daumé III, Alexandra Olteanu, Emily Sheng, Dan Vann, Hanna Wallach. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Eve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett, Hal Daumé III, Alexandra Olteanu, Emily Sheng, Dan Vann, Hanna M. Wallach |
ACL (1) | 5 |
| 2023 | Factual or Contextual? Disentangling Error Types in Entity Description GenerationabstractIn the task of entity description generation, given a context and a specified entity, a model must describe that entity correctly and in a contextually-relevant way.In this task, as well as broader language generation tasks, the generation of a nonfactual description (factual error) versus an incongruous description (contextual error) is fundamentally different, yet often conflated.We develop an evaluation paradigm that enables us to disentangle these two types of errors in naturally occurring textual contexts.We find that factuality and congruity are often at odds, and that models specifically struggle with accurate descriptions of entities that are less familiar to people.This shortcoming of language models raises concerns around the trustworthiness of such models, since factual errors on less well-known entities are exactly those that a human reader will not recognize.1 1 The code and data used in the paper is available at https: //github.com/navitagoyal/Factual-or-Contextual-E rrors-in-LM-Desc-Gen. Navita Goyal, Ani Nenkova, Hal Daumé III |
ACL (1) | 3 |
| 2023 | Towards Conceptualization of "Fair Explanation": Disparate Impacts of anti-Asian Hate Speech Explanations on Content ModeratorsabstractRecent research at the intersection of AI explainability and fairness has focused on how explanations can improve human-plus-AI task performance as assessed by fairness measures.We propose to characterize what constitutes an explanation that is itself "fair" -an explanation that does not adversely impact specific populations.We formulate a novel evaluation method of "fair explanations" using not just accuracy and label time, but also psychological impact of explanations on different user groups across many metrics (mental discomfort, stereotype activation, and perceived workload).We apply this method in the context of content moderation of potential hate speech, and its differential impact on Asian vs. non-Asian proxy moderators, across explanation approaches (saliency map and counterfactual explanation).We find that saliency maps generally perform better and show less evidence of disparate impact (group) and individual unfairness than counterfactual explanations.1 Content warning: This paper contains examples of hate speech and racially discriminatory language.The authors do not support such content.Please consider your risk of discomfort carefully before continuing reading! Tin Nguyen 0005, Jiannan Xu, Aayushi Roy, Hal Daumé III, Marine Carpuat |
EMNLP | 4 |
| 2023 | What Else Do I Need to Know? The Effect of Background Information on Users' Reliance on QA SystemsabstractNavita Goyal, Eleftheria Briakou, Amanda Liu, Connor Baumler, Claire Bonial, Jeffrey Micher, Clare Voss, Marine Carpuat, Hal Daumé III. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Navita Goyal, Eleftheria Briakou, Amanda Liu, Connor Baumler, Claire Bonial, Jeffrey Micher, Clare R. Voss, Marine Carpuat, Hal Daumé III |
EMNLP | 9 |
| 2023 | A Rose by Any Other Name would not Smell as Sweet: Social Bias in Names MistranslationabstractWe ask the question: Are there widespread disparities in machine translations of names across race/ethnicity, and gender?We hypothesize that the translation quality of names and surrounding context will be lower for names associated with US racial and ethnic minorities due to these systems' tendencies to standardize language to predominant language patterns.We develop a dataset of names that are strongly demographically aligned and propose a translation evaluation procedure based on round-trip translation.We analyze the effect of name demographics on translation quality using generalized linear mixed effects models and find that the ability of translation systems to correctly translate female-associated names is significantly lower than male-associated names.This effect is particularly pronounced for femaleassociated names that are also associated with racial (Black) and ethnic (Hispanic) minorities.This disparity in translation quality between social groups for something as personal as someone's name has significant implications for people's professional, personal and cultural identities, self-worth and ease of communication.Our findings suggest that more MT research is needed to improve the translation of names and to provide high-quality service for users regardless of gender, race, and ethnicity. Sandra Sandoval, Jieyu Zhao 0001, Marine Carpuat, Hal Daumé III |
EMNLP | 4 |
| 2023 | ASL Citizen: A Community-Sourced Dataset for Advancing Isolated Sign Language RecognitionabstractSign languages are used as a primary language by approximately 70 million D/deaf people world-wide. However, most communication technologies operate in spoken and written languages, creating inequities in access. To help tackle this problem, we release ASL Citizen, the first crowdsourced Isolated Sign Language Recognition (ISLR) dataset, collected with consent and containing 83,399 videos for 2,731 distinct signs filmed by 52 signers in a variety of environments. We propose that this dataset be used for sign language dictionary retrieval for American Sign Language (ASL), where a user demonstrates a sign to their webcam to retrieve matching signs from a dictionary. We show that training supervised machine learning classifiers with our dataset advances the state-of-the-art on metrics relevant for dictionary retrieval, achieving 63\% accuracy and a recall-at-10 of 91\%, evaluated entirely on videos of users who are not present in the training or validation sets. Aashaka Desai, Lauren Berger, Fyodor O. Minakov, Nessa Milano, Chinmay Singh, Kriston Pumphrey, Richard E. Ladner, Hal Daumé III, Alex Lu 0002, Naomi Caselli, Danielle Bragg |
NeurIPS | 8 |
| 2023 | TACO: Temporal Latent Action-Driven Contrastive Loss for Visual Reinforcement Learning
Ruijie Zheng, Yanchao Sun, Jieyu Zhao 0001, Huazhe Xu, Hal Daumé III, Furong Huang |
NeurIPS | 7 |
| 2022 | A Framework for Learning to Request Rich and Contextually Useful Information from HumansabstractWhen deployed, AI agents will encounter problems that are beyond their autonomous problem-solving capabilities. Leveraging human assistance can help agents overcome their inherent limitations and robustly cope with unfamiliar situations. We present a general interactive framework that enables an agent to request and interpret rich, contextually useful information from an assistant that has knowledge about the task and the environment. We demonstrate the practicality of our framework on a simulated human-assisted navigation problem. Aided with an assistance-requesting policy learned by our method, a navigation agent achieves up to a 7{\texttimes} improvement in success rate on tasks that take place in previously unseen environments, compared to fully autonomous behavior. We show that the agent can take advantage of different types of information depending on the context, and analyze the benefits and challenges of learning the assistance-requesting policy when the assistant can recursively decompose tasks into subtasks. Khanh X. Nguyen, Yonatan Bisk, Hal Daumé III |
ICML | 3 |
| 2022 | Theory-Grounded Measurement of U.S. Social Stereotypes in English Language ModelsabstractYang Cao, Anna Sotnikova, Hal Daumé III, Rachel Rudinger, Linda Zou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yang Trista Cao, Anna Sotnikova, Hal Daumé III, Rachel Rudinger, Linda Zou |
NAACL-HLT | 3 |
| 2022 | Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their ImplicationsabstractKaitlyn Zhou, Su Lin Blodgett, Adam Trischler, Hal Daumé III, Kaheer Suleman, Alexandra Olteanu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Kaitlyn Zhou, Su Lin Blodgett, Adam Trischler, Hal Daumé III, Kaheer Suleman, Alexandra Olteanu |
NAACL-HLT | 4 |
| 2022 | Spoken language interaction with robots: Recommendations for future researchabstractWith robotics rapidly advancing, more effective human–robot interaction is increasingly needed to realize the full potential of robots for society. While spoken language must be part of the solution, our ability to provide spoken language interaction capabilities is still very limited. In this article, based on the report of an interdisciplinary workshop convened by the National Science Foundation, we identify key scientific and engineering advances needed to enable effective spoken language interaction with robotics. We make 25 recommendations, involving eight general themes: putting human needs first, better modeling the social and interactive aspects of language, improving robustness, creating new methods for rapid adaptation, better integrating speech and language with other communication modalities, giving speech and language components access to rich representations of the robot’s current knowledge and state, making all components operate in real time, and improving research infrastructure and resources. Research and development that prioritizes these topics will, we believe, provide a solid foundation for the creation of speech-capable robots that are easy and effective for humans to work with. Matthew Marge, Carol Y. Espy-Wilson, Nigel G. Ward, Abeer Alwan, Yoav Artzi, Mohit Bansal, Gilmer L. Blankenship, Joyce Y. Chai, Hal Daumé III, Debadeepta Dey, Mary P. Harper, Thomas Howard, Casey Kennington, Ivana Kruijff-Korbayová, Dinesh Manocha, Cynthia Matuszek, Ross Mead, Raymond J. Mooney, Roger K. Moore, Mari Ostendorf, Heather Pon-Barry, Alexander I. Rudnicky, Matthias Scheutz, Robert St. Amant, Stefanie Tellex, David R. Traum, Zhou Yu 0005 |
Comput. Speech Lang. | 9 |
| 2022 | Heterogeneous Supervised Topic ModelsabstractAbstract Researchers in the social sciences are often interested in the relationship between text and an outcome of interest, where the goal is to both uncover latent patterns in the text and predict outcomes for unseen texts. To this end, this paper develops the heterogeneous supervised topic model (HSTM), a probabilistic approach to text analysis and prediction. HSTMs posit a joint model of text and outcomes to find heterogeneous patterns that help with both text analysis and prediction. The main benefit of HSTMs is that they capture heterogeneity in the relationship between text and the outcome across latent topics. To fit HSTMs, we develop a variational inference algorithm based on the auto-encoding variational Bayes framework. We study the performance of HSTMs on eight datasets and find that they consistently outperform related methods, including fine-tuned black-box models. Finally, we apply HSTMs to analyze news articles labeled with pro- or anti-tone. We find evidence of differing language used to signal a pro- and anti-tone. Dhanya Sridhar, Hal Daumé III, David M. Blei |
Trans. Assoc. Comput. Linguistics | 2 |
| 2021 | Meta-Learning Effective Exploration Strategies for Contextual Bandits
Amr Sharaf, Hal Daumé III |
AAAI | 2 |
| 2021 | A Novice-Reviewer Experiment to Address Scarcity of Qualified Reviewers in Large ConferencesabstractConference peer review constitutes a human-computation process whose importance cannot be overstated: not only it identifies the best submissions for acceptance, but, ultimately, it impacts the future of the whole research area by promoting some ideas and restraining others. A surge in the number of submissions received by leading AI conferences has challenged the sustainability of the review process by increasing the burden on the pool of qualified reviewers which is growing at a much slower rate. In this work, we consider the problem of reviewer recruiting with a focus on the scarcity of qualified reviewers in large conferences. Specifically, we design a procedure for (i) recruiting reviewers from the population not typically covered by major conferences and (ii) guiding them through the reviewing pipeline. In conjunction with the ICML 2020 --- a large, top-tier machine learning conference --- we recruit a small set of reviewers through our procedure and compare their performance with the general population of ICML reviewers. Our experiment reveals that a combination of the recruiting and guiding mechanisms allows for a principled enhancement of the reviewer pool and results in reviews of superior quality compared to the conventional pool of reviews as evaluated by senior members of the program committee (meta-reviewers). Ivan Stelmakh, Nihar B. Shah, Aarti Singh, Hal Daumé III |
AAAI | 4 |
| 2021 | Distantly-Supervised Dense Retrieval Enables Open-Domain Question Answering without Evidence AnnotationabstractOpen-domain question answering answers a question based on evidence retrieved from a large corpus.State-of-the-art neural approaches require intermediate evidence annotations for training.However, such intermediate annotations are expensive, and methods that rely on them cannot transfer to the more common setting, where only questionanswer pairs are available.This paper investigates whether models can learn to find evidence from a large corpus, with only distant supervision from answer labels for model training, thereby generating no additional annotation cost.We introduce a novel approach (DISTDR) that iteratively improves over a weak retriever by alternately finding evidence from the up-to-date model and encouraging the model to learn the most likely evidence.Without using any evidence labels, DISTDR is on par with fully-supervised state-of-theart methods on both multi-hop and singlehop QA benchmarks.Our analysis confirms that DISTDR finds more accurate evidence over iterations, which leads to model improvements.The code is available at https:// github.com/henryzhao5852/DistDR. Chen Zhao 0013, Chenyan Xiong, Jordan L. Boyd-Graber, Hal Daumé III |
EMNLP (1) | 4 |
| 2021 | From Human Explanation to Model Interpretability: A Framework Based on Weight of EvidenceabstractWe take inspiration from the study of human explanation to inform the design and evaluation of interpretability methods in machine learning. First, we survey the literature on human explanation in philosophy, cognitive science, and the social sciences, and propose a list of design principles for machine-generated explanations that are meaningful to humans. Using the concept of weight of evidence from information theory, we develop a method for generating explanations that adhere to these principles. We show that this method can be adapted to handle high-dimensional, multi-class settings, yielding a flexible framework for generating explanations. We demonstrate that these explanations can be estimated accurately from finite samples and are robust to small perturbations of the inputs. We also evaluate our method through a qualitative user study with machine learning practitioners, where we observe that the resulting explanations are usable despite some participants struggling with background concepts like prior class probabilities. Finally, we conclude by surfacing design implications for interpretability tools in general. David Alvarez-Melis, Harmanpreet Kaur, Hal Daumé III, Hanna M. Wallach, Jennifer Wortman Vaughan |
HCOMP | 3 |
| 2021 | Multi-Step Reasoning Over Unstructured Text with Beam Dense RetrievalabstractChen Zhao, Chenyan Xiong, Jordan Boyd-Graber, Hal Daumé III. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Chen Zhao 0013, Chenyan Xiong, Jordan L. Boyd-Graber, Hal Daumé III |
NAACL-HLT | 4 |
| 2021 | Toward Gender-Inclusive Coreference Resolution: An Analysis of Gender and Bias Throughout the Machine Learning LifecycleabstractAbstract Correctly resolving textual mentions of people fundamentally entails making inferences about those people. Such inferences raise the risk of systematic biases in coreference resolution systems, including biases that can harm binary and non-binary trans and cis stakeholders. To better understand such biases, we foreground nuanced conceptualizations of gender from sociology and sociolinguistics, and investigate where in the machine learning pipeline such biases can enter a coreference resolution system. We inspect many existing data sets for trans-exclusionary biases, and develop two new data sets for interrogating bias in both crowd annotations and in existing coreference resolution systems. Through these studies, conducted on English text, we confirm that without acknowledging and building systems that recognize the complexity of gender, we will build systems that fail for: quality of service, stereotyping, and over- or under-representation, especially for binary and non-binary trans users. Yang Trista Cao, Hal Daumé III |
Comput. Linguistics | 2 |
| 2021 | Prior and Prejudice: The Novice Reviewers' Bias against Resubmissions in Conference Peer ReviewabstractModern machine learning and computer science conferences are experiencing a surge in the number of submissions that challenges the quality of peer review as the number of competent reviewers is growing at a much slower rate. To curb this trend and reduce the burden on reviewers, several conferences have started encouraging or even requiring authors to declare the previous submission history of their papers. Such initiatives have been met with skepticism among authors, who raise the concern about a potential bias in reviewers' recommendations induced by this information. In this work, we investigate whether reviewers exhibit a bias caused by the knowledge that the submission under review was previously rejected at a similar venue, focusing on a population of novice reviewers who constitute a large fraction of the reviewer pool in leading machine learning and computer science conferences. We design and conduct a randomized controlled trial closely replicating the relevant components of the peer-review pipeline with $133$ reviewers (master's, junior PhD students, and recent graduates of top US universities) writing reviews for $19$ papers. The analysis reveals that reviewers indeed become negatively biased when they receive a signal about paper being a resubmission, giving almost 1 point lower overall score on a 10-point Likert item (Δ = -0.78, 95% CI = [-1.30, -0.24]) than reviewers who do not receive such a signal. Looking at specific criteria scores (originality, quality, clarity and significance), we observe that novice reviewers tend to underrate quality the most. Ivan Stelmakh, Nihar B. Shah, Aarti Singh, Hal Daumé III |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2020 | Language (Technology) is Power: A Critical Survey of "Bias" in NLPabstractWe survey 146 papers analyzing "bias" in NLP systems, fnding that their motivations are often vague, inconsistent, and lacking in normative reasoning, despite the fact that analyzing "bias" is an inherently normative process.We further fnd that these papers' proposed quantitative techniques for measuring or mitigating "bias" are poorly matched to their motivations and do not engage with the relevant literature outside of NLP.Based on these fndings, we describe the beginnings of a path forward by proposing three recommendations that should guide work analyzing "bias" in NLP systems.These recommendations rest on a greater recognition of the relationships between language and social hierarchies, encouraging researchers and practitioners to articulate their conceptualizations of "bias"-i.e., what kinds of system behaviors are harmful, in what ways, to whom, and why, as well as the normative reasoning underlying these statements-and to center work around the lived experiences of members of communities affected by NLP systems, while interrogating and reimagining the power relations between technologists and such communities.NLP task Papers Embeddings (type-level or contextualized) 54 Coreference resolution 20 Language modeling or dialogue generation 17 Hate-speech detection 17 Sentiment analysis 15 Machine translation 8 Tagging or parsing 5 Surveys, frameworks, and meta-analyses 20 Other 22 Su Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. Wallach |
ACL | 3 |
| 2020 | Active Imitation Learning with Noisy GuidanceabstractImitation learning algorithms provide state-ofthe-art results on many structured prediction tasks by learning near-optimal search policies.Such algorithms assume training-time access to an expert that can provide the optimal action at any queried state; unfortunately, the number of such queries is often prohibitive, frequently rendering these approaches impractical.To combat this query complexity, we consider an active learning setting in which the learning algorithm has additional access to a much cheaper noisy heuristic that provides noisy guidance.Our algorithm, LEAQI, learns a difference classifier that predicts when the expert is likely to disagree with the heuristic, and queries the expert only when necessary.We apply LEAQI to three sequence labeling tasks, demonstrating significantly fewer queries to the expert and comparable (or better) accuracies over a passive approach. Kianté Brantley, Hal Daumé III, Amr Sharaf |
ACL | 2 |
| 2020 | Toward Gender-Inclusive Coreference ResolutionabstractCorrectly resolving textual mentions of people fundamentally entails making inferences about those people.Such inferences raise the risk of systemic biases in coreference resolution systems, including biases that can harm binary and non-binary trans and cis stakeholders.To better understand such biases, we foreground nuanced conceptualizations of gender from sociology and sociolinguistics, and develop two new datasets for interrogating bias in crowd annotations and in existing coreference resolution systems.Through these studies, conducted on English text, we confirm that without acknowledging and building systems that recognize the complexity of gender, we build systems that lead to many potential harms. Yang Trista Cao, Hal Daumé III |
ACL | 2 |
| 2020 | Operationalizing the Legal Principle of Data Minimization for PersonalizationabstractArticle 5(1)(c) of the European Union's General Data Protection Regulation (GDPR) requires that "personal data shall be [...] adequate, relevant, and limited to what is necessary in relation to the purposes for which they are processed (`data minimisation')". To date, the legal and computational definitions of 'purpose limitation' and 'data minimization' remain largely unclear. In particular, the interpretation of these principles is an open issue for information access systems that optimize for user experience through personalization and do not strictly require personal data collection for the delivery of basic service. Asia J. Biega, Peter Potash, Hal Daumé III, Fernando Diaz 0001, Michèle Finck |
SIGIR | 3 |
| 2019 | Improving Fairness in Machine Learning Systems: What Do Industry Practitioners Need?abstractThe potential for machine learning (ML) systems to amplify social inequities and unfairness is receiving increasing popular and academic attention. A surge of recent work has focused on the development of algorithmic tools to assess and mitigate such unfairness. If these tools are to have a positive impact on industry practice, however, it is crucial that their design be informed by an understanding of real-world needs. Through 35 semi-structured interviews and an anonymous survey of 267 ML practitioners, we conduct the first systematic investigation of commercial product teams' challenges and needs for support in developing fairer ML systems. We identify areas of alignment and disconnect between the challenges faced by teams in practice and the solutions proposed in the fair ML research literature. Based on these findings, we highlight directions for future ML and HCI research that will better address practitioners' needs. Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miroslav Dudík, Hanna M. Wallach |
CHI | 3 |
| 2019 | Help, Anna! Visual Navigation with Natural Multimodal Assistance via Retrospective Curiosity-Encouraging Imitation LearningabstractKhanh Nguyen, Hal Daumé III. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Hal Daumé III |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Comparing and Developing Tools to Measure the Readability of Domain-Specific TextsabstractElissa Redmiles, Lisa Maszkiewicz, Emily Hwang, Dhruv Kuchhal, Everest Liu, Miraida Morales, Denis Peskov, Sudha Rao, Rock Stevens, Kristina Gligorić, Sean Kross, Michelle Mazurek, Hal Daumé III. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Elissa M. Redmiles, Lisa N. Maszkiewicz, Emily Hwang, Dhruv Kuchhal, Everest Liu, Miraida Morales, Denis Peskov, Sudha Rao, Rock Stevens, Kristina Gligoric, Sean Kross, Michelle L. Mazurek, Hal Daumé III |
EMNLP/IJCNLP (1) | 13 |
| 2019 | Contextual Memory TreesabstractWe design and study a Contextual Memory Tree (CMT), a learning memory controller that inserts new memories into an experience store of unbounded size. It operates online and is designed to efficiently query for memories from that store, supporting logarithmic time insertion and retrieval operations. Hence CMT can be integrated into existing statistical learning algorithms as an augmented memory unit without substantially increasing training and inference computation. Furthermore CMT operates as a reduction to classification, allowing it to benefit from advances in representation or architecture. We demonstrate the efficacy of CMT by augmenting existing multi-class and multi-label classification algorithms with CMT and observe statistical improvement. We also test CMT learning on several image-captioning tasks to demonstrate that it performs computationally better than a simple nearest neighbors memory system while benefitting from reward learning. Wen Sun 0002, Alina Beygelzimer, Hal Daumé III, John Langford 0001, Paul Mineiro |
ICML | 3 |
| 2019 | Non-Monotonic Sequential Text GenerationabstractStandard sequential generation methods assume a pre-specified generation order, such as text generation methods which generate words from left to right. In this work, we propose a framework for training models of text generation that operate in non-monotonic orders; the model directly learns good orders, without any additional annotation. Our framework operates by generating a word at an arbitrary position, and then recursively generating words to its left and then words to its right, yielding a binary tree. Learning is framed as imitation learning, including a coaching method which moves from imitating an oracle to reinforcing the policy’s own preferences. Experimental results demonstrate that using the proposed method, it is possible to learn policies which generate text without pre-specifying a generation order, while achieving competitive performance with conventional left-to-right generation. Sean Welleck, Kianté Brantley, Hal Daumé III, Kyunghyun Cho |
ICML | 3 |
| 2019 | Warm-starting Contextual Bandits: Robustly Combining Supervised and Bandit FeedbackabstractWe investigate the feasibility of learning from both fully-labeled supervised data and contextual bandit data. We specifically consider settings in which the underlying learning signal may be different between these two data sources. Theoretically, we state and prove no-regret algorithms for learning that is robust to divergences between the two sources. Empirically, we evaluate some of these algorithms on a large selection of datasets, showing that our approaches are feasible, and helpful in practice. Chicheng Zhang, Alekh Agarwal, Hal Daumé III, John Langford 0001, Sahand Negahban |
ICML | 3 |
| 2019 | Reinforcement Learning with Convex ConstraintsabstractIn standard reinforcement learning (RL), a learning agent seeks to optimize the overall reward. However, many key aspects of a desired behavior are more naturally expressed as constraints. For instance, the designer may want to limit the use of unsafe actions, increase the diversity of trajectories to enable exploration, or approximate expert trajectories when rewards are sparse. In this paper, we propose an algorithmic scheme that can handle a wide class of constraints in RL tasks: specifically, any constraints that require expected values of some vector measurements (such as the use of an action) to lie in a convex set. This captures previously studied constraints (such as safety and proximity to an expert), but also enables new classes of constraints (such as diversity). Our approach comes with rigorous theoretical guarantees and only relies on the ability to approximately solve standard RL tasks. As a result, it can be easily adapted to work with any model-free or model-based RL. In our experiments, we show that it matches previous algorithms that enforce safety via constraints, but can also enforce new properties that these algorithms do not incorporate, such as diversity. Sobhan Miryoosefi, Kianté Brantley, Hal Daumé III, Miroslav Dudík, Robert E. Schapire |
NeurIPS | 3 |
| 2019 | Active Learning for Cost-Sensitive ClassificationabstractWe design an active learning algorithm for cost-sensitive multiclass classification: problems where different errors have different costs. Our algorithm, COAL, makes predictions by regressing to each label's cost and predicting the smallest. On a new example, it uses a set of regressors that perform well on past data to estimate possible costs for each label. It queries only the labels that could be the best, ignoring the sure losers. We prove COAL can be efficiently implemented for any regression family that admits squared loss optimization; it also enjoys strong guarantees with respect to predictive performance and labeling effort. We empirically compare COAL to passive learning and several active learning baselines, showing significant improvements in labeling effort and test cost on real-world datasets. Akshay Krishnamurthy, Alekh Agarwal, Tzu-Kuo Huang, Hal Daumé III, John Langford 0001 |
J. Mach. Learn. Res. | 4 |
| 2018 | Learning to Ask Good Questions: Ranking Clarification Questions using Neural Expected Value of Perfect InformationabstractInquiry is fundamental to communication, and machines cannot effectively collaborate with humans unless they can ask questions.In this work, we build a neural network model for the task of ranking clarification questions.Our model is inspired by the idea of expected value of perfect information: a good question is one whose expected answer will be useful.We study this problem using data from StackExchange, a plentiful online resource in which people routinely ask clarifying questions to posts so that they can better offer assistance to the original poster.We create a dataset of clarification questions consisting of ∼77K posts paired with a clarification question (and answer) from three domains of StackExchange: askubuntu, unix and superuser.We evaluate our model on 500 samples of this dataset against expert human judgments and demonstrate significant improvements over controlled baselines. Sudha Rao, Hal Daumé III |
ACL (1) | 2 |
| 2018 | Content Selection in Deep Learning Models of SummarizationabstractWe carry out experiments with deep learning models of summarization across the domains of news, personal stories, meetings, and medical articles in order to understand how content selection is performed.We find that many sophisticated features of state of the art extractive summarizers do not improve performance over simpler models.These results suggest that it is easier to create a summarizer for a new domain than previous work suggests and bring into question the benefit of deep learning models for summarization for those domains that do have massive datasets (i.e., news).At the same time, they suggest important questions for new research in summarization; namely, new forms of sentence representations or external knowledge sources are needed that are better suited to the sumarization task. Chris Kedzie, Kathy McKeown, Hal Daumé III |
EMNLP | 3 |
| 2018 | Residual Loss Prediction: Reinforcement Learning With No Incremental Feedback
Hal Daumé III, John Langford 0001, Amr Sharaf |
ICLR (Poster) | 1 |
| 2018 | Hierarchical Imitation and Reinforcement LearningabstractWe study how to effectively leverage expert feedback to learn sequential decision-making policies. We focus on problems with sparse rewards and long time horizons, which typically pose significant challenges in reinforcement learning. We propose an algorithmic framework, called hierarchical guidance, that leverages the hierarchical structure of the underlying problem to integrate different modes of expert interaction. Our framework can incorporate different combinations of imitation learning (IL) and reinforcement learning (RL) at different levels, leading to dramatic reductions in both expert effort and cost of exploration. Using long-horizon benchmarks, including Montezuma’s Revenge, we demonstrate that our approach can learn significantly faster than hierarchical RL, and be significantly more label-efficient than standard IL. We also theoretically analyze labeling cost for certain instantiations of our framework. Hoang Minh Le 0002, Nan Jiang 0008, Alekh Agarwal, Miroslav Dudík, Yisong Yue, Hal Daumé III |
ICML | 6 |
| 2018 | When Does Machine Learning FAIL? Generalized Transferability for Evasion and Poisoning Attacks
Octavian Suciu, Radu Marginean, Yigitcan Kaya, Hal Daumé III, Tudor Dumitras |
USENIX Security Symposium | 4 |
| 2017 | Unsupervised Learning of Evolving Relationships Between Literary CharactersabstractUnderstanding inter-character relationships is fundamental for understanding character intentions and goals in a narrative. This paper addresses unsupervised modeling of relationships between characters. We model relationships as dynamic phenomenon, represented as evolving sequences of latent states empirically learned from data. Unlike most previous work our approach is completely unsupervised. This enables data-driven inference of inter-character relationship types beyond simple sentiment polarities, by incorporating lexical and semantic representations, and leveraging large quantities of raw text. We present three models based on rich sets of linguistic features that capture various cues about relationships. We compare these models with existing techniques and also demonstrate that relationship categories learned by our model are semantically coherent. Snigdha Chaturvedi, Mohit Iyyer, Hal Daumé III |
AAAI | 3 |
| 2017 | The Amazing Mysteries of the Gutter: Drawing Inferences Between Panels in Comic Book NarrativesabstractVisual narrative is often a combination of explicit information and judicious omissions, relying on the viewer to supply missing details. In comics, most movements in time and space are hidden in the gutters between panels. To follow the story, readers logically connect panels together by inferring unseen actions through a process called closure. While computers can now describe the content of natural images, in this paper we examine whether they can understand the closure-driven narratives conveyed by stylized artwork and dialogue in comic book panels. We collect a dataset, COMICS, that consists of over 1.2 million panels (120 GB) paired with automatic textbox transcriptions. An in-depth analysis of COMICS demonstrates that neither text nor image alone can tell a comic book story, so a computer must understand both modalities to keep up with the plot. We introduce three cloze-style tasks that ask models to predict narrative and character-centric aspects of a panel given n preceding panels as context. Various deep neural architectures underperform human baselines on these tasks, suggesting that COMICS contains fundamental challenges for both vision and language. Mohit Iyyer, Varun Manjunatha, Anupam Guha, Yogarshi Vyas, Jordan L. Boyd-Graber, Hal Daumé III, Larry Davis 0001 |
CVPR | 6 |
| 2017 | Reinforcement Learning for Bandit Neural Machine Translation with Simulated Human FeedbackabstractMachine translation is a natural candidate problem for reinforcement learning from human feedback: users provide quick, dirty ratings on candidate translations to guide a system to improve.Yet, current neural machine translation training focuses on expensive human-generated reference translations.We describe a reinforcement learning algorithm that improves neural machine translation systems from simulated human feedback.Our algorithm combines the advantage actor-critic algorithm (Mnih et al., 2016) with the attention-based neural encoderdecoder architecture (Luong et al., 2015).This algorithm (a) is well-designed for problems with a large action space and delayed rewards, (b) effectively optimizes traditional corpus-level machine translation metrics, and (c) is robust to skewed, high-variance, granular feedback modeled after actual human behaviors. Hal Daumé III, Jordan L. Boyd-Graber |
EMNLP | 2 |
| 2017 | Logarithmic Time One-Against-SomeabstractWe create a new online reduction of multiclass classification to binary classification for which training and prediction time scale logarithmically with the number of classes. We show that several simple techniques give rise to an algorithm which is superior to previous logarithmic time classification approaches while competing with one-against-all in space. The core construction is based on using a tree to select a small subset of labels with high recall, which are then scored using a one-against-some structure with high precision. Hal Daumé III, Nikos Karampatziakis, John Langford 0001, Paul Mineiro |
ICML | 1 |
| 2017 | Active Learning for Cost-Sensitive ClassificationabstractWe design an active learning algorithm for cost-sensitive multiclass classification: problems where different errors have different costs. Our algorithm, COAL, makes predictions by regressing to each label’s cost and predicting the smallest. On a new example, it uses a set of regressors that perform well on past data to estimate possible costs for each label. It queries only the labels that could be the best, ignoring the sure losers. We prove COAL can be efficiently implemented for any regression family that admits squared loss optimization; it also enjoys strong guarantees with respect to predictive performance and labeling effort. Our experiment with COAL show significant improvements in labeling effort and test cost over passive and active baselines. Akshay Krishnamurthy, Alekh Agarwal, Tzu-Kuo Huang, Hal Daumé III, John Langford 0001 |
ICML | 4 |
| 2016 | Short Text Representation for Detecting Churn in MicroblogsabstractChurn happens when a customer leaves a brand or stop using its services. Brands reduce their churn rates by identifying and retaining potential churners through customer retention campaigns. In this paper, we consider the problem of classifying micro-posts as churny or non-churny with respect to a given brand. Motivated by the recent success of recurrent neural networks (RNNs) in word representation, we propose to utilize RNNs to learn micro-post and churn indicator representations. We show that such representations improve the performance of churn detection in microblogs and lead to more accurate ranking of churny contents. Furthermore, in this researchwe show that state-of-the-art sentiment analysis approaches fail to identify churny contents. Experiments on Twitter data about three telco brands show the utility of our approach for this task. Hadi Amiri, Hal Daumé III |
AAAI | 2 |
| 2016 | Ask, and Shall You Receive? Understanding Desire Fulfillment in Natural Language TextabstractThe ability to comprehend wishes or desires and their fulfillment is important to Natural Language Understanding. This paper introduces the task of identifying if a desire expressed by a subject in a given short piece of text was fulfilled. We propose various unstructured and structured models that capture fulfillment cues such as the subject's emotional state and actions. Our experiments with two different datasets demonstrate the importance of understanding the narrative and discourse structure to address this task. Snigdha Chaturvedi, Dan Goldwasser, Hal Daumé III |
AAAI | 3 |
| 2016 | Modeling Evolving Relationships Between Characters in Literary NovelsabstractStudying characters plays a vital role in computationally representing and interpreting narratives. Unlike previous work, which has focused on inferring character roles, we focus on the problem of modeling their relationships. Rather than assuming a fixed relationship for a character pair, we hypothesize that relationships temporally evolve with the progress of the narrative, and formulate the problem of relationship modeling as a structured prediction problem. We propose a semi-supervised framework to learn relationship sequences from fully as well as partially labeled data. We present a Markovian model capable of accumulating historical beliefs about the relationship and status changes. We use a set of rich linguistic and semantically motivated features that incorporate world knowledge to investigate the textual content of narrative. We empirically demonstrate that such a framework outperforms competitive baselines. Snigdha Chaturvedi, Hal Daumé III, Chris Dyer |
AAAI | 3 |
| 2016 | Learning Text Pair Similarity with Context-sensitive Autoencoders
Hadi Amiri, Philip Resnik, Jordan L. Boyd-Graber, Hal Daumé III |
ACL (1) | 4 |
| 2016 | Interpretese vs. Translationese: The Uniqueness of Human Strategies in Simultaneous InterpretationabstractComputational approaches to simultaneous interpretation are stymied by how little we know about the tactics human interpreters use.We produce a parallel corpus of translated and simultaneously interpreted text and study differences between them through a computational approach.Our analysis reveals that human interpreters regularly apply several effective tactics to reduce translation latency, including sentence segmentation and passivization.In addition to these unique, clever strategies, we show that limited human memory also causes other idiosyncratic properties of human interpretation such as generalization and omission of source content. He He 0001, Jordan L. Boyd-Graber, Hal Daumé III |
HLT-NAACL | 3 |
| 2016 | Feuding Families and Former Friends: Unsupervised Learning for Dynamic Fictional RelationshipsabstractMohit Iyyer, Anupam Guha, Snigdha Chaturvedi, Jordan Boyd-Graber, Hal Daumé III. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Mohit Iyyer, Anupam Guha, Snigdha Chaturvedi, Jordan L. Boyd-Graber, Hal Daumé III |
HLT-NAACL | 5 |
| 2016 | A Credit Assignment Compiler for Joint PredictionabstractMany machine learning applications involve jointly predicting multiple mutually dependent output variables. Learning to search is a family of methods where the complex decision problem is cast into a sequence of decisions via a search space. Although these methods have shown promise both in theory and in practice, implementing them has been burdensomely awkward. In this paper, we show the search space can be defined by an arbitrary imperative program, turning learning to search into a credit assignment compiler. Altogether with the algorithmic improvements for the compiler, we radically reduce the complexity of programming and the running time. We demonstrate the feasibility of our approach on multiple joint prediction tasks. In all cases, we obtain accuracies as high as alternative approaches, at drastically reduced execution and programming time. Kai-Wei Chang 0001, He He 0001, Stéphane Ross, Hal Daumé III, John Langford 0001 |
NIPS | 4 |
| 2016 | Large Scale Retrieval and Generation of Image Descriptions
Vicente Ordonez, Xufeng Han, Polina Kuznetsova, Girish Kulkarni, Margaret Mitchell, Kota Yamaguchi, Karl Stratos, Amit Goyal 0001, Jesse Dodge, Alyssa C. Mensch, Hal Daumé III, Alexander C. Berg, Yejin Choi 0001, Tamara L. Berg |
Int. J. Comput. Vis. | 11 |
| 2016 | Predicting the impact of scientific concepts using full-text featuresabstractNew scientific concepts, interpreted broadly, are continuously introduced in the literature, but relatively few concepts have a long‐term impact on society. The identification of such concepts is a challenging prediction task that would help multiple parties—including researchers and the general public—focus their attention within the vast scientific literature. In this paper we present a system that predicts the future impact of a scientific concept, represented as a technical term, based on the information available from recently published research articles. We analyze the usefulness of rich features derived from the full text of the articles through a variety of approaches, including rhetorical sentence analysis, information extraction, and time‐series analysis. The results from two large‐scale experiments with 3.8 million full‐text articles and 48 million metadata records support the conclusion that full‐text features are significantly more useful for prediction than metadata‐only features and that the most accurate predictions result from combining the metadata and full‐text features. Surprisingly, these results hold even when the metadata features are available for a much larger number of documents than are available for the full‐text features. Kathy McKeown, Hal Daumé III, Snigdha Chaturvedi, John Paparrizos, Kapil Thadani, Pablo Barrio 0002, Or Biran, Suvarna Bothe, Michael Collins 0001, Kenneth R. Fleischmann, Luis Gravano, Rahul Jha, Ben King, Kevin McInerney, Taesun Moon, Arvind Neelakantan, Diarmuid Ó Séaghdha, Dragomir R. Radev, Thomas Clay Templeton, Simone Teufel |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2016 | Learning Reductions That Really WorkabstractIn this paper, we provide a summary of the mathematical and computational techniques that have enabled learning reductions to effectively address a wide class of tasks, and show that this approach to solving machine learning problems can be broadly useful. Our work is instantiated and tested in a machine learning library, Vowpal Wabbit, to prove that the techniques discussed here are fully viable in practice. Alina Beygelzimer, Hal Daumé III, John Langford 0001, Paul Mineiro |
Proc. IEEE | 2 |
| 2015 | Target-Dependent Churn Classification in MicroblogsabstractWe consider the problem of classifying micro-posts as churny or non-churny with respect to a given brand. Using Twitter data about three brands, we find that standard machine learning techniques clearly outperform keyword based approaches. However, the three machine learning techniques we employed (linear classification, support vector machines, and logistic regression) do not perform as well on churn classification as on other text classification problems. We investigate demographic, content, and context churn indicators in microblogs and examine factors that make this problem more challenging. Experimental results show an average F1 performance of 75% for target-dependent churn classification in microblogs. Hadi Amiri, Hal Daumé III |
AAAI | 2 |
| 2015 | Deep Unordered Composition Rivals Syntactic Methods for Text ClassificationabstractMohit Iyyer, Varun Manjunatha, Jordan Boyd-Graber, Hal Daumé III. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Mohit Iyyer, Varun Manjunatha, Jordan L. Boyd-Graber, Hal Daumé III |
ACL (1) | 4 |
| 2015 | Why discourse affects speakers' choice of referring expressionsabstractNaho Orita, Eliana Vornov, Naomi Feldman, Hal Daumé III. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Naho Orita, Eliana Vornov, Naomi Feldman, Hal Daumé III |
ACL (1) | 4 |
| 2015 | Syntax-based Rewriting for Simultaneous Machine TranslationabstractDivergent word order between languages causes delay in simultaneous machine translation.We present a sentence rewriting method that generates more monotonic translations to improve the speedaccuracy tradeoff.We design grammaticality and meaning-preserving syntactic transformation rules that operate on constituent parse trees.We apply the rules to reference translations to make their word order closer to the source language word order.On Japanese-English translation (two languages with substantially different structure), incorporating the rewritten, more monotonic reference translation into a phrase-based machine translation system enables better translations faster than a baseline system that only uses gold reference translations. He He 0001, Alvin Grissom II, John Morgan, Jordan L. Boyd-Graber, Hal Daumé III |
EMNLP | 5 |
| 2015 | On Correcting Inputs: Inverse Optimization for Online Structured PredictionabstractAlgorithm designers typically assume that the input data is correct, and then proceed to find "optimal" or "sub-optimal" solutions using this input data. However this assumption of correct data does not always hold in practice, especially in the context of online learning systems where the objective is to learn appropriate feature weights given some training samples. Such scenarios necessitate the study of inverse optimization problems where one is given an input instance as well as a desired output and the task is to adjust the input data so that the given output is indeed optimal. Motivated by learning structured prediction models, in this paper we consider inverse optimization with a margin, i.e., we require the given output to be better than all other feasible outputs by a desired margin. We consider such inverse optimization problems for maximum weight matroid basis, matroid intersection, perfect matchings, minimum cost maximum flows, and shortest paths and derive the first known results for such problems with a non-zero margin. The effectiveness of these algorithmic approaches to online learning for structured prediction is also discussed. Hal Daumé III, Samir Khuller, Manish Purohit, Gregory Sanders |
FSTTCS | 1 |
| 2015 | Learning to Search Better than Your TeacherabstractMethods for learning to search for structured prediction typically imitate a reference policy, with existing theoretical guarantees demonstrating low regret compared to that reference. This is unsatisfactory in many applications where the reference policy is suboptimal and the goal of learning is to improve upon it. Can learning to search work even when the reference is poor? We provide a new learning to search algorithm, LOLS, which does well relative to the reference policy, but additionally guarantees low regret compared to deviations from the learned policy: a local-optimality guarantee. Consequently, LOLS can improve upon the reference policy, unlike previous algorithms. This enables us to develop structured contextual bandits, a partial information structured prediction setting with many potential applications. Kai-Wei Chang 0001, Akshay Krishnamurthy, Alekh Agarwal, Hal Daumé III, John Langford 0001 |
ICML | 4 |
| 2015 | Hands-on Learning to Search for Structured PredictionabstractHal Daumé III, John Langford, Kai-Wei Chang, He He, Sudha Rao. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Tutorial Abstracts. 2015. Hal Daumé III, John Langford 0001, Kai-Wei Chang 0001, He He 0001, Sudha Rao |
HLT-NAACL | 1 |
| 2015 | Dialogue focus tracking for zero pronoun resolutionabstractSudha Rao, Allyson Ettinger, Hal Daumé III, Philip Resnik. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Sudha Rao, Allyson Ettinger, Hal Daumé III, Philip Resnik |
HLT-NAACL | 3 |
| 2014 | Learning Latent Engagement Patterns of Students in Online CoursesabstractMaintaining and cultivating student engagement is critical for learning. Understanding factors affecting student engagement will help in designing better courses and improving student retention. The large number of participants in massive open online courses (MOOCs) and data collected from their interaction with the MOOC open up avenues for studying student engagement at scale. In this work, we develop a framework for modeling and understanding student engagement in online courses based on student behavioral cues. Our first contribution is the abstraction of student engagement types using latent representations and using that in a probabilistic model to connect student behavior with course completion. We demonstrate that the latent formulation for engagement helps in predicting student survival across three MOOCs. Next, in order to initiate better instructor interventions, we need to be able to predict student survival early in the course. We demonstrate that we can predict student survival early in the course reliably using the latent model. Finally, we perform a closer quantitative analysis of user interaction with the MOOC and identify student activities that are good indicators for survival at different points in the course. Arti Ramesh, Dan Goldwasser, Bert Huang, Hal Daumé III, Lise Getoor |
AAAI | 4 |
| 2014 | Predicting Instructor's Intervention in MOOC forumsabstractInstructor intervention in student discussion forums is a vital component in Massive Open Online Courses (MOOCs), where personalized interaction is limited. This paper introduces the problem of predicting instructor interventions in MOOC forums. We propose several prediction models designed to capture unique aspects of MOOCs, combining course information, forum structure and posts content. Our models abstract contents of individual posts of threads using latent categories, learned jointly with the binary intervention prediction problem. Experiments over data from two Coursera MOOCs demonstrate that incorporating the structure of threads into the learning problem leads to better predictive performance. Snigdha Chaturvedi, Dan Goldwasser, Hal Daumé III |
ACL (1) | 3 |
| 2014 | A Unified Model for Soft Linguistic Reordering Constraints in Statistical Machine TranslationabstractThis paper explores a simple and effective unified framework for incorporating soft linguistic reordering constraints into a hierarchical phrase-based translation system: 1) a syntactic reordering model that explores reorderings for context free grammar rules; and 2) a semantic reordering model that focuses on the reordering of predicate-argument structures.We develop novel features based on both models and use them as soft constraints to guide the translation process.Experiments on Chinese-English translation show that the reordering approach can significantly improve a state-of-the-art hierarchical phrase-based translation system.However, the gain achieved by the semantic reordering model is limited in the presence of the syntactic reordering model, and we therefore provide a detailed analysis of the behavior differences between the two. Yuval Marton, Philip Resnik, Hal Daumé III |
ACL (1) | 4 |
| 2014 | "I Object!" Modeling Latent Pragmatic Effects in Courtroom DialoguesabstractUnderstanding the actionable outcomes of a dialogue requires effectively modeling situational roles of dialogue participants, the structure of the dialogue and the relevance of each utterance to an eventual action.We develop a latent-variable model that can capture these notions and apply it in the context of courtroom dialogues, in which the objection speech act is used as binary supervision to drive the learning process.We demonstrate quantitatively and qualitatively that our model is able to uncover natural discourse structure from this distant supervision. Dan Goldwasser, Hal Daumé III |
EACL | 2 |
| 2014 | Don't Until the Final Verb Wait: Reinforcement Learning for Simultaneous Machine TranslationabstractWe introduce a reinforcement learningbased approach to simultaneous machine translation-producing a translation while receiving input wordsbetween languages with drastically different word orders: from verb-final languages (e.g., German) to verb-medial languages (English).In traditional machine translation, a translator must "wait" for source material to appear before translation begins.We remove this bottleneck by predicting the final verb in advance.We use reinforcement learning to learn when to trust predictions about unseen, future portions of the sentence.We also introduce an evaluation metric to measure expeditiousness and quality.We show that our new translation model outperforms batch and monotone translation strategies. Alvin Grissom II, He He 0001, Jordan L. Boyd-Graber, John Morgan, Hal Daumé III |
EMNLP | 5 |
| 2014 | A Neural Network for Factoid Question Answering over ParagraphsabstractText classification methods for tasks like factoid question answering typi-cally use manually defined string match-ing rules or bag of words representa-tions. These methods are ineffective when question text contains very few individual words (e.g., named entities) that are indicative of the answer. We introduce a recursive neural network (rnn) model that can reason over such input by modeling textual composition-ality. We apply our model, qanta, to a dataset of questions from a trivia competition called quiz bowl. Unlike previous rnn models, qanta learns word and phrase-level representations that combine across sentences to reason about entities. The model outperforms multiple baselines and, when combined with information retrieval methods, ri-vals the best human players. 1 Mohit Iyyer, Jordan L. Boyd-Graber, Leonardo Max Batista Claudino, Richard Socher, Hal Daumé III |
EMNLP | 5 |
| 2014 | Uncovering hidden engagement patterns for predicting learner performance in MOOCsabstractMaintaining and cultivating student engagement is a prerequisite for MOOCs to have broad educational impact. Understanding student engagement as a course progresses helps characterize student learning patterns and can aid in minimizing dropout rates, initiating instructor intervention. In this paper, we construct a probabilistic model connecting student behavior and class performance, formulating student engagement types as latent variables. We show that our model identifies course success indicators that can be used by instructors to initiate interventions and assist students. Arti Ramesh, Dan Goldwasser, Bert Huang, Hal Daumé III, Lise Getoor |
L@S | 4 |
| 2014 | Learning to Search in Branch and Bound Algorithms
He He 0001, Hal Daumé III, Jason Eisner |
NIPS | 2 |
| 2014 | Guest Editor's Introduction to the Special Issue on Domain Adaptation for Vision Applications
Dong Xu 0001, Rama Chellappa, Trevor Darrell, Hal Daumé III |
Int. J. Comput. Vis. | 4 |
| 2013 | SenseSpotting: Never let your parallel data tie you to an old domain
Marine Carpuat, Hal Daumé III, Katharine Henry, Ann Irvine, Jagadeesh Jagarlamudi, Rachel Rudinger |
ACL (1) | 2 |
| 2013 | Dynamic Feature Selection for Dependency ParsingabstractFeature computation and exhaustive search have significantly restricted the speed of graph-based dependency parsing.We propose a faster framework of dynamic feature selection, where features are added sequentially as needed, edges are pruned early, and decisions are made online for each sentence.We model this as a sequential decision-making problem and solve it by imitation learning techniques.We test our method on 7 languages.Our dynamic parser can achieve accuracies comparable or even superior to parsers using a full set of features, while computing fewer than 30% of the feature templates. He He 0001, Hal Daumé III, Jason Eisner |
EMNLP | 2 |
| 2013 | Monolingual Marginal Matching for Translation Model AdaptationabstractWhen using a machine translation (MT) model trained on OLD-domain parallel data to translate NEW-domain text, one major challenge is the large number of out-of-vocabulary (OOV) and new-translation-sense words.We present a method to identify new translations of both known and unknown source language words that uses NEW-domain comparable document pairs.Starting with a joint distribution of source-target word pairs derived from the OLD-domain parallel corpus, our method recovers a new joint distribution that matches the marginal distributions of the NEW-domain comparable document pairs, while minimizing the divergence from the OLD-domain distribution.Adding learned translations to our French-English MT model results in gains of about 2 BLEU points over strong baselines. Ann Irvine, Chris Quirk, Hal Daumé III |
EMNLP | 3 |
| 2013 | Kernel regression for Head-Related Transfer Function interpolation and spectral extrema extractionabstractHead-Related Transfer Function (HRTF) representation and interpolation is an important problem in spatial audio. We present a kernel regression method based on Gaussian process (GP) modeling of the joint spatial-frequency relationship between HRTF measurements and obtain a smooth non-linear representation based on data measured over both arbitrary and structured spherical measurement grids. This representation is further extended to the problem of extracting spectral extrema (notches and peaks). We perform HRTF interpolation and spectral extrema extraction using freely available CIPIC HRTF data. Experimental results are shown. Yuancheng Luo, Dmitry N. Zotkin, Hal Daumé III, Ramani Duraiswami |
ICASSP | 3 |
| 2013 | Discriminatively Enhanced Topic ModelsabstractThis paper proposes a space-efficient, discriminatively enhanced topic model: a V structured topic model with an embedded log-linear component. The discriminative log-linear component reduces the number of parameters to be learnt while outperforming baseline generative models. At the same time, the explanatory power of the generative component is not compromised. We establish its superiority over a purely generative model by applying it to two different ranking tasks: (a) In the first task, we look at the problem of proposing alternative citations given textual and bibliographic evidence. We solve it as a ranking problem in itself and as a platform for further qualitative analysis of convergence of scientific phenomenon. (b) In the second task we address the problem of ranking potential email recipients based on email content and sender information. Snigdha Chaturvedi, Hal Daumé III, Taesun Moon |
ICDM | 2 |
| 2013 | Predictable Dual-View HashingabstractWe propose a Predictable Dual-View Hashing (PDH) algorithm which embeds proximity of data samples in the original spaces. We create a cross-view hamming space with the ability to compare information from previously incomparable domains with a notion of ‘predictability’. By performing comparative experimental analysis on two large datasets, PASCAL-Sentence and SUN-Attribute, we demonstrate the superiority of our method to the state-of-the-art dual-view binary code learning algorithms. Mohammad Rastegari, Shobeir Fakhraei, Hal Daumé III, Larry Davis 0001 |
ICML (3) | 4 |
| 2013 | Modeling Syntactic and Semantic Structures in Hierarchical Phrase-based Translation
Philip Resnik, Hal Daumé III |
HLT-NAACL | 3 |
| 2013 | Binary to Bushy: Bayesian Hierarchical Clustering with the Beta CoalescentabstractDiscovering hierarchical regularities in data is a key problem in interacting with large datasets, modeling cognition, and encoding knowledge. A previous Bayesian solution---Kingman's coalescent---provides a convenient probabilistic model for data represented as a binary tree. Unfortunately, this is inappropriate for data better described by bushier trees. We generalize an existing belief propagation framework of Kingman's coalescent to the beta coalescent, which models a wider range of tree structures. Because of the complex combinatorial search over possible structures, we develop new sampling schemes using sequential Monte Carlo and Dirichlet process mixture models, which render inference efficient and tractable. We present results on both synthetic and real data that show the beta coalescent outperforms Kingman's coalescent on real datasets and is qualitatively better at capturing data in bushy hierarchies. Yuening Hu, Jordan L. Boyd-Graber, Hal Daumé III, Z. Irene Ying |
NIPS | 3 |
| 2013 | A Computational Model for Plot UnitsabstractThis research revisitsplot units, which were developed in the 1980s as a conceptual knowledge structure to represent the affect states of and emotional tensions between characters in narrative stories. We present a fully automated system, called AESOP, that generates plot unit representations for narrative texts. AESOP performs four steps: affect state recognition, character identification, affect state projection, and link creation. We also identify a type of knowledge that seems to be missing from existing lexical resources: verbs that impart positive or negative polarity onto their patients (e.g., “eat” imparts negative polarity because being eaten is bad, whereas “fed” imparts positive polarity because being fed is good). We develop two techniques to automatically harvest these “patient polarity verbs” (PPVs) from a Web corpus, and show that the PPVs improve affect state recognition. Finally, we evaluate AESOP’s performance on a set of fables, and present several analyses to shed light on the capabilities and limitations of current natural language processing technology for plot unit generation. Amit Goyal 0001, Ellen Riloff, Hal Daumé III |
Comput. Intell. | 3 |
| 2013 | Improving performance of natural language processing part-of-speech tagging on clinical narratives through domain adaptationabstractOBJECTIVE: Natural language processing (NLP) tasks are commonly decomposed into subtasks, chained together to form processing pipelines. The residual error produced in these subtasks propagates, adversely affecting the end objectives. Limited availability of annotated clinical data remains a barrier to reaching state-of-the-art operating characteristics using statistically based NLP tools in the clinical domain. Here we explore the unique linguistic constructions of clinical texts and demonstrate the loss in operating characteristics when out-of-the-box part-of-speech (POS) tagging tools are applied to the clinical domain. We test a domain adaptation approach integrating a novel lexical-generation probability rule used in a transformation-based learner to boost POS performance on clinical narratives. METHODS: Two target corpora from independent healthcare institutions were constructed from high frequency clinical narratives. Four leading POS taggers with their out-of-the-box models trained from general English and biomedical abstracts were evaluated against these clinical corpora. A high performing domain adaptation method, Easy Adapt, was compared to our newly proposed method ClinAdapt. RESULTS: The evaluated POS taggers drop in accuracy by 8.5-15% when tested on clinical narratives. The highest performing tagger reports an accuracy of 88.6%. Domain adaptation with Easy Adapt reports accuracies of 88.3-91.0% on clinical texts. ClinAdapt reports 93.2-93.9%. CONCLUSIONS: ClinAdapt successfully boosts POS tagging performance through domain adaptation requiring a modest amount of annotated clinical data. Improving the performance of critical NLP subtasks is expected to reduce pipeline error propagation leading to better overall results on complex processing tasks. Jeffrey P. Ferraro, Hal Daumé III, Scott L. DuVall, Wendy W. Chapman, Henk Harkema, Peter J. Haug |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | Measuring Machine Translation Errors in New DomainsabstractWe develop two techniques for analyzing the effect of porting a machine translation system to a new domain. One is a macro-level analysis that measures how domain shift affects corpus-level evaluation; the second is a micro-level analysis for word-level errors. We apply these methods to understand what happens when a Parliament-trained phrase-based machine translation system is applied in four very different domains: news, medical texts, scientific articles and movie subtitles. We present quantitative and qualitative experiments that highlight opportunities for future research in domain adaptation for machine translation. Ann Irvine, John Morgan, Marine Carpuat, Hal Daumé III, Dragos Stefan Munteanu |
Trans. Assoc. Comput. Linguistics | 4 |
| 2012 | Efficient Protocols for Distributed Classification and Optimization
Hal Daumé III, Jeff M. Phillips, Avishek Saha, Suresh Venkatasubramanian |
ALT | 1 |
| 2012 | Understanding and predicting importance in imagesabstractWhat do people care about in an image? To drive computational visual recognition toward more human-centric outputs, we need a better understanding of how people perceive and judge the importance of content in images. In this paper, we explore how a number of factors relate to human perception of importance. Proposed factors fall into 3 broad types: 1) factors related to composition, e.g. size, location, 2) factors related to semantics, e.g. category of object or scene, and 3) contextual factors related to the likelihood of attribute-object, or object-scene pairs. We explore these factors using what people describe as a proxy for importance. Finally, we build models to predict what will be described about an image given either known image content, or image content estimated automatically by recognition systems. Alexander C. Berg, Tamara L. Berg, Hal Daumé III, Jesse Dodge, Amit Goyal 0001, Xufeng Han, Alyssa C. Mensch, Margaret Mitchell, Aneesh Sood, Karl Stratos, Kota Yamaguchi |
CVPR | 3 |
| 2012 | Generalized Multiview Analysis: A discriminative latent spaceabstractThis paper presents a general multi-view feature extraction approach that we call Generalized Multiview Analysis or GMA. GMA has all the desirable properties required for cross-view classification and retrieval: it is supervised, it allows generalization to unseen classes, it is multi-view and kernelizable, it affords an efficient eigenvalue based solution and is applicable to any domain. GMA exploits the fact that most popular supervised and unsupervised feature extraction techniques are the solution of a special form of a quadratic constrained quadratic program (QCQP), which can be solved efficiently as a generalized eigenvalue problem. GMA solves a joint, relaxed QCQP over different feature spaces to obtain a single (non)linear subspace. Intuitively, GMA is a supervised extension of Canonical Correlational Analysis (CCA), which is useful for cross-view classification and retrieval. The proposed approach is general and has the potential to replace CCA whenever classification or retrieval is the purpose and label information is available. We outperform previous approaches for textimage retrieval on Pascal and Wiki text-image data. We report state-of-the-art results for pose and lighting invariant face recognition on the MultiPIE face dataset, significantly outperforming other approaches. Abhishek Sharma 0001, Abhishek Kumar 0001, Hal Daumé III, David Jacobs 0001 |
CVPR | 3 |
| 2012 | Incorporating Lexical Priors into Topic Models
Jagadeesh Jagarlamudi, Hal Daumé III, Raghavendra Udupa |
EACL | 2 |
| 2012 | Midge: Generating Image Descriptions From Computer Vision Detections
Margaret Mitchell, Jesse Dodge, Amit Goyal 0001, Kota Yamaguchi, Karl Stratos, Xufeng Han, Alyssa C. Mensch, Alexander C. Berg, Tamara L. Berg, Hal Daumé III |
EACL | 10 |
| 2012 | Besting the Quiz Master: Crowdsourcing Incremental Classification Games
Jordan L. Boyd-Graber, Brianna Satinoff, He He 0001, Hal Daumé III |
EMNLP-CoNLL | 4 |
| 2012 | Sketch Algorithms for Estimating Point Queries in NLP
Amit Goyal 0001, Hal Daumé III, Graham Cormode |
EMNLP-CoNLL | 2 |
| 2012 | Fast Large-Scale Approximate Graph Construction for NLP
Amit Goyal 0001, Hal Daumé III, Raul Guerra |
EMNLP-CoNLL | 2 |
| 2012 | Regularized Interlingual Projections: Evaluation on Multilingual Transliteration
Jagadeesh Jagarlamudi, Hal Daumé III |
EMNLP-CoNLL | 2 |
| 2012 | Learning Task Grouping and Overlap in Multi-task Learning
Abhishek Kumar 0001, Hal Daumé III |
ICML | 2 |
| 2012 | A Binary Classification Framework for Two-Stage Multiple Kernel Learning
Abhishek Kumar 0001, Alexandru Niculescu-Mizil, Koray Kavukcuoglu, Hal Daumé III |
ICML | 4 |
| 2012 | Flexible Modeling of Latent Task Structures in Multitask Learning
Alexandre Tachard Passos, Piyush Rai, Jacques Wainer, Hal Daumé III |
ICML | 4 |
| 2012 | Towards a Watson that sees: Language-guided action recognition for robotsabstractFor robots of the future to interact seamlessly with humans, they must be able to reason about their surroundings and take actions that are appropriate to the situation. Such reasoning is only possible when the robot has knowledge of how the World functions, which must either be learned or hard-coded. In this paper, we propose an approach that exploits language as an important resource of high-level knowledge that a robot can use, akin to IBM's Watson in Jeopardy!. In particular, we show how language can be leveraged to reduce the ambiguity that arises from recognizing actions involving hand-tools from video data. Starting from the premise that tools and actions are intrinsically linked, with one explaining the existence of the other, we trained a language model over a large corpus of English newswire text so that we can extract this relationship directly. This model is then used as a prior to select the best tool and action that explains the video. We formalize the approach in the context of 1) an unsupervised recognition and 2) a supervised classification scenario by an EM formulation for the former and integrating language features for the latter. Results are validated over a new hand-tool action dataset, and comparisons with state of the art STIP features showed significantly improved results when language is used. In addition, we discuss the implications of these results and how it provides a framework for integrating language into vision on other robotic applications. Ching Lik Teo, Yezhou Yang, Hal Daumé III, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 3 |
| 2012 | Detecting Visual Text
Jesse Dodge, Amit Goyal 0001, Xufeng Han, Alyssa C. Mensch, Margaret Mitchell, Karl Stratos, Kota Yamaguchi, Yejin Choi 0001, Hal Daumé III, Alexander C. Berg, Tamara L. Berg |
HLT-NAACL | 9 |
| 2012 | Low-Dimensional Discriminative Reranking
Jagadeesh Jagarlamudi, Hal Daumé III |
HLT-NAACL | 2 |
| 2012 | Imitation Learning by CoachingabstractImitation Learning has been shown to be successful in solving many challenging real-world problems. Some recent approaches give strong performance guarantees by training the policy iteratively. However, it is important to note that these guarantees depend on how well the policy we found can imitate the oracle on the training data. When there is a substantial difference between the oracle's ability and the learner's policy space, we may fail to find a policy that has low error on the training set. In such cases, we propose to use a coach that demonstrates easy-to-learn actions for the learner and gradually approaches the oracle. By a reduction of learning by demonstration to online learning, we prove that coaching can yield a lower regret bound than using the oracle. We apply our algorithm to a novel cost-sensitive dynamic feature selection problem, a hard decision problem that considers a user-specified accuracy-cost trade-off. Experimental results on UCI datasets show that our method outperforms state-of-the-art imitation learning methods in dynamic features selection and two static feature selection methods. He He 0001, Hal Daumé III, Jason Eisner |
NIPS | 2 |
| 2012 | Learned Prioritization for Trading Off Accuracy and SpeedabstractUsers want natural language processing (NLP) systems to be both fast and accurate, but quality often comes at the cost of speed. The field has been manually exploring various speed-accuracy tradeoffs (for particular problems and datasets). We aim to explore this space automatically, focusing here on the case of agenda-based syntactic parsing \cite{kay-1986}. Unfortunately, off-the-shelf reinforcement learning techniques fail to learn good policies: the state space is simply too large to explore naively. An attempt to counteract this by applying imitation learning algorithms also fails: the ``teacher'' is far too good to successfully imitate with our inexpensive features. Moreover, it is not specifically tuned for the known reward function. We propose a hybrid reinforcement/apprenticeship learning algorithm that, even with only a few inexpensive features, can automatically learn weights that achieve competitive accuracies at significant improvements in speed over state-of-the-art baselines. Jiarong Jiang, Adam R. Teichert, Hal Daumé III, Jason Eisner |
NIPS | 3 |
| 2012 | Simultaneously Leveraging Output and Task Structures for Multiple-Output RegressionabstractMultiple-output regression models require estimating multiple functions, one for each output. To improve parameter estimation in such models, methods based on structural regularization of the model parameters are usually needed. In this paper, we present a multiple-output regression model that leverages the covariance structure of the functions (i.e., how the multiple functions are related with each other) as well as the conditional covariance structure of the outputs. This is in contrast with existing methods that usually take into account only one of these structures. More importantly, unlike most of the other existing methods, none of these structures need be known a priori in our model, and are learned from the data. Several previously proposed structural regularization based multiple-output regression models turn out to be special cases of our model. Moreover, in addition to being a rich model for multiple-output regression, our model can also be used in estimating the graphical model structure of a set of variables (multivariate outputs) conditioned on another set of variables (inputs). Experimental results on both synthetic and real datasets demonstrate the effectiveness of our method. Piyush Rai, Abhishek Kumar 0001, Hal Daumé III |
NIPS | 3 |
| 2012 | Leveraging Social Bookmarks from Partially Tagged Corpus for Improved Web Page ClusteringabstractAutomatic clustering of Web pages helps a number of information retrieval tasks, such as improving user interfaces, collection clustering, introducing diversity in search results, etc. Typically, Web page clustering algorithms use only features extracted from the page-text. However, the advent of social-bookmarking Web sites, such as StumbleUpon.com and Delicious.com, has led to a huge amount of user-generated content such as the social tag information that is associated with the Web pages. In this article, we present a subspace based feature extraction approach that leverages the social tag information to complement the page-contents of a Web page for extracting beter features, with the goal of improved clustering performance. In our approach, we consider page-text and tags as two separate views of the data, and learn a shared subspace that maximizes the correlation between the two views. Any clustering algorithm can then be applied in this subspace. We then present an extension that allows our approach to be applicable even if the Web page corpus is only partially tagged, that is, when the social tags are present for not all, but only for a small number of Web pages. We compare our subspace based approach with a number of baselines that use tag information in various other ways, and show that the subspace based approach leads to improved performance on the Web page clustering task. We also discuss some possible future work including an active learning extension that can help in choosing which Web pages to get tags for, if we only can get the social tags for only a small number of Web pages. Anusua Trivedi, Piyush Rai, Hal Daumé III, Scott L. DuVall |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2011 | Lossy Conservative Update (LCU) Sketch: Succinct Approximate Count StorageabstractIn this paper, we propose a variant of the conservativeupdate Count-Min sketch to further reduce the overestimation error incurred. Inspired by ideas from lossy counting, we divide a stream of items into multiple windows, and decrement certain counts in the sketch at window boundaries. We refer to this approach as a lossy conservative update (LCU). The reduction in overestimation error of counts comes at the cost of introducing under-estimation error in counts. However, in our intrinsic evaluations, we show that the reduction in overestimation is much greater than the under-estimation error introduced by our method LCU. We apply our LCU framework to scale distributional similarity computations to web-scale corpora. We show that this technique is more efficient in terms of memory, and time, and more robust than conservative update with Count-Min (CU) sketch on this task. Amit Goyal 0001, Hal Daumé III |
AAAI | 2 |
| 2011 | Approximate Scalable Bounded Space Sketch for Large Data NLP
Amit Goyal 0001, Hal Daumé III |
EMNLP | 2 |
| 2011 | Improving Bilingual Projections via Sparse Covariance Matrices
Jagadeesh Jagarlamudi, Raghavendra Udupa, Hal Daumé III, Abhijit Bhole |
EMNLP | 3 |
| 2011 | Corpus-Guided Sentence Generation of Natural Images
Yezhou Yang, Ching Lik Teo, Hal Daumé III, Yiannis Aloimonos |
EMNLP | 3 |
| 2011 | A Co-training Approach for Multi-view Spectral Clustering
Abhishek Kumar 0001, Hal Daumé III |
ICML | 2 |
| 2011 | Beam Search based MAP Estimates for the Indian Buffet Process
Piyush Rai, Hal Daumé III |
ICML | 2 |
| 2011 | A Geometric View of Conjugate Priors
Arvind Agarwal, Hal Daumé III |
IJCAI | 2 |
| 2011 | Message-Passing for Approximate MAP Inference with Latent VariablesabstractWe consider a general inference setting for discrete probabilistic graphical models where we seek maximum a posteriori (MAP) estimates for a subset of the random variables (max nodes), marginalizing over the rest (sum nodes). We present a hybrid message-passing algorithm to accomplish this. The hybrid algorithm passes a mix of sum and max messages depending on the type of source node (sum or max). We derive our algorithm by showing that it falls out as the solution of a particular relaxation of a variational framework. We further show that the Expectation Maximization algorithm can be seen as an approximation to our algorithm. Experimental results on synthetic and real-world datasets, against several baselines, demonstrate the efficacy of our proposed algorithm. Jiarong Jiang, Piyush Rai, Hal Daumé III |
NIPS | 3 |
| 2011 | Co-regularized Multi-view Spectral ClusteringabstractIn many clustering problems, we have access to multiple views of the data each of which could be individually used for clustering. Exploiting information from multiple views, one can hope to find a clustering that is more accurate than the ones obtained using the individual views. Since the true clustering would assign a point to the same cluster irrespective of the view, we can approach this problem by looking for clusterings that are consistent across the views, i.e., corresponding data points in each view should have same cluster membership. We propose a spectral clustering framework that achieves this goal by co-regularizing the clustering hypotheses, and propose two co-regularization schemes to accomplish this. Experimental comparisons with a number of baselines on two synthetic and three real-world datasets establish the efficacy of our proposed approaches. Abhishek Kumar 0001, Piyush Rai, Hal Daumé III |
NIPS | 3 |
| 2011 | Active Supervised Domain Adaptation
Avishek Saha, Piyush Rai, Hal Daumé III, Suresh Venkatasubramanian, Scott L. DuVall |
ECML/PKDD (3) | 3 |
| 2010 | Kernelized Sorting for Natural Language ProcessingabstractKernelized sorting is an approach for matching objects from two sources (or domains) that does not require any prior notion of similarity between objects across the two sources. Unfortunately, this technique is highly sensitive to initialization and high dimensional data. We present variants of kernelized sorting to increase its robustness and performance on several Natural Language Processing (NLP) tasks: document matching from parallel and comparable corpora, machine transliteration and even image processing. Empirically we show that, on these tasks, a semi-supervised variant of kernelized sorting outperforms matching canonical correlation analysis. Jagadeesh Jagarlamudi, Seth Juarez, Hal Daumé III |
AAAI | 3 |
| 2010 | Extracting Multilingual Topics from Unaligned Comparable Corpora
Jagadeesh Jagarlamudi, Hal Daumé III |
ECIR | 2 |
| 2010 | Automatically Producing Plot Unit Representations for Narrative Text
Amit Goyal 0001, Ellen Riloff, Hal Daumé III |
EMNLP | 3 |
| 2010 | Learning Multiple Tasks using Manifold RegularizationabstractWe present a novel method for multitask learning (MTL) based on {\it manifold regularization}: assume that all task parameters lie on a manifold. This is the generalization of a common assumption made in the existing literature: task parameters share a common {\it linear} subspace. One proposed method uses the projection distance from the manifold to regularize the task parameters. The manifold structure and the task parameters are learned using an alternating optimization framework. When the manifold structure is fixed, our method decomposes across tasks which can be learnt independently. An approximation of the manifold regularization scheme is presented that preserves the convexity of the single task learning problem, and makes the proposed MTL framework efficient and easy to implement. We show the efficacy of our method on several datasets. Arvind Agarwal, Hal Daumé III, Samuel Gerber |
NIPS | 2 |
| 2010 | Co-regularization Based Semi-supervised Domain AdaptationabstractThis paper presents a co-regularization based approach to semi-supervised domain adaptation. Our proposed approach (EA++) builds on the notion of augmented space (introduced in EASYADAPT (EA) [1]) and harnesses unlabeled data in target domain to further enable the transfer of information from source to target. This semi-supervised approach to domain adaptation is extremely simple to implement and can be applied as a pre-processing step to any supervised learner. Our theoretical analysis (in terms of Rademacher complexity) of EA and EA++ show that the hypothesis class of EA++ has lower complexity (compared to EA) and hence results in tighter generalization bounds. Experimental results on sentiment analysis tasks reinforce our theoretical findings and demonstrate the efficacy of the proposed method when compared to EA as well as a few other baseline approaches. Hal Daumé III, Abhishek Kumar 0001, Avishek Saha |
NIPS | 1 |
| 2010 | A geometric view of conjugate priors
Arvind Agarwal, Hal Daumé III |
Mach. Learn. | 2 |
| 2009 | Unsupervised search-based structured predictionabstractWe describe an adaptation and application of a search-based structured prediction algorithm “Searn” to unsupervised learning problems. We show that it is possible to reduce unsupervised learning to supervised learning and demonstrate a high-quality unsupervised shift-reduce parsing model. We additionally show a close connection between unsupervised Searn and expectation maximization. Finally, we demonstrate the efficacy of a semi-supervised extension. The key idea that enables this is an application of the predict-self idea for unsupervised learning. Hal Daumé III |
ICML | 1 |
| 2009 | Exponential Family Hybrid Semi-Supervised Learning
Arvind Agarwal, Hal Daumé III |
IJCAI | 2 |
| 2009 | Streamed Learning: One-Pass SVMs
Piyush Rai, Hal Daumé III, Suresh Venkatasubramanian |
IJCAI | 2 |
| 2009 | Non-Parametric Bayesian Areal Linguistics
Hal Daumé III |
HLT-NAACL | 1 |
| 2009 | Streaming for large scale NLP: Language Modeling
Amit Goyal 0001, Hal Daumé III, Suresh Venkatasubramanian |
HLT-NAACL | 2 |
| 2009 | Multi-Label Prediction via Sparse Infinite CCAabstractCanonical Correlation Analysis (CCA) is a useful technique for modeling dependencies between two (or more) sets of variables. Building upon the recently suggested probabilistic interpretation of CCA, we propose a nonparametric, fully Bayesian framework that can automatically select the number of correlation components, and effectively capture the sparsity underlying the projections. In addition, given (partially) labeled data, our algorithm can also be used as a (semi)supervised dimensionality reduction technique, and can be applied to learn useful predictive features in the context of learning a set of related tasks. Experimental results demonstrate the efficacy of the proposed approach for both CCA as a stand-alone problem, and when applied to multi-label prediction. Piyush Rai, Hal Daumé III |
NIPS | 2 |
| 2009 | Bayesian Multitask Learning with Latent Hierarchies
Hal Daumé III |
UAI | 1 |
| 2009 | Search-based structured prediction
Hal Daumé III, John Langford 0001, Daniel Marcu |
Mach. Learn. | 1 |
| 2008 | Name Translation in Statistical Machine Translation - Learning When to Transliterate
Ulf Hermjakob, Kevin Knight, Hal Daumé III |
ACL | 3 |
| 2008 | Cross-Task Knowledge-Constrained Self Training
Hal Daumé III |
EMNLP | 1 |
| 2008 | Structure compilation: trading structure for featuresabstractStructured models often achieve excellent performance but can be slow at test time. We investigate structure compilation, where we replace structure with features, which are often computationally simpler but unfortunately statistically more complex. We analyze this tradeoff theoretically and empirically on three natural language processing tasks. We also introduce a simple method to transfer predictive power from structure to features via unlabeled data, while incurring a minimal statistical penalty. Percy Liang, Hal Daumé III, Daniel Klein 0001 |
ICML | 2 |
| 2008 | The Infinite Hierarchical Factor Regression ModelabstractWe propose a nonparametric Bayesian factor regression model that accounts for uncertainty in the number of factors, and the relationship between factors. To accomplish this, we propose a sparse variant of the Indian Buffet Process and couple this with a hierarchical model over factors, based on Kingman's coalescent. We apply this model to two problems (factor analysis and factor regression) in gene-expression data analysis. Piyush Rai, Hal Daumé III |
NIPS | 2 |
| 2007 | Frustratingly Easy Domain Adaptation
Hal Daumé III |
ACL | 1 |
| 2007 | A Bayesian Model for Discovering Typological Implications
Hal Daumé III, Lyle Campbell |
ACL | 1 |
| 2007 | Bayesian Agglomerative Clustering with CoalescentsabstractWe introduce a new Bayesian model for hierarchical clustering based on a prior over trees called Kingman’s coalescent. We develop novel greedy and sequential Monte Carlo inferences which operate in a bottom-up agglomerative fashion. We show experimentally the superiority of our algorithms over the state-of-the-art, and demonstrate our approach in document clustering and phylolinguistics. Yee Whye Teh, Hal Daumé III, Daniel M. Roy 0001 |
NIPS | 2 |
| 2006 | Bayesian Query-Focused SummarizationabstractWe present BAYESUM (for "Bayesian summarization"), a model for sentence extraction in query-focused summarization.BAYESUM leverages the common case in which multiple documents are relevant to a single query.Using these documents as reinforcement for query terms, BAYESUM is not afflicted by the paucity of information in short queries.We show that approximate inference in BAYESUM is possible on large data sets and results in a stateof-the-art summarization system.Furthermore, we show how BAYESUM can be understood as a justified query expansion technique in the language modeling for IR framework. Hal Daumé III, Daniel Marcu |
ACL | 1 |
| 2006 | Beyond EM: Bayesian Techniques for Human Language Technology Researchers
Hal Daumé III |
HLT-NAACL | 1 |
| 2006 | Domain Adaptation for Statistical ClassifiersabstractThe most basic assumption used in statistical learning theory is that training data and test data are drawn from the same underlying distribution. Unfortunately, in many applications, the "in-domain" test data is drawn from a distribution that is related, but not identical, to the "out-of-domain" distribution of the training data. We consider the common case in which labeled out-of-domain data is plentiful, but labeled in-domain data is scarce. We introduce a statistical formulation of this problem in terms of a simple mixture model and present an instantiation of this framework to maximum entropy classifiers and their linear chain counterparts. We present efficient inference algorithms for this special case based on the technique of conditional expectation maximization. Our experimental results show that our approach leads to improved performance on three real world tasks on four different data sets from the natural language processing domain. Hal Daumé III, Daniel Marcu |
J. Artif. Intell. Res. | 1 |
| 2005 | Learning as search optimization: approximate large margin methods for structured predictionabstractMappings to structured output spaces (strings, trees, partitions, etc.) are typically learned using extensions of classification algorithms to simple graphical structures (eg., linear chains) in which search and parameter estimation can be performed exactly. Unfortunately, in many complex problems, it is rare that exact search or parameter estimation is tractable. Instead of learning exact models and searching via heuristic means, we embrace this difficulty and treat the structured output problem in terms of approximate search. We present a framework for learning as search optimization, and two parameter updates with convergence the-orems and bounds. Empirical evidence shows that our integrated approach to learning and decoding can outperform exact models at smaller computational cost. Hal Daumé III, Daniel Marcu |
ICML | 1 |
| 2005 | Induction of Word and Phrase Alignments for Automatic Document SummarizationabstractCurrent research in automatic single-document summarization is dominated by two effective, yet naïve approaches: summarization by sentence extraction and headline generation via bagof-words models. While successful in some tasks, neither of these models is able to adequately capture the large set of linguistic devices utilized by humans when they produce summaries. One possible explanation for the widespread use of these models is that good techniques have been developed to extract appropriate training data for them from existing document/abstract and document/ headline corpora. We believe that future progress in automatic summarization will be driven both by the development of more sophisticated, linguistically informed models, as well as a more effective leveraging of document/abstract corpora. In order to open the doors to simultaneously achieving both of these goals, we have developed techniques for automatically producing word-to-word and phrase-to-phrase alignments between documents and their human-written abstracts. These alignments make explicit the correspondences that exist in such document/abstract pairs and create a potentially rich data source from which complex summarization algorithms may learn. This paper describes experiments we have carried out to analyze the ability of humans to perform such alignments, and based on these analyses, we describe experiments for creating them automatically. Our model for the alignment task is based on an extension of the standard hidden Markov model and learns to create alignments in a completely unsupervised fashion. We describe our model in detail and present experimental results that show that our model is able to learn to reliably identify word- and phrase-level alignments in a corpus of (document, abstract) pairs. Hal Daumé III, Daniel Marcu |
Comput. Linguistics | 1 |
| 2005 | A Bayesian Model for Supervised Clustering with the Dirichlet Process PriorabstractWe develop a Bayesian framework for tackling the supervised clustering problem, the generic problem encountered in tasks such as reference matching, coreference resolution, identity uncertainty and record linkage. Our clustering model is based on the Dirichlet process prior, which enables us to define distributions over the countably infinite sets that naturally arise in this problem. We add supervision to our model by positing the existence of a set of unobserved random variables (we call these "reference types") that are generic across all clusters. Inference in our framework, which requires integrating over infinitely many parameters, is solved using Markov chain Monte Carlo techniques. We present algorithms for both conjugate and non-conjugate priors. We present a simple---but general---parameterization of our model based on a Gaussian assumption. We evaluate this model on one artificial task and three real-world tasks, comparing it against both unsupervised and state-of-the-art supervised algorithms. Our results show that our model is able to outperform other models across a variety of tasks and performance metrics. Hal Daumé III, Daniel Marcu |
J. Mach. Learn. Res. | 1 |
| 2004 | A Phrase-Based HMM Approach to Document/Abstract Alignment
Hal Daumé III, Daniel Marcu |
EMNLP | 1 |
| 2004 | NP Bracketing by Maximum Entropy Tagging and SVM Reranking
Hal Daumé III, Daniel Marcu |
EMNLP | 1 |
| 2004 | Book Review, Inderjeet Mani: Automatic Summarization, John Benjamins Publishing, Amsterdam, The Netherlands, 2001, xi + 286 pp
Hal Daumé III |
Mach. Transl. | 1 |
| 2002 | A Noisy-Channel Model for Document CompressionabstractWe present a document compression system that uses a hierarchical noisy-channel model of text production. Our compression system first automatically derives the syntactic structure of each sentence and the overall discourse structure of the text given as input. The system then uses a statistical hierarchical model of text production in order to drop non-important syntactic and discourse constituents so as to generate coherent, grammatical document compressions of arbitrary length. The system outperforms both a baseline and a sentence-based compression system that operates by simplifying sequentially all sentences in a text. Our results support the claim that discourse knowledge plays an important role in document summarization. Hal Daumé III, Daniel Marcu |
ACL | 1 |
| 2002 | The Importance of Lexicalized Syntax Models for Natural Language Generation Tasks
Hal Daumé III, Kevin Knight, Irene Langkilde-Geary, Daniel Marcu, Kenji Yamada |
INLG | 1 |