EDBT 2026 Demo / reviewers in the wild / expert
Jacob Eisenstein
dblp:82/2305
· DBLP profile ↗
87ranked-venue papers
27as first author
17since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 67 · 19 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 16 · 8 first-authorDatabases, data management, data science and information retrieval · 7 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rewarding Progress: Scaling Automated Process Verifiers for LLM ReasoningabstractA promising approach for improving reasoning in large language models is to use process reward models (PRMs). PRMs provide feedback at each step of a multi-step reasoning trace, improving credit assignment over outcome reward models (ORMs) that only provide feedback at the final step. However, collecting dense, per-step human labels is not scalable, and training PRMs from automatically-labeled data has thus far led to limited gains. With the goal of using PRMs to improve a *base* policy via test-time search and reinforcement learning (RL), we ask: ``How should we design process rewards?'' Our key insight is that, to be effective, the process reward for a step should measure
*progress*: a change in the likelihood of producing a correct response in the future, before and after taking the step, as measured under a *prover* policy distinct from the base policy. Such progress values can {distinguish} good and bad steps generated by the base policy, even though the base policy itself cannot. Theoretically, we show that even weaker provers can improve the base policy, as long as they distinguish steps without being too misaligned with the base policy. Our results show that process rewards defined as progress under such provers improve the efficiency of exploration during test-time search and online RL. We empirically validate our claims by training **process advantage verifiers (PAVs)** to measure progress under such provers and show that compared to ORM, they are >8% more accurate, and 1.5-5x more compute-efficient. Equipped with these insights, our PAVs enable **one of the first results** showing a 6x gain in sample efficiency for a policy trained using online RL with PRMs vs. ORMs. Amrith Setlur, Chirag Nagpal, Adam Fisch, Xinyang Geng, Jacob Eisenstein, Rishabh Agarwal, Alekh Agarwal, Jonathan Berant, Aviral Kumar |
ICLR | 5 |
| 2025 | InfAlign: Inference-aware language model alignmentabstractLanguage model alignment is a critical step
in training modern generative language models.
Alignment targets to improve win rate of a sample
from the aligned model against the base model.
Today, we are increasingly using inference-time
algorithms (e.g., Best-of-$N$ , controlled decoding, tree search) to decode from language models
rather than standard sampling. We show that this
train/test mismatch makes standard RLHF framework sub-optimal in view of such inference-time
methods. To this end, we propose a framework for
inference-aware alignment (InfAlign), which
aims to optimize *inference-time win rate* of the
aligned policy against the base model. We prove
that for any inference-time decoding procedure,
the optimal aligned policy is the solution to the
standard RLHF problem with a *transformation*
of the reward. This motivates us to provide the
calibrate-and-transform RL (InfAlign-CTRL)
algorithm to solve this problem, which involves
a reward calibration step and a KL-regularized
reward maximization step with a transformation
of the calibrated reward. For best-of-$N$ sampling
and best-of-$N$ jailbreaking, we propose specific
transformations offering up to 3-8% improvement
on inference-time win rates. Finally, we also show
that our proposed reward calibration method is a
strong baseline for optimizing standard win rate. Ananth Balashankar, Ziteng Sun, Jonathan Berant, Jacob Eisenstein, Michael Collins 0001, Adrian Hutter, Jong Lee, Chirag Nagpal, Flavien Prost, Aradhana Sinha, Ananda Theertha Suresh, Ahmad Beirami |
ICML | 4 |
| 2025 | Theoretical guarantees on the best-of-n alignment policyabstractA simple and effective method for the inference-time alignment of generative models is the best-of-$n$ policy, where $n$ samples are drawn from a reference policy, ranked based on a reward function, and the highest ranking one is selected. A commonly used analytical expression in the literature claims that the KL divergence between the best-of-$n$ policy and the reference policy is equal to $\log (n) - (n-1)/n.$ We disprove the validity of this claim, and show that it is an upper bound on the actual KL divergence. We also explore the tightness of this upper bound in different regimes, and propose a new estimator for the KL divergence and empirically show that it provides a tight approximation. We also show that the win rate of the best-of-$n$ policy against the reference policy is upper bounded by $n/(n+1)$ and derive bounds on the tightness of this characterization. We conclude with analyzing the tradeoffs between win rate and KL divergence of the best-of-$n$ alignment policy, which demonstrate that very good tradeoffs are achievable with $n < 1000$. Ahmad Beirami, Alekh Agarwal, Jonathan Berant, Alexander D'Amour, Jacob Eisenstein, Chirag Nagpal, Ananda Theertha Suresh |
ICML | 5 |
| 2024 | Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual AlignmentabstractAligning language models (LMs) based on human-annotated preference data is a crucial step in obtaining practical and performant LMbased systems.However, multilingual human preference data are difficult to obtain at scale, making it challenging to extend this framework to diverse languages.In this work, we evaluate a simple approach for zero-shot crosslingual alignment, where a reward model is trained on preference data in one source language and directly applied to other target languages.On summarization and open-ended dialog generation, we show that this method is consistently successful under comprehensive evaluation settings, including human evaluation: cross-lingually aligned models are preferred by humans over unaligned models on up to >70% of evaluation instances.We moreover find that a different-language reward model sometimes yields better aligned models than a same-language reward model.We also identify best practices when there is no languagespecific data for even supervised finetuning, another component in alignment.en de en es en ru en tr en vi de en es en ru en tr en vi en 0 20 40 60 ROUGE-L (a) Summarization, unaligned SFT model Target-Language SFT Data Zhaofeng Wu, Ananth Balashankar, Jacob Eisenstein, Ahmad Beirami |
EMNLP | 4 |
| 2024 | Transforming and Combining Rewards for Aligning Large Language ModelsabstractA common approach for aligning language models to human preferences is to first learn a reward model from preference data, and then use this reward model to update the language model. We study two closely related problems that arise in this approach. First, any monotone transformation of the reward model preserves preference ranking; is there a choice that is "better" than others? Second, we often wish to align language models to multiple properties: how should we combine multiple reward models? Using a probabilistic interpretation of the alignment procedure, we identify a natural choice for transformation for (the common case of) rewards learned from Bradley-Terry preference models. The derived transformation is straightforward: we apply a log-sigmoid function to the centered rewards, a method we term "LSC-transformation" (log-sigmoid-centered transformation). This transformation has two important properties. First, it emphasizes improving poorly-performing outputs, rather than outputs that already score well. This mitigates both underfitting (where some prompts are not improved) and reward hacking (where the model learns to exploit misspecification of the reward model). Second, it enables principled aggregation of rewards by linking summation to logical conjunction: the sum of transformed rewards corresponds to the probability that the output is "good" in all measured properties, in a sense we make precise. Experiments aligning language models to be both helpful and harmless using RLHF show substantial improvements over the baseline (non-transformed) approach. Chirag Nagpal, Jonathan Berant, Jacob Eisenstein, Alexander D'Amour, Oluwasanmi Koyejo, Victor Veitch |
ICML | 4 |
| 2023 | Dialect-robust Evaluation of Generated TextabstractJiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, Sebastian Gehrmann. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, Sebastian Gehrmann |
ACL (1) | 8 |
| 2023 | Selectively Answering Ambiguous QuestionsabstractTrustworthy language models should abstain from answering questions when they do not know the answer.However, the answer to a question can be unknown for a variety of reasons.Prior research has focused on the case in which the question is clear and the answer is unambiguous but possibly unknown.But the answer to a question can also be unclear due to uncertainty of the questioner's intent or context.We investigate question answering from this perspective, focusing on answering a subset of questions with a high degree of accuracy, from a set of questions in which many are inherently ambiguous.In this setting, we find that the most reliable approach to decide when to abstain involves quantifying repetition within sampled model outputs, rather than the model's likelihood or self-verification as used in prior work.We find this to be the case across different types of uncertainty and model scales, and with or without instruction tuning.Our results suggest that sampling-based confidence scores help calibrate answers to relatively unambiguous questions, with more dramatic improvements on ambiguous questions. Jeremy R. Cole, Michael J. Q. Zhang, Daniel Gillick, Julian Martin Eisenschlos, Bhuwan Dhingra, Jacob Eisenstein |
EMNLP | 6 |
| 2023 | MD3: The Multi-Dialect Dataset of Dialogues
Jacob Eisenstein, Vinodkumar Prabhakaran, Clara Rivera, Dorottya Demszky, Devyani Sharma |
INTERSPEECH | 1 |
| 2022 | The MultiBERTs: BERT Reproductions for Robustness Analysis
Thibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei, Naomi Saphra, Alexander D'Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das 0001, Ellie Pavlick |
ICLR | 10 |
| 2022 | Informativeness and Invariance: Two Perspectives on Spurious Correlations in Natural LanguageabstractSpurious correlations are a threat to the trustworthiness of natural language processing systems, motivating research into methods for identifying and eliminating them.However, addressing the problem of spurious correlations requires more clarity on what they are and how they arise in language data.Gardner et al. (2021)argue that the compositional nature of language implies that all correlations between labels and individual "input features" are spurious.This paper analyzes this proposal in the context of a toy example, demonstrating three distinct conditions that can give rise to feature-label correlations in a simple PCFG.Linking the toy example to a structured causal model shows that (1) feature-label correlations can arise even when the label is invariant to interventions on the feature, and (2) feature-label correlations may be absent even when the label is sensitive to interventions on the feature.Because input features will be individually correlated with labels in all but very rare circumstances, domain knowledge must be applied to identify spurious correlations that pose genuine robustness threats. Jacob Eisenstein |
NAACL-HLT | 1 |
| 2022 | Underspecification Presents Challenges for Credibility in Modern Machine LearningabstractMachine learning (ML) systems often exhibit unexpectedly poor behavior when they are deployed in real-world domains. We identify underspecification in ML pipelines as a key reason for these failures. An ML pipeline is the full procedure followed to train and validate a predictor. Such a pipeline is underspecified when it can return many distinct predictors with equivalently strong test performance. Underspecification is common in modern ML pipelines that primarily validate predictors on held-out data that follow the same distribution as the training data. Predictors returned by underspecified pipelines are often treated as equivalent based on their training domain performance, but we show here that such predictors can behave very differently in deployment domains. This ambiguity can lead to instability and poor model behavior in practice, and is a distinct failure mode from previously identified issues arising from structural mismatch between training and deployment domains. We provide evidence that underspecfication has substantive implications for practical ML pipelines, using examples from computer vision, medical imaging, natural language processing, clinical risk prediction based on electronic health records, and medical genomics. Our results show the need to explicitly account for underspecification in modeling pipelines that are intended for real-world deployment in any domain. Alexander D'Amour, Katherine A. Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew Hoffman 0001, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yi-An Ma, Cory Y. McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Zachary Nado, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman 0003, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin G. Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Xuezhi Wang 0002, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai, D. Sculley |
J. Mach. Learn. Res. | 9 |
| 2022 | Time-Aware Language Models as Temporal Knowledge BasesabstractAbstract Many facts come with an expiration date, from the name of the President to the basketball team Lebron James plays for. However, most language models (LMs) are trained on snapshots of data collected at a specific moment in time. This can limit their utility, especially in the closed-book setting where the pretraining corpus must contain the facts the model should memorize. We introduce a diagnostic dataset aimed at probing LMs for factual knowledge that changes over time and highlight problems with LMs at either end of the spectrum—those trained on specific slices of temporal data, as well as those trained on a wide range of temporal data. To mitigate these problems, we propose a simple technique for jointly modeling text with its timestamp. This improves memorization of seen facts from the training time period, as well as calibration on predictions about unseen facts from future time periods. We also show that models trained with temporal context can be efficiently “refreshed” as new data arrives, without the need for retraining from scratch. Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, William W. Cohen |
Trans. Assoc. Comput. Linguistics | 5 |
| 2022 | Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and BeyondabstractAbstract A fundamental goal of scientific research is to learn about causal relationships. However, despite its critical role in the life and social sciences, causality has not had the same importance in Natural Language Processing (NLP), which has traditionally placed more emphasis on predictive tasks. This distinction is beginning to fade, with an emerging area of interdisciplinary research at the convergence of causal inference and language processing. Still, research on causality in NLP remains scattered across domains without unified definitions, benchmark datasets and clear articulations of the challenges and opportunities in the application of causal inference to the textual domain, with its unique properties. In this survey, we consolidate research across academic areas and situate it in the broader NLP landscape. We introduce the statistical challenge of estimating causal effects with text, encompassing settings where text is used as an outcome, treatment, or to address confounding. In addition, we explore potential uses of causal inference to improve the robustness, fairness, and interpretability of NLP models. We thus provide a unified overview of causal inference for the NLP community.1 Amir Feder, Katherine A. Keith, Emaad Manzoor, Reid Pryzant, Dhanya Sridhar, Zach Wood-Doughty, Jacob Eisenstein, Justin Grimmer, Roi Reichart, Margaret E. Roberts, Brandon M. Stewart, Victor Veitch, Diyi Yang |
Trans. Assoc. Comput. Linguistics | 7 |
| 2021 | Learning to Recognize Dialect FeaturesabstractDorottya Demszky, Devyani Sharma, Jonathan Clark, Vinodkumar Prabhakaran, Jacob Eisenstein. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Dorottya Demszky, Devyani Sharma, Jonathan H. Clark, Vinodkumar Prabhakaran, Jacob Eisenstein |
NAACL-HLT | 5 |
| 2021 | Counterfactual Invariance to Spurious Correlations in Text ClassificationabstractInformally, a 'spurious correlation' is the dependence of a model on some aspect of the input data that an analyst thinks shouldn't matter. In machine learning, these have a know-it-when-you-see-it character; e.g., changing the gender of a sentence's subject changes a sentiment predictor's output. To check for spurious correlations, we can 'stress test' models by perturbing irrelevant parts of input data and seeing if model predictions change. In this paper, we study stress testing using the tools of causal inference. We introduce counterfactual invariance as a formalization of the requirement that changing irrelevant parts of the input shouldn't change model predictions. We connect counterfactual invariance to out-of-domain model performance, and provide practical schemes for learning (approximately) counterfactual invariant predictors (without access to counterfactual examples). It turns out that both the means and implications of counterfactual invariance depend fundamentally on the true underlying causal structure of the data---in particular, whether the label causes the features or the features cause the label. Distinct causal structures require distinct regularization schemes to induce counterfactual invariance. Similarly, counterfactual invariance implies different domain shift guarantees depending on the underlying causal structure. This theory is supported by empirical results on text classification. Victor Veitch, Alexander D'Amour, Steve Yadlowsky, Jacob Eisenstein |
NeurIPS | 4 |
| 2021 | Follow the leader: Documents on the leading edge of semantic change get more citationsabstractAbstract Diachronic word embeddings—vector representations of words over time—offer remarkable insights into the evolution of language and provide a tool for quantifying sociocultural change from text documents. Prior work has used such embeddings to identify shifts in the meaning of individual words. However, simply knowing that a word has changed in meaning is insufficient to identify the instances of word usage that convey the historical meaning or the newer meaning. In this study, we link diachronic word embeddings to documents, by situating those documents as leaders or laggards with respect to ongoing semantic changes. Specifically, we propose a novel method to quantify the degree of semantic progressiveness in each word usage, and then show how these usages can be aggregated to obtain scores for each document. We analyze two large collections of documents, representing legal opinions and scientific articles. Documents that are scored as semantically progressive receive a larger number of citations, indicating that they are especially influential. Our work thus provides a new technique for identifying lexical semantic leaders and demonstrates a new link between progressive use of language and influence in a citation network. Sandeep Soni, Kristina Lerman, Jacob Eisenstein |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2021 | Sparse, Dense, and Attentional Representations for Text RetrievalabstractAbstract Dual encoders perform retrieval by encoding documents and queries into dense low-dimensional vectors, scoring each document by its inner product with the query. We investigate the capacity of this architecture relative to sparse bag-of-words models and attentional neural networks. Using both theoretical and empirical analysis, we establish connections between the encoding dimension, the margin between gold and lower-ranked documents, and the document length, suggesting limitations in the capacity of fixed-length encodings to support precise retrieval of long documents. Building on these insights, we propose a simple neural model that combines the efficiency of dual encoders with some of the expressiveness of more costly attentional architectures, and explore sparse-dense hybrids to capitalize on the precision of sparse retrieval. These models outperform strong alternatives in large-scale retrieval. Yi Luan, Jacob Eisenstein, Kristina Toutanova, Michael Collins 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | AdvAug: Robust Adversarial Augmentation for Neural Machine TranslationabstractIn this paper, we propose a new adversarial augmentation method for Neural Machine Translation (NMT).The main idea is to minimize the vicinal risk over virtual sentences sampled from two vicinity distributions, of which the crucial one is a novel vicinity distribution for adversarial sentences that describes a smooth interpolated embedding space centered around observed training sentence pairs.We then discuss our approach, AdvAug, to train NMT models using the embeddings of virtual sentences in sequence-tosequence learning.Experiments on Chinese-English, English-French, and English-German translation benchmarks show that AdvAug achieves significant improvements over the Transformer (up to 4.9 BLEU points), and substantially outperforms other data augmentation techniques (e.g.back-translation) without using extra corpora. Yong Cheng 0003, Lu Jiang 0004, Wolfgang Macherey, Jacob Eisenstein |
ACL | 4 |
| 2020 | Characterizing Collective Attention via Descriptor Context: A Case Study of Public Discussions of Crisis Events
Diyi Yang, Jacob Eisenstein |
ICWSM | 3 |
| 2019 | The Referential Reader: A Recurrent Entity Network for Anaphora ResolutionabstractWe present a new architecture for storing and accessing entity mentions during online text processing.While reading the text, entity references are identified, and may be stored by either updating or overwriting a cell in a fixedlength memory.The update operation implies coreference with the other mentions that are stored in the same cell; the overwrite operation causes these mentions to be forgotten.By encoding the memory operations as differentiable gates, it is possible to train the model end-to-end, using both a supervised anaphora resolution objective as well as a supplementary language modeling objective.Evaluation on a dataset of pronoun-name anaphora demonstrates strong performance with purely incremental text processing. Fei Liu 0023, Luke Zettlemoyer, Jacob Eisenstein |
ACL (1) | 3 |
| 2019 | Unsupervised Domain Adaptation of Contextualized Embeddings for Sequence LabelingabstractXiaochuang Han, Jacob Eisenstein. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiaochuang Han, Jacob Eisenstein |
EMNLP/IJCNLP (1) | 2 |
| 2018 | Predicting Semantic Relations using Global Graph PropertiesabstractSemantic graphs, such as WordNet, are resources which curate natural language on two distinguishable layers.On the local level, individual relations between synsets (semantic building blocks) such as hypernymy and meronymy enhance our understanding of the words used to express their meanings.Globally, analysis of graph-theoretic properties of the entire net sheds light on the structure of human language as a whole.In this paper, we combine global and local properties of semantic graphs through the framework of Max-Margin Markov Graph Models (M3GM), a novel extension of Exponential Random Graph Model (ERGM) that scales to large multi-relational graphs.We demonstrate how such global modeling improves performance on the local task of predicting semantic relations between synsets, yielding new state-ofthe-art results on the WN18RR dataset, a challenging version of WordNet link prediction in which "easy" reciprocal cases are removed.In addition, the M3GM model identifies multirelational motifs that are characteristic of wellformed lexical semantic ontologies. Yuval Pinter, Jacob Eisenstein |
EMNLP | 2 |
| 2018 | Making "fetch" happen: The influence of social and linguistic context on nonstandard word growth and declineabstractIn an online community, new words come and go: today's haha may be replaced by tomorrow's lol.Changes in online writing are usually studied as a social process, with innovations diffusing through a network of individuals in a speech community.But unlike other types of innovation, language change is shaped and constrained by the grammatical system in which it takes part.To investigate the role of social and structural factors in language change, we undertake a large-scale analysis of the frequencies of nonstandard words in Reddit.Dissemination across many linguistic contexts is a predictor of success: words that appear in more linguistic contexts grow faster and survive longer.Furthermore, social dissemination plays a less important role in explaining word growth and decline than previously hypothesized. Jacob Eisenstein |
EMNLP | 2 |
| 2018 | Explainable Prediction of Medical Codes from Clinical TextabstractJames Mullenbach, Sarah Wiegreffe, Jon Duke, Jimeng Sun, Jacob Eisenstein. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. James Mullenbach, Sarah Wiegreffe, Jon Duke, Jimeng Sun 0001, Jacob Eisenstein |
NAACL-HLT | 5 |
| 2018 | Interactional Stancetaking in Online ForumsabstractLanguage is shaped by the relationships between the speaker/writer and the audience, the object of discussion, and the talk itself. In turn, language is used to reshape these relationships over the course of an interaction. Computational researchers have succeeded in operationalizing sentiment, formality, and politeness, but each of these constructs captures only some aspects of social and relational meaning. Theories of interactional stancetaking have been put forward as holistic accounts, but until now, these theories have been applied only through detailed qualitative analysis of (portions of) a few individual conversations. In this article, we propose a new computational operationalization of interpersonal stancetaking. We begin with annotations of three linked stance dimensions—affect, investment, and alignment—on 68 conversation threads from the online platform Reddit. Using these annotations, we investigate thread structure and linguistic properties of stancetaking in online conversations. We identify lexical features that characterize the extremes along each stancetaking dimension, and show that these stancetaking properties can be predicted with moderate accuracy from bag-of-words features, even with a relatively small labeled training set. These quantitative analyses are supplemented by extensive qualitative analysis, highlighting the compatibility of computational and qualitative methods in synthesizing evidence about the creation of interactional meaning. Scott F. Kiesling, Umashanthi Pavalanathan, Jim Fitzpatrick, Xiaochuang Han, Jacob Eisenstein |
Comput. Linguistics | 5 |
| 2018 | The Internet's Hidden Rules: An Empirical Study of Reddit Norm Violations at Micro, Meso, and Macro ScalesabstractNorms are central to how online communities are governed. Yet, norms are also emergent, arise from interaction, and can vary significantly between communities---making them challenging to study at scale. In this paper, we study community norms on Reddit in a large-scale, empirical manner. Via 2.8M comments removed by moderators of 100 top subreddits over 10 months, we use both computational and qualitative methods to identify three types of norms: macro norms that are universal to most parts of Reddit; meso norms that are shared across certain groups of subreddits; and micro norms that are specific to individual, relatively unique subreddits. Given the size of Reddit's user base---and the wide range of topics covered by different subreddits---we argue this represents the first large-scale census of the norms in broader internet culture. In other words, these findings shed light on what Reddit values, and how widely-held those values are. We conclude by discussing implications for the design of new and existing online communities. Eshwar Chandrasekharan, Mattia Samory, Shagun Jhaver, Hunter Charvat, Amy S. Bruckman, Cliff Lampe, Jacob Eisenstein, Eric Gilbert |
Proc. ACM Hum. Comput. Interact. | 7 |
| 2018 | Mind Your POV: Convergence of Articles and Editors Towards Wikipedia's Neutrality NormabstractWikipedia has a strong norm of writing in a "neutral point of view" (NPOV). Articles that violate this norm are tagged, and editors are encouraged to make corrections. But the impact of this tagging system has not been quantitatively measured. Does NPOV tagging help articles to converge to the desired style? Do NPOV corrections encourage editors to adopt this style? We study these questions using a corpus of NPOV-tagged articles and a set of lexicons associated with biased language. An interrupted time series analysis shows that after an article is tagged for NPOV, there is a significant decrease in biased language in the article, as measured by several lexicons. However, for individual editors, NPOV corrections and talk page discussions yield no significant change in the usage of words in most of these lexicons, including Wikipedia's own list of "words to watch." This suggests that NPOV tagging and discussion does improve content, but has less success enculturating editors to the site's linguistic norms. Umashanthi Pavalanathan, Xiaochuang Han, Jacob Eisenstein |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2017 | Unsupervised Learning for Lexicon-Based ClassificationabstractIn lexicon-based classification, documents are assigned labels by comparing the number of words that appear from two opposed lexicons, such as positive and negative sentiment. Creating such words lists is often easier than labeling instances, and they can be debugged by non-experts if classification performance is unsatisfactory. However, there is little analysis or justification of this classification heuristic. This paper describes a set of assumptions that can be used to derive a probabilistic justification for lexicon-based classification, as well as an analysis of its expected accuracy. One key assumption behind lexicon-based classification is that all words in each lexicon are equally predictive. This is rarely true in practice, which is why lexicon-based approaches are usually outperformed by supervised classifiers that learn distinct weights on each word from labeled instances. This paper shows that it is possible to learn such weights without labeled data, by leveraging co-occurrence statistics across the lexicons. This offers the best of both worlds: light supervision in the form of lexicons, and data-driven classification with higher accuracy than traditional word-counting heuristics. Jacob Eisenstein |
AAAI | 1 |
| 2017 | A Multidimensional Lexicon for Interpersonal StancetakingabstractThe sociolinguistic construct of stancetaking describes the activities through which discourse participants create and signal relationships to their interlocutors, to the topic of discussion, and to the talk itself.Stancetaking underlies a wide range of interactional phenomena, relating to formality, politeness, affect, and subjectivity.We present a computational approach to stancetaking, in which we build a theoretically-motivated lexicon of stance markers, and then use multidimensional analysis to identify a set of underlying stance dimensions.We validate these dimensions intrinsically and extrinsically, showing that they are internally coherent, match pre-registered hypotheses, and correlate with social phenomena. Umashanthi Pavalanathan, Jim Fitzpatrick, Scott F. Kiesling, Jacob Eisenstein |
ACL (1) | 4 |
| 2017 | #Anorexia, #anarexia, #anarexyia: Characterizing online community practices with orthographic variationabstractDistinctive linguistic practices help communities build solidarity and differentiate themselves from outsiders. In an online community, one such practice is variation in orthography, which includes spelling, punctuation, and capitalization. Using a dataset of over two million Instagram posts, we investigate orthographic variation in a community that shares pro-eating disorder (pro-ED) content. We find that not only does orthographic variation grow more frequent over time, it also becomes more profound or “deep,” with variants becoming increasingly distant from the original: as, for example, #anarexyia is more distant than #anarexia from the original spelling #anorexia. We find that the these changes are driven by newcomers, who adopt the most extreme linguistic practices as they enter the community. Moreover, this behavior correlates with engagement with the community: the newcomers that adopt deeper variant orthography tend to remain active for longer in the community, and posts with deeper variation receive more positive feedback in the form of “likes.” Previous work has linked community membership change with language change, and our work casts this connection in a new light, with newcomers driving an evolving practice rather than adapting to it. We also demonstrate the utility of orthographic variation as a new lens to study sociolinguistic change in online communities, particularly when the change results from an exogenous force such as a content ban. Stevie Chancellor, Munmun De Choudhury, Jacob Eisenstein |
IEEE BigData | 4 |
| 2017 | Mimicking Word Embeddings using Subword RNNsabstractWord embeddings improve generalization over lexical features by placing each word in a lower-dimensional space, using distributional information obtained from unlabeled data.However, the effectiveness of word embeddings for downstream NLP tasks is limited by out-of-vocabulary (OOV) words, for which embeddings do not exist.In this paper, we present MIM-ICK, an approach to generating OOV word embeddings compositionally, by learning a function from spellings to distributional embeddings.Unlike prior work, MIMICK does not require re-training on the original word embedding corpus; instead, learning is performed at the type level.Intrinsic and extrinsic evaluations demonstrate the power of this simple approach.On 23 languages, MIMICK improves performance over a word-based baseline for tagging part-of-speech and morphosyntactic attributes.It is competitive with (and complementary to) a supervised characterbased model in low-resource settings. Yuval Pinter, Robert Guthrie, Jacob Eisenstein |
EMNLP | 3 |
| 2017 | A Kernel Independence Test for Geographical Language VariationabstractQuantifying the degree of spatial dependence for linguistic variables is a key task for analyzing dialectal variation. However, existing approaches have important drawbacks. First, they are based on parametric models of dependence, which limits their power in cases where the underlying parametric assumptions are violated. Second, they are not applicable to all types of linguistic data: Some approaches apply only to frequencies, others to boolean indicators of whether a linguistic variable is present. We present a new method for measuring geographical language variation, which solves both of these problems. Our approach builds on Reproducing Kernel Hilbert Space (RKHS) representations for nonparametric statistics, and takes the form of a test statistic that is computed from pairs of individual geotagged observations without aggregation into predefined geographical bins. We compare this test with prior work using synthetic data as well as a diverse set of real data sets: a corpus of Dutch tweets, a Dutch syntactic atlas, and a data set of letters to the editor in North American newspapers. Our proposed test is shown to support robust inferences across a broad range of scenarios and types of data. Dong Nguyen 0002, Jacob Eisenstein |
Comput. Linguistics | 2 |
| 2017 | You Can't Stay Here: The Efficacy of Reddit's 2015 Ban Examined Through Hate SpeechabstractIn 2015, Reddit closed several subreddits-foremost among them r/fatpeoplehate and r/CoonTown-due to violations of Reddit's anti-harassment policy. However, the effectiveness of banning as a moderation approach remains unclear: banning might diminish hateful behavior, or it may relocate such behavior to different parts of the site. We study the ban of r/fatpeoplehate and r/CoonTown in terms of its effect on both participating users and affected subreddits. Working from over 100M Reddit posts and comments, we generate hate speech lexicons to examine variations in hate speech usage via causal inference methods. We find that the ban worked for Reddit. More accounts than expected discontinued using the site; those that stayed drastically decreased their hate speech usage-by at least 80%. Though many subreddits saw an influx of r/fatpeoplehate and r/CoonTown "migrants," those subreddits saw no significant changes in hate speech usage. In other words, other subreddits did not inherit the problem. We conclude by reflecting on the apparent success of the ban, discussing implications for online moderation, Reddit and internet communities more broadly. Eshwar Chandrasekharan, Umashanthi Pavalanathan, Anirudh Srinivasan, Adam Glynn, Jacob Eisenstein, Eric Gilbert |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2017 | Overcoming Language Variation in Sentiment Analysis with Social AttentionabstractVariation in language is ubiquitous, particularly in newer forms of writing such as social media. Fortunately, variation is not random; it is often linked to social properties of the author. In this paper, we show how to exploit social networks to make sentiment analysis more robust to social language variation. The key idea is linguistic homophily: the tendency of socially linked individuals to use language in similar ways. We formalize this idea in a novel attention-based neural network architecture, in which attention is divided among several basis models, depending on the author’s position in the social network. This has the effect of smoothing the classification function across the social network, and makes it possible to induce personalized classifiers even for authors for whom there is no labeled data or demographic metadata. This model significantly improves the accuracies of sentiment analysis on Twitter and on review data. Yi Yang 0038, Jacob Eisenstein |
Trans. Assoc. Comput. Linguistics | 2 |
| 2016 | Morphological Priors for Probabilistic Neural Word EmbeddingsabstractWord embeddings allow natural language processing systems to share statistical information across related words.These embeddings are typically based on distributional statistics, making it difficult for them to generalize to rare or unseen words.We propose to improve word embeddings by incorporating morphological information, capturing shared sub-word features.Unlike previous work that constructs word embeddings directly from morphemes, we combine morphological and distributional information in a unified probabilistic framework, in which the word embedding is a latent variable.The morphological information provides a prior distribution on the latent word embeddings, which in turn condition a likelihood function over an observed corpus.This approach yields improvements on intrinsic word similarity evaluations, and also in the downstream task of part-of-speech tagging. Parminder Bhatia, Robert Guthrie, Jacob Eisenstein |
EMNLP | 3 |
| 2016 | Toward Socially-Infused Information Extraction: Embedding Authors, Mentions, and EntitiesabstractEntity linking is the task of identifying mentions of entities in text, and linking them to entries in a knowledge base. This task is especially difficult in microblogs, as there is little additional text to provide disambiguating context; rather, authors rely on an implicit common ground of shared knowledge with their readers. In this paper, we attempt to capture some of this implicit context by exploiting the social network structure in microblogs. We build on the theory of homophily, which implies that socially linked individuals share interests, and are therefore likely to mention the same sorts of entities. We implement this idea by encoding authors, mentions, and entities in a continuous vector space, which is constructed so that socially-connected authors have similar vector representations. These vectors are incorporated into a neural structured prediction model, which captures structural constraints that are inherent in the entity linking task. Together, these design decisions yield F1 improvements of 1%-5% on benchmark datasets, as compared to the previous state-of-the-art. Yi Yang 0038, Ming-Wei Chang, Jacob Eisenstein |
EMNLP | 3 |
| 2016 | A Latent Variable Recurrent Neural Network for Discourse-Driven Language ModelsabstractThis paper presents a novel latent variable recurrent neural network architecture for jointly modeling sequences of words and (possibly latent) discourse relations between adjacent sentences.A recurrent neural network generates individual words, thus reaping the benefits of discriminatively-trained vector representations.The discourse relations are represented with a latent variable, which can be predicted or marginalized, depending on the task.The resulting model can therefore employ a training objective that includes not only discourse relation classification, but also word prediction.As a result, it outperforms state-ofthe-art alternatives for two tasks: implicit discourse relation classification in the Penn Discourse Treebank, and dialog act classification in the Switchboard corpus.Furthermore, by marginalizing over latent discourse relations at test time, we obtain a discourse informed language model, which improves over a strong LSTM baseline. Yangfeng Ji, Gholamreza Haffari, Jacob Eisenstein |
HLT-NAACL | 3 |
| 2016 | Part-of-Speech Tagging for Historical EnglishabstractAs more historical texts are digitized, there is interest in applying natural language processing tools to these archives.However, the performance of these tools is often unsatisfactory, due to language change and genre differences.Spelling normalization heuristics are the dominant solution for dealing with historical texts, but this approach fails to account for changes in usage and vocabulary.In this empirical paper, we assess the capability of domain adaptation techniques to cope with historical texts, focusing on the classic benchmark task of part-of-speech tagging.We evaluate several domain adaptation methods on the task of tagging Early Modern English and Modern British English texts in the Penn Corpora of Historical English.We demonstrate that the Feature Embedding method for unsupervised domain adaptation outperforms word embeddings and Brown clusters, showing the importance of embedding the entire feature space, rather than just individual words.Feature Embeddings also give better performance than spelling normalization, but the combination of the two methods is better still, yielding a 5% raw improvement in tagging accuracy on Early Modern English texts. Yi Yang 0038, Jacob Eisenstein |
HLT-NAACL | 2 |
| 2015 | Better Document-level Sentiment Analysis from RST Discourse ParsingabstractDiscourse structure is the hidden link between surface features and document-level properties, such as sentiment polarity.We show that the discourse analyses produced by Rhetorical Structure Theory (RST) parsers can improve document-level sentiment analysis, via composition of local information up the discourse tree.First, we show that reweighting discourse units according to their position in a dependency representation of the rhetorical structure can yield substantial improvements on lexicon-based sentiment analysis.Next, we present a recursive neural network over the RST structure, which offers significant improvements over classificationbased methods. Parminder Bhatia, Yangfeng Ji, Jacob Eisenstein |
EMNLP | 3 |
| 2015 | Closing the Gap: Domain Adaptation from Explicit to Implicit Discourse RelationsabstractMany discourse relations are explicitly marked with discourse connectives, and these examples could potentially serve as a plentiful source of training data for recognizing implicit discourse relations.However, there are important linguistic differences between explicit and implicit discourse relations, which limit the accuracy of such an approach.We account for these differences by applying techniques from domain adaptation, treating implicitly and explicitly-marked discourse relations as separate domains.The distribution of surface features varies across these two domains, so we apply a marginalized denoising autoencoder to induce a dense, domain-general representation.The label distribution is also domain-specific, so we apply a resampling technique that is similar to instance weighting.In combination with a set of automatically-labeled data, these improvements eliminate more than 80% of the transfer loss incurred by training an implicit discourse relation classifier on explicitly-marked discourse relations. Yangfeng Ji, Jacob Eisenstein |
EMNLP | 3 |
| 2015 | Confounds and Consequences in Geotagged Twitter DataabstractTwitter is often used in quantitative studies that identify geographically-preferred topics, writing styles, and entities.These studies rely on either GPS coordinates attached to individual messages, or on the user-supplied location field in each profile.In this paper, we compare these data acquisition techniques and quantify the biases that they introduce; we also measure their effects on linguistic analysis and textbased geolocation.GPS-tagging and selfreported locations yield measurably different corpora, and these linguistic differences are partially attributable to differences in dataset composition by age and gender.Using a latent variable model to induce age and gender, we show how these demographic variables interact with geography to affect language use.We also show that the accuracy of text-based geolocation varies with population demographics, giving the best results for men above the age of 40. Umashanthi Pavalanathan, Jacob Eisenstein |
EMNLP | 2 |
| 2015 | Psychological Effects of Urban Crime Gleaned from Social Media
Jose Manuel Delgado Valdes, Jacob Eisenstein, Munmun De Choudhury |
ICWSM | 2 |
| 2015 | "You're Mr. Lebowski, I'm the Dude": Inducing Address Term Formality in Signed Social NetworksabstractWe present an unsupervised model for inducing signed social networks from the content exchanged across network edges.Inference in this model solves three problems simultaneously: (1) identifying the sign of each edge;(2) characterizing the distribution over content for each edge type; (3) estimating weights for triadic features that map to theoretical models such as structural balance.We apply this model to the problem of inducing the social function of address terms, such as Madame, comrade, and dude.On a dataset of movie scripts, our system obtains a coherent clustering of address terms, while at the same time making intuitively plausible judgments of the formality of social relations in each film.As an additional contribution, we provide a bootstrapping technique for identifying and tagging address terms in dialogue.1 Vinodh Krishnan Elangovan, Jacob Eisenstein |
HLT-NAACL | 2 |
| 2015 | Unsupervised Multi-Domain Adaptation with Feature EmbeddingsabstractRepresentation learning is the dominant technique for unsupervised domain adaptation, but existing approaches have two major weaknesses.First, they often require the specification of "pivot features" that generalize across domains, which are selected by taskspecific heuristics.We show that a novel but simple feature embedding approach provides better performance, by exploiting the feature template structure common in NLP problems.Second, unsupervised domain adaptation is typically treated as a task of moving from a single source to a single target domain.In reality, test data may be diverse, relating to the training data in some ways but not others.We propose an alternative formulation, in which each instance has a vector of domain attributes, can be used to learn distill the domain-invariant properties of each feature. 1 Yi Yang 0038, Jacob Eisenstein |
HLT-NAACL | 2 |
| 2015 | Identifying visual attributes for object recognition from text and taxonomy
Caglar Tirkaz, Jacob Eisenstein, Tevfik Metin Sezgin, Berrin A. Yanikoglu |
Comput. Vis. Image Underst. | 2 |
| 2015 | One Vector is Not Enough: Entity-Augmented Distributed Semantics for Discourse RelationsabstractDiscourse relations bind smaller linguistic units into coherent texts. Automatically identifying discourse relations is difficult, because it requires understanding the semantics of the linked arguments. A more subtle challenge is that it is not enough to represent the meaning of each argument of a discourse relation, because the relation may depend on links between lowerlevel components, such as entity mentions. Our solution computes distributed meaning representations for each discourse argument by composition up the syntactic parse tree. We also perform a downward compositional pass to capture the meaning of coreferent entity mentions. Implicit discourse relations are then predicted from these two representations, obtaining substantial improvements on the Penn Discourse Treebank. Yangfeng Ji, Jacob Eisenstein |
Trans. Assoc. Comput. Linguistics | 2 |
| 2014 | Representation Learning for Text-level Discourse ParsingabstractText-level discourse parsing is notoriously difficult, as distinctions between discourse relations require subtle semantic judg-ments that are not easily captured using standard features. In this paper, we present a representation learning approach, in which we transform surface features into a latent space that facilitates RST dis-course parsing. By combining the machin-ery of large-margin transition-based struc-tured prediction with representation learn-ing, our method jointly learns to parse dis-course while at the same time learning a discourse-driven projection of surface fea-tures. The resulting shift-reduce discourse parser obtains substantial improvements over the previous state-of-the-art in pre-dicting relations and nuclearity on the RST Treebank. 1 Yangfeng Ji, Jacob Eisenstein |
ACL (1) | 2 |
| 2013 | Discriminative Improvements to Distributional Sentence SimilarityabstractMatrix and tensor factorization have been applied to a number of semantic relatedness tasks, including paraphrase identification.The key idea is that similarity in the latent space implies semantic relatedness.We describe three ways in which labeled data can improve the accuracy of these approaches on paraphrase classification.First, we design a new discriminative term-weighting metric called TF-KLD, which outperforms TF-IDF.Next, we show that using the latent representation from matrix factorization as features in a classification algorithm substantially improves accuracy.Finally, we combine latent features with fine-grained n-gram overlap features, yielding performance that is 3% more accurate than the prior state-of-the-art. Yangfeng Ji, Jacob Eisenstein |
EMNLP | 2 |
| 2013 | A Log-Linear Model for Unsupervised Text NormalizationabstractWe present a unified unsupervised statistical model for text normalization.The relationship between standard and non-standard tokens is characterized by a log-linear model, permitting arbitrary features.The weights of these features are trained in a maximumlikelihood framework, employing a novel sequential Monte Carlo training algorithm to overcome the large label space, which would be impractical for traditional dynamic programming solutions.This model is implemented in a normalization system called UNLOL, which achieves the best known results on two normalization datasets, outperforming more complex systems.We use the output of UNLOL to automatically normalize a large corpus of social media text, revealing a set of coherent orthographic styles that underlie online language variation. Yi Yang 0038, Jacob Eisenstein |
EMNLP | 2 |
| 2013 | What to do about bad language on the internet
Jacob Eisenstein |
HLT-NAACL | 1 |
| 2013 | Discourse Connectors for Latent Subjectivity in Sentiment Analysis
Rakshit S. Trivedi, Jacob Eisenstein |
HLT-NAACL | 2 |
| 2013 | Automated text mining for requirements analysis of policy documentsabstractBusinesses and organizations in jurisdictions around the world are required by law to provide their customers and users with information about their business practices in the form of policy documents. Requirements engineers analyze these documents as sources of requirements, but this analysis is a time-consuming and mostly manual process. Moreover, policy documents contain legalese and present readability challenges to requirements engineers seeking to analyze them. In this paper, we perform a large-scale analysis of 2,061 policy documents, including policy documents from the Google Top 1000 most visited websites and the Fortune 500 companies, for three purposes: (1) to assess the readability of these policy documents for requirements engineers; (2) to determine if automated text mining can indicate whether a policy document contains requirements expressed as either privacy protections or vulnerabilities; and (3) to establish the generalizability of prior work in the identification of privacy protections and vulnerabilities from privacy policies to other policy documents. Our results suggest that this requirements analysis technique, developed on a small set of policy documents in two domains, may generalize to other domains. Aaron K. Massey, Jacob Eisenstein, Annie I. Antón, Peter P. Swire |
RE | 2 |
| 2012 | Bootstrapping a Unified Model of Lexical and Phonetic Acquisition
Micha Elsner, Sharon Goldwater, Jacob Eisenstein |
ACL (1) | 3 |
| 2012 | Document hierarchies from text and linksabstractHierarchical taxonomies provide a multi-level view of large document collections, allowing users to rapidly drill down to fine-grained distinctions in topics of interest. We show that automatically induced taxonomies can be made more robust by combining text with relational links. The underlying mechanism is a Bayesian generative model in which a latent hierarchical structure explains the observed data --- thus, finding hierarchical groups of documents with similar word distributions and dense network connections. As a nonparametric Bayesian model, our approach does not require pre-specification of the branching factor at each non-terminal, but finds the appropriate level of detail directly from the data. Unlike many prior latent space models of network structure, the complexity of our approach does not grow quadratically in the number of documents, enabling application to networks with more than ten thousand nodes. Experimental results on hypertext and citation network corpora demonstrate the advantages of our hierarchical, multimodal approach. Qirong Ho, Jacob Eisenstein, Eric P. Xing |
WWW | 2 |
| 2011 | Discovering Sociolinguistic Associations with Structured Sparsity
Jacob Eisenstein, Noah A. Smith, Eric P. Xing |
ACL | 1 |
| 2011 | Sparse Additive Generative Models of Text
Jacob Eisenstein, Amr Ahmed 0001, Eric P. Xing |
ICML | 1 |
| 2011 | Unified analysis of streaming newsabstractNews clustering, categorization and analysis are key components of any news portal. They require algorithms capable of dealing with dynamic data to cluster, interpret and to temporally aggregate news articles. These three tasks are often solved separately. In this paper we present a unified framework to group incoming news articles into temporary but tightly-focused storylines, to identify prevalent topics and key entities within these stories, and to reveal the temporal structure of stories as they evolve. We achieve this by building a hybrid clustering and topic model. To deal with the available wealth of data we build an efficient parallel inference algorithm by sequential Monte Carlo estimation. Time and memory costs are nearly constant in the length of the history, and the approach scales to hundreds of thousands of documents. We demonstrate the efficiency and accuracy on the publicly available TDT dataset and data of a major internet news site. Amr Ahmed 0001, Qirong Ho, Jacob Eisenstein, Eric P. Xing, Alexander J. Smola, Choon Hui Teo |
WWW | 3 |
| 2010 | A Latent Variable Model for Geographic Lexical Variation
Jacob Eisenstein, Brendan T. O'Connor 0001, Noah A. Smith, Eric P. Xing |
EMNLP | 1 |
| 2009 | Reading to Learn: Constructing Features from Semantic Abstracts
Jacob Eisenstein, James Clarke, Dan Goldwasser, Dan Roth 0001 |
EMNLP | 1 |
| 2009 | Hierarchical Text Segmentation from Multi-Scale Lexical Cohesion
Jacob Eisenstein |
HLT-NAACL | 1 |
| 2009 | Adding More Languages Improves Unsupervised Multilingual Part-of-Speech Tagging: a Bayesian Non-Parametric Approach
Benjamin Snyder, Tahira Naseem, Jacob Eisenstein, Regina Barzilay |
HLT-NAACL | 3 |
| 2009 | Learning Document-Level Semantic Properties from Free-Text AnnotationsabstractThis paper presents a new method for inferring the semantic properties of documents by leveraging free-text keyphrase annotations. Such annotations are becoming increasingly abundant due to the recent dramatic growth in semi-structured, user-generated online content. One especially relevant domain is product reviews, which are often annotated by their authors with pros/cons keyphrases such as ``a real bargain'' or ``good value.'' These annotations are representative of the underlying semantic properties; however, unlike expert annotations, they are noisy: lay authors may use different labels to denote the same property, and some labels may be missing. To learn using such noisy annotations, we find a hidden paraphrase structure which clusters the keyphrases. The paraphrase structure is linked with a latent topic model of the review texts, enabling the system to predict the properties of unannotated documents and to effectively aggregate the semantic properties of multiple reviews. Our approach is implemented as a hierarchical Bayesian model with joint inference. We find that joint inference increases the robustness of the keyphrase clustering and encourages the latent topics to correlate with semantically meaningful properties. Multiple evaluations demonstrate that our model substantially outperforms alternative approaches for summarizing single and multiple documents into a set of semantically salient keyphrases. S. R. K. Branavan, Harr Chen, Jacob Eisenstein, Regina Barzilay |
J. Artif. Intell. Res. | 3 |
| 2009 | Multilingual Part-of-Speech Tagging: Two Unsupervised ApproachesabstractWe demonstrate the effectiveness of multilingual learning for unsupervised part-of-speech tagging. The central assumption of our work is that by combining cues from multiple languages, the structure of each becomes more apparent. We consider two ways of applying this intuition to the problem of unsupervised part-of-speech tagging: a model that directly merges tag structures for a pair of languages into a single sequence and a second model which instead incorporates multilingual context using latent variables. Both approaches are formulated as hierarchical Bayesian models, using Markov Chain Monte Carlo sampling techniques for inference. Our results demonstrate that by incorporating multilingual evidence we can achieve impressive performance gains across a range of scenarios. We also found that performance improves steadily as the number of available languages increases. Tahira Naseem, Benjamin Snyder, Jacob Eisenstein, Regina Barzilay |
J. Artif. Intell. Res. | 3 |
| 2008 | Discourse Topic and Gestural Form
Jacob Eisenstein, Regina Barzilay, Randall Davis |
AAAI | 1 |
| 2008 | Learning Document-Level Semantic Properties from Free-Text Annotations
S. R. K. Branavan, Harr Chen, Jacob Eisenstein, Regina Barzilay |
ACL | 3 |
| 2008 | Gestural Cohesion for Topic Segmentation
Jacob Eisenstein, Regina Barzilay, Randall Davis |
ACL | 1 |
| 2008 | Bayesian Unsupervised Topic Segmentation
Jacob Eisenstein, Regina Barzilay |
EMNLP | 1 |
| 2008 | Unsupervised Multilingual Learning for POS Tagging
Benjamin Snyder, Tahira Naseem, Jacob Eisenstein, Regina Barzilay |
EMNLP | 3 |
| 2008 | Gesture Salience as a Hidden Variable for Coreference Resolution and Keyframe ExtractionabstractGesture is a non-verbal modality that can contribute crucial information to the understanding of natural language. But not all gestures are informative, and non-communicative hand motions may confuse natural language processing (NLP) and impede learning. People have little diffculty ignoring irrelevant hand movements and focusing on meaningful gestures, suggesting that an automatic system could also be trained to perform this task. However, the informativeness of a gesture is context-dependent and labeling enough data to cover all cases would be expensive. We present conditional modality fusion, a conditional hidden-variable model that learns to predict which gestures are salient for coreference resolution, the task of determining whether two noun phrases refer to the same semantic entity. Moreover, our approach uses only coreference annotations, and not annotations of gesture salience itself. We show that gesture features improve performance on coreference resolution, and that by attending only to gestures that are salient, our method achieves further significant gains. In addition, we show that the model of gesture salience learned in the context of coreference accords with human intuition, by demonstrating that gestures judged to be salient by our model can be used successfully to create multimedia keyframe summaries of video. These summaries are similar to those created by human raters, and significantly outperform summaries produced by baselines from the literature. Jacob Eisenstein, Regina Barzilay, Randall Davis |
J. Artif. Intell. Res. | 1 |
| 2007 | Turning Lectures into Comic Books Using Linguistically Salient Gestures
Jacob Eisenstein, Regina Barzilay, Randall Davis |
AAAI | 1 |
| 2007 | Conditional Modality Fusion for Coreference Resolution
Jacob Eisenstein, Randall Davis |
ACL | 1 |
| 2006 | Interacting with communication appliances: an evaluation of two computer vision-based selection techniquesabstractCommunication appliances, intended for home settings, require intuitive forms of interaction. Computer vision offers a potential solution, but is not yet sufficiently accurate.As interaction designers, we need to know more than the absolute accuracy of such techniques: we must also be able to compare how they will work in our design settings, especially if we allow users to collaborate in the interpretation of their actions. We conducted a 2x4 within-subjects experiment to compare two interaction techniques based on computer vision: motion sensing, with EyeToy®-like feedback, and object tracking. Both techniques were 100% accurate with 2 or 5 choices. With 21 choices, object-tracking had significantly fewer errors and took less time for an accurate selection. Participants' subjective preferences were divided equally between the two techniques. This study compares these techniques as they would be used in real-world applications, with integrated user feedback, allowing interface designers to choose the one that best suits the specific user requirements for their particular application. Jacob Eisenstein, Wendy E. Mackay |
CHI | 1 |
| 2006 | Semantic Back-Pointers from Gesture
Jacob Eisenstein |
HLT-NAACL | 1 |
| 2006 | Gesture Improves Coreference Resolution
Jacob Eisenstein, Randall Davis |
HLT-NAACL | 1 |
| 2004 | Gestural cues for speech understandingabstractNo abstract available. Jacob Eisenstein |
ICMI | 1 |
| 2004 | Visual and linguistic information in gesture classificationabstractClassification of natural hand gestures is usually approached by applying pattern recognition to the movements of the hand. However, the gesture categories most frequently cited in the psychology literature are fundamentally multimodal; the definitions make reference to the surrounding linguistic context. We address the question of whether gestures are naturally multimodal, or whether they can be classified from hand-movement data alone. First, we describe an empirical study showing that the removal of auditory information significantly impairs the ability of human raters to classify gestures. Then we present an automatic gesture classification system based solely on an n-gram model of linguistic context; the system is intended to supplement a visual classifier, but achieves 66% accuracy on a three-class classification problem on its own. This represents higher accuracy than human raters achieve when presented with the same information. Jacob Eisenstein, Randall Davis |
ICMI | 1 |
| 2004 | A Salience-Based Approach to Gesture-Speech Alignment
Jacob Eisenstein, C. Mario Christoudias |
HLT-NAACL | 1 |
| 2003 | Device Independence and Extensibility in Gesture RecognitionabstractGesture recognition techniques often suffer from being highly device-dependent and hard to extend. If a system is trained using data from a specific glove input device, that system is typically unusable with any other input device. The set of gestures that a system is trained to recognize is typically not extensible, without retraining the entire system. We propose a novel gesture recognition framework to address these problems. This framework is based on a multi-layered view of gesture recognition. Only the lowest layer is device dependent, it converts raw sensor values produced by the glove to a glove-independent semantic description of the hand. The higher layers of our framework can be reused across gloves, and are easily extensible to include new gestures. We have experimentally evaluated our framework and found that it yields comparable performance to conventional techniques, while substantiating our claims of device independence and extensibility. Jacob Eisenstein, Shahram Ghandeharizadeh, Leana Golubchik, Cyrus Shahabi, Donghui Yan, Roger Zimmermann |
VR | 1 |
| 2002 | Agents and GUIs from task modelsabstractThis work unifies two important threads of research in intelligent user interfaces which share the common element of explicit task modeling. On the one hand, longstanding research on task-centered GUI design (sometimes called model-based design) has explored the benefits of explicitly modeling the task to be performed by an interface and using this task model as an integral part of the interface design process. More recently, research on collaborative interface agents has shown how an explicit task model can be used to control the behavior of a software agent that helps a user perform tasks using a GUI. This paper describes a collection of tools we have implemented which generate both a GUI and a collaborative interface agent from the same task model. Our task-centered GUI design tool incorporates a number of novel features which help the designer to integrate the task model into the design process without being unduly distracted. Our implementation of collaborative interface agents is built on top of the COLLAGEN middleware for collaborative interface agents. Jacob Eisenstein, Charles Rich |
IUI | 1 |
| 2002 | A GUI editor that generates tutoring agentsabstractTutoring agents can provide a dynamic and engaging way to help users understand an application. However, integrating tutoring agents into applications is difficult. It requires the expertise to create the tutoring agent, and also an understanding of the inner workings of the application itself. This demo presents a task-based GUI editor that produces a software agent tutor for free. The designer need only create a task model, and then use the editor to produce the GUI. A tutoring agent will automatically be included in the new application. Jacob Eisenstein, Charles Rich |
IUI | 1 |
| 2002 | XIML: a common representation for interaction data abstractWe introduce XIML (eXtensible Interface Markup Language), a proposed common representation for interaction data. We claim that XIML fulfills the requirements that we have found essential for a language of its type: (1) it supports design, operation, organization, and evaluation functions, (2) it is able to relate the abstract and concrete data elements of an interface, and (3) it enables knowledge-based systems to exploit the captured data. Angel R. Puerta, Jacob Eisenstein |
IUI | 2 |
| 2001 | Alternative Representations and Abstractions for Moving Sensors DatabasesabstractMoving sensors refers to an emerging class of data intensive applications that inpacts disciplines such as communication, health-care, scientific applications, etc. These applications consist of a fixed number of sensors that move and produce streams of data as a function of time. They may require the system to match these streams against stored streams to retrieve relevant data (patterns). With communication, for example, a speaking impaired individual might utilize a haptic glove that translates hand signs into written (spoken) words. The glove consists of sensors for different finger joints. These sensors report their location and values as a function of time, producing streams of data. These streams are matched against a repository of spatio-temporal streams to retrieve the corresponding English character or word.The contributions of this study are two fold. First, it introduces a framework to store and retrieve "moving sensors" data. The framework advocates physical data independence and software-reuse. Second, we investigate alternative representations for storage and retrieve of data in support of query processing. We quantify the tradeoff associated with these alternatives using empirical data RoboCup soccer matches. Jacob Eisenstein, Shahram Ghandeharizadeh, Cyrus Shahabi, Gautam Shanbhag, Roger Zimmermann |
CIKM | 1 |
| 2001 | Applying model-based techniques to the development of UIs for mobile computersabstractMobile computing poses a series of unique challenges for user interface design and development: user interfaces must now accommodate the capabilities of various access devices and be suitable for different contexts of use, while preserving consistency and usability. We propose a set of techniques that will aid UI designers who are working in the domain of mobile computing. These techniques will allow designers to build UIs across several platforms, while respecting the unique constraints posed by each platform. In addition, these techniques will help designers to recognize and accommodate the unique contexts in which mobile computing occurs. Central to our approach is the development of a user-interface model that serves to isolate those features that are common to the various contexts of use, and to specify how the user-interface should adjust when the context changes. We claim that without some abstract description of the UI, it is likely that the design and the development of user-interfaces for mobile computing will be very time consuming, error-prone or even doomed to failure. Jacob Eisenstein, Jean Vanderdonckt, Angel R. Puerta |
IUI | 1 |
| 2000 | Adaptation in automated user-interface designabstractDesign problems involve issues of stylistic preference and flexible standards of success; human designers often proceed by intuition and are unaware of following any strict rule-based procedures. These features make design tasks especially difficult to automate. Adaptation is proposed as a means to overcome these challenges. We describe a system that applies an adaptive algorithm to automated user interface design within the framework of the MOBI-D (Model-Based Interface Designer) interface development environment. Preliminary experiments indicate that adaptation improves the performance of the automated user interface design system. Jacob Eisenstein, Angel R. Puerta |
IUI | 1 |
| 1999 | Individual and/versus social creativity (panel session)abstractArticle Free Access Share on Individual and/versus social creativity (panel session) Authors: Ernest Edmonds LUTCHI Research Centre, Loughborough University, Leicestershire, LE11 3TU, UK LUTCHI Research Centre, Loughborough University, Leicestershire, LE11 3TU, UKView Profile , Linda Candy Loughborough University, UK Loughborough University, UKView Profile , Geoff Cox University of Plymouth, UK University of Plymouth, UKView Profile , Jacob Eisenstein Stanford University Stanford UniversityView Profile , Gerhard Fischer University of Colorado University of ColoradoView Profile , Bob Hughes Bristol, UK Bristol, UKView Profile , Tom Hewett Drexel University Drexel UniversityView Profile Authors Info & Claims C&C '99: Proceedings of the 3rd conference on Creativity & cognitionOctober 1999Pages 36–39https://doi.org/10.1145/317561.317570Published:01 October 1999Publication History 8citation646DownloadsMetricsTotal Citations8Total Downloads646Last 12 Months24Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Ernest A. Edmonds, Linda Candy, Geoff Cox, Jacob Eisenstein, Gerhard Fischer, Bob Hughes, Thomas T. Hewett |
Creativity & Cognition | 4 |
| 1999 | Towards a General Computational Framework for Model-Based Interface Development SystemsabstractModel-based interface development systems have not been able to progress beyond producing narrowly focused interface designs of restricted applicability. We identify a level-of-abstraction mismatch in interface models, which we call the mapping problem, as the cause of the limitations in the usefulness of model-based systems. We propose a general computational framework for solving the mapping problem in model-based systems. We show an implementation of the framework within the MOBI-D (Model-Based Interface Designer) interface development environment. The MOBI-D approach to solving the mapping problem enables for the first time with modelbased technology the design of a wide variety of types of user interfaces. Keywords Model-based interface development, interface models, knowledge-based user interface design, user interface development tools Angel R. Puerta, Jacob Eisenstein |
IUI | 2 |
| 1999 | Towards a general computational framework for model-based interface development systems
Angel R. Puerta, Jacob Eisenstein |
Knowl. Based Syst. | 2 |