Jason Baldridge

dblp:90/6617 · also Jason M. Baldridge · DBLP profile ↗
← Back
69ranked-venue papers
9as first author
18since 2021 · last 2024
0000-0002-4712-1841ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 67 · 9 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Can Generative Multimodal Models Count to Ten?
Sunayana Rane, Alexander Ku, Jason Baldridge, Ian Tenney, Thomas L. Griffiths 0001, Been Kim
CogSci3
2024 Where Do We Go From Here? Multi-scale Allocentric Relational Inferencefrom Natural Spatial Descriptions
abstract
Tzuf Paz-Argaman, John Palowitch, Sayali Kulkarni, Jason Baldridge, Reut Tsarfaty. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Tzuf Paz-Argaman, John Palowitch, Sayali Kulkarni, Jason Baldridge, Reut Tsarfaty
EACL (1)4
2024 DOCCI: Descriptions of Connected and Contrasting Images
Yasumasa Onoe, Sunayana Rane, Zachary Berger, Yonatan Bitton, Jaemin Cho 0001, Roopal Garg, Alexander Ku, Zarana Parekh, Jordi Pont-Tuset, Garrett Tanzer, Su Wang 0001, Jason Baldridge
ECCV (60)12
2024 ImageInWords: Unlocking Hyper-Detailed Image Descriptions
abstract
Roopal Garg, Andrea Burns, Burcu Karagol Ayan, Yonatan Bitton, Ceslee Montgomery, Yasumasa Onoe, Andrew Bunner, Ranjay Krishna, Jason Michael Baldridge, Radu Soricut. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Roopal Garg, Andrea Burns, Burcu Karagol Ayan, Yonatan Bitton, Ceslee Montgomery, Yasumasa Onoe, Andrew Bunner, Ranjay Krishna, Jason Baldridge, Radu Soricut
EMNLP9
2024 Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
abstract
Evaluating text-to-image models is notoriously difficult. A strong recent approach for assessing text-image faithfulness is based on QG/A (question generation and answering), which uses pre-trained foundational models to automatically generate a set of questions and answers from the prompt, and output images are scored based on whether these answers extracted with a visual question answering model are consistent with the prompt-based answers. This kind of evaluation is naturally dependent on the quality of the underlying QG and VQA models. We identify and address several reliability challenges in existing QG/A work: (a) QG questions should respect the prompt (avoiding hallucinations, duplications, and omissions) and (b) VQA answers should be consistent (not asserting that there is no motorcycle in an image while also claiming the motorcycle is blue). We address these issues with Davidsonian Scene Graph (DSG), an empirically grounded evaluation framework inspired by formal semantics, which is adaptable to any QG/A frameworks. DSG produces atomic and unique questions organized in dependency graphs, which (i) ensure appropriate semantic coverage and (ii) sidestep inconsistent answers. With extensive experimentation and human evaluation on a range of model configurations (LLM, VQA, and T2I), we empirically demonstrate that DSG addresses the challenges noted above. Finally, we present DSG-1k, an open-sourced evaluation benchmark that includes 1,060 prompts, covering a wide range of fine-grained semantic categories with a balanced distribution. We release the DSG-1k prompts and the corresponding DSG questions.
Jaemin Cho 0001, Yushi Hu, Jason Baldridge, Roopal Garg, Ranjay Krishna, Mohit Bansal, Jordi Pont-Tuset, Su Wang 0001
ICLR3
2024 CoBIT: A Contrastive Bi-directional Image-Text Generation Model
abstract
The field of Vision-and-Language (VL) has witnessed a proliferation of pretrained foundation models. Current techniques typically employ only one type of training objective, whether it's (1) contrastive objectives (like CLIP), (2) image-to-text generative objectives (like PaLI), or (3) text-to-image generative objectives (like Parti). However, all these three objectives are mutually relevant and are all based on image-text pairs. Intuitively, the first two objectives can be considered as complementary projections between two modalities, and contrastive learning can preserve global alignment and generations facilitate fine-grained understanding. Inspired by this, we present a Contrastive Bi-directional Image-Text generation model (CoBIT) to first time unify the three pre-training objectives in one framework. Specifically, CoBIT employs a novel unicoder-decoder structure consisting of an image unicoder, a text unicoder, and a cross-modal decoder. The image/text unicoders can switch between encoding and decoding in different tasks, enabling flexibility and shared knowledge that benefits both image-to-text and text-to-image generations. CoBIT achieves superior performance in image understanding, image-text understanding (Retrieval, Captioning, VQA, SNLI-VE), and text-based content creation, particularly in zero-shot scenarios.
Haoxuan You, Mandy Guo, Zhecan Wang, Kai-Wei Chang 0001, Jason Baldridge
ICLR5
2023 Simple and Effective Synthesis of Indoor 3D Scenes
abstract
We study the problem of synthesizing immersive 3D indoor scenes from one or a few images. Our aim is to generate high-resolution images and videos from novel viewpoints, including viewpoints that extrapolate far beyond the input images while maintaining 3D consistency. Existing approaches are highly complex, with many separately trained stages and components. We propose a simple alternative: an image-to-image GAN that maps directly from reprojections of incomplete point clouds to full high-resolution RGB-D images. On the Matterport3D and RealEstate10K datasets, our approach significantly outperforms prior work when evaluated by humans, as well as on FID scores. Further, we show that our model is useful for generative data augmentation. A vision-and-language navigation (VLN) agent trained with trajectories spatially-perturbed by our model improves success rate by up to 1.5% over a state of the art baseline on the mature R2R benchmark. Our code will be made available to facilitate generative data augmentation and applications to downstream robotics and embodied AI tasks.
Jing Yu Koh, Harsh Agrawal, Dhruv Batra, Austin Waters, Honglak Lee, Yinfei Yang, Jason Baldridge
AAAI8
2023 Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting
abstract
Text-guided image editing can have a transformative impact in supporting creative applications. A key challenge is to generate edits that are faithful to input text prompts, while consistent with input images. We present Imagen Editor, a cascaded diffusion model built, by fine-tuning Imagen [36] on text-guided image inpainting. Imagen Editor's edits are faithful to the text prompts, which is accomplished by using object detectors to propose inpainting masks during training. In addition, Imagen Editor captures fine details in the input image by conditioning the cascaded pipeline on the original high resolution image. To improve qualitative and quantitative evaluation, we introduce EditBench, a systematic benchmark for text-guided image inpainting. EditBench evaluates inpainting edits on natural and generated images exploring objects, attributes, and scenes. Through extensive human evaluation on EditBench, we find that object-masking during training leads to across-the-board improvements in text-image alignment – such that Imagen Editor is preferred over DALL-E 2 [31] and Stable Diffusion [33] – and, as a cohort, these models are better at object-rendering than text-rendering, and handle material/color/size attributes better than count/shape attributes.
Su Wang 0001, Chitwan Saharia, Ceslee Montgomery, Jordi Pont-Tuset, Shai Noy, Stefano Pellegrini, Yasumasa Onoe, Sarah Laszlo, David J. Fleet, Radu Soricut, Jason Baldridge, Mohammad Norouzi 0002
CVPR11
2023 A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation Learning
abstract
Recent studies in Vision-and-Language Navigation (VLN) train RL agents to execute natural-language navigation instructions in photorealistic environments, as a step towards robots that can follow human instructions. However, given the scarcity of human instruction data and limited diversity in the training environments, these agents still struggle with complex language grounding and spatial language understanding. Pretraining on large text and image-text datasets from the web has been extensively explored but the improvements are limited. We investigate large-scale augmentation with synthetic instructions. We take 500+ indoor environments captured in densely-sampled 360 ° panoramas, construct navigation trajectories through these panoramas, and generate a visually-grounded instruction for each trajectory using Marky [63], a high-quality multilingual navigation instruction generator. We also synthesize image observations from novel viewpoints using an image-to-image GAN [27]. The resulting dataset of 4.2M instruction-trajectory pairs is two orders of magnitude larger than existing human-annotated datasets, and contains a wider variety of environments and viewpoints. To efficiently leverage data at this scale, we train a simple transformer agent with imitation learning. On the challenging RxR dataset, our approach outperforms all existing RL agents, improving the state-of-the-art NDTW from 71.1 to 79.1 in seen environments, and from 64.6 to 66.8 in unseen test environments. Our work points to a new path to improving instruction-following agents, emphasizing large-scale training on near-human quality synthetic instructions.
Aishwarya Kamath, Su Wang 0001, Jing Yu Koh, Alexander Ku, Austin Waters, Yinfei Yang, Jason Baldridge, Zarana Parekh
CVPR8
2023 Gaussian Process Probes (GPP) for Uncertainty-Aware Probing
abstract
Understanding which concepts models can and cannot represent has been fundamental to many tasks: from effective and responsible use of models to detecting out of distribution data. We introduce Gaussian process probes (GPP), a unified and simple framework for probing and measuring uncertainty about concepts represented by models. As a Bayesian extension of linear probing methods, GPP asks what kind of distribution over classifiers (of concepts) is induced by the model. This distribution can be used to measure both what the model represents and how confident the probe is about what the model represents. GPP can be applied to any pre-trained model with vector representations of inputs (e.g., activations). It does not require access to training data, gradients, or the architecture. We validate GPP on datasets containing both synthetic and real images. Our experiments show it can (1) probe a model's representations of concepts even with a very small number of examples, (2) accurately measure both epistemic uncertainty (how confident the probe is) and aleatory uncertainty (how fuzzy the concepts are to the model), and (3) detect out of distribution data using those uncertainty measures as well as classic methods do. By using Gaussian processes to expand what probing can offer, GPP provides a data-efficient, versatile and uncertainty-aware tool for understanding and evaluating the capabilities of machine learning models.
Alexander Ku, Jason Baldridge, Thomas L. Griffiths 0001, Been Kim
NeurIPS3
2022 Less is More: Generating Grounded Navigation Instructions from Landmarks
abstract
We study the automatic generation of navigation instructions from 360° images captured on indoor routes. Existing generators suffer from poor visual grounding, causing them to rely on language priors and hallucinate objects. Our Marky-mt5 system addresses this by focusing on visual landmarks; it comprises a first stage landmark detector and a second stage generator-a multimodal, multilingual, multi-task encoder-decoder. To train it, we bootstrap grounded landmark annotations on top of the Room-across-Room (RxR) dataset. Using text parsers, weak supervision from RxR's pose traces, and a multilingual image-text encoder trained on 1.8b images, we identify 971k English, Hindi and Telugu landmark descriptions and ground them to specific regions in panoramas. On Room-to-Room, human wayfind-ers obtain success rates (SR) of 71% following Marky-mt5's instructions, just shy of their 75% SR following human instructions-and well above SRs with other genera-tors. Evaluations on RxR's longer, diverse paths obtain 61-64% SRs on three languages. Generating such high-quality navigation instructions in novel environments is a step to-wards conversational navigation tools and could facilitate larger-scale training of instruction-following agents.
Su Wang 0001, Ceslee Montgomery, Jordi Orbay, Vighnesh Birodkar, Aleksandra Faust, Izzeddin Gur, Natasha Jaques, Austin Waters, Jason Baldridge
CVPR9
2022 Vector-quantized Image Modeling with Improved VQGAN
Jing Yu Koh, Han Zhang 0010, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge
ICLR9
2021 Cross-Modal Contrastive Learning for Text-to-Image Generation
abstract
The output of text-to-image synthesis systems should be coherent, clear, photo-realistic scenes with high semantic fidelity to their conditioned text descriptions. Our Cross-Modal Contrastive Generative Adversarial Network (XMC-GAN) addresses this challenge by maximizing the mutual information between image and text. It does this via multiple contrastive losses which capture inter-modality and intra-modality correspondences. XMC-GAN uses an attentional self-modulation generator, which enforces strong text-image correspondence, and a contrastive discriminator, which acts as a critic as well as a feature encoder for contrastive learning. The quality of XMC-GAN’s output is a major step up from previous models, as we show on three challenging datasets. On MS-COCO, not only does XMC-GAN improve state-of-the-art FID from 24.70 to 9.33, but– more importantly–people prefer XMC-GAN by 77.3% for image quality and 74.1% for image-text alignment, compared to three other recent models. XMC-GAN also generalizes to the challenging Localized Narratives dataset (which has longer, more detailed descriptions), improving state-of-the-art FID from 48.70 to 14.12. Lastly, we train and evaluate XMC-GAN on the challenging Open Images data, establishing a strong benchmark FID score of 26.91.
Han Zhang 0010, Jing Yu Koh, Jason Baldridge, Honglak Lee, Yinfei Yang
CVPR3
2021 Crisscrossed Captions: Extended Intramodal and Intermodal Semantic Similarity Judgments for MS-COCO
abstract
By supporting multi-modal retrieval training and evaluation, image captioning datasets have spurred remarkable progress on representation learning.Unfortunately, datasets have limited cross-modal associations: images are not paired with other images, captions are only paired with other captions of the same image, there are no negative associations and there are missing positive cross-modal associations.This undermines research into how inter-modality learning impacts intra-modality tasks.We address this gap with Crisscrossed Captions (CxC), an extension of the MS-COCO dataset with human semantic similarity judgments for 267,095 intra-and intermodality pairs.We report baseline results on CxC for strong existing unimodal and multimodal models.We also evaluate a multitask dual encoder trained on both image-caption and caption-caption pairs that crucially demonstrates CxC's value for measuring the influence of intra-and inter-modality learning.
Zarana Parekh, Jason Baldridge, Daniel M. Cer, Austin Waters, Yinfei Yang
EACL2
2021 On the Evaluation of Vision-and-Language Navigation Instructions
abstract
Ming Zhao, Peter Anderson, Vihan Jain, Su Wang, Alexander Ku, Jason Baldridge, Eugene Ie. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Vihan Jain, Su Wang 0001, Alexander Ku, Jason Baldridge, Eugene Ie
EACL6
2021 Pathdreamer: A World Model for Indoor Navigation
abstract
People navigating in unfamiliar buildings take advantage of myriad visual, spatial and semantic cues to efficiently achieve their navigation goals. Towards equipping computational agents with similar capabilities, we introduce Pathdreamer, a visual world model for agents navigating in novel indoor environments. Given one or more previous visual observations, Pathdreamer generates plausible high-resolution 360° visual observations (RGB, semantic segmentation and depth) for viewpoints that have not been visited, in buildings not seen during training. In regions of high uncertainty (e.g. predicting around corners, imagining the contents of an unseen room), Pathdreamer can predict diverse scenes, allowing an agent to sample multiple realistic outcomes for a given trajectory. We demonstrate that Pathdreamer encodes useful and accessible visual, spatial and semantic knowledge about human environments by using it in the downstream task of Vision-and-Language Navigation (VLN). Specifically, we show that planning ahead with Pathdreamer brings about half the benefit of looking ahead at actual observations from unobserved parts of the environment. We hope that Pathdreamer will help unlock model-based approaches to challenging embodied navigation tasks such as navigating to specified objects and VLN.
Jing Yu Koh, Honglak Lee, Yinfei Yang, Jason Baldridge
ICCV4
2021 Talk, Don't Write: A Study of Direct Speech-Based Image Retrieval
abstract
Speech-based image retrieval has been studied as a proxy for joint representation learning, usually without emphasis on retrieval itself.As such, it is unclear how well speech-based retrieval can work in practice -both in an absolute sense and versus alternative strategies that combine automatic speech recognition (ASR) with strong text encoders.In this work, we extensively study and expand choices of encoder architectures, training methodology (including unimodal and multimodal pretraining), and other factors.Our experiments cover different types of speech in three datasets: Flickr Audio, Places Audio, and Localized Narratives.Our best model configuration achieves large gains over state of the art, e.g., pushing recall-atone from 21.8% to 33.2% for Flickr Audio and 27.6% to 53.4% for Places Audio.We also show our best speech-based models can match or exceed cascaded ASR-to-text encoding when speech is spontaneous, accented, or otherwise hard to automatically transcribe.
Ramon Sanabria, Austin Waters, Jason Baldridge
Interspeech3
2021 Text-to-Image Generation Grounded by Fine-Grained User Attention
abstract
Localized Narratives [28] is a dataset with detailed natural language descriptions of images paired with mouse traces that provide a sparse, fine-grained visual grounding for phrases. We propose TRECS, a sequential model that exploits this grounding to generate images. TRECS uses descriptions to retrieve segmentation masks and predict object labels aligned with mouse traces. These alignments are used to select and position masks to generate a fully covered segmentation canvas; the final image is produced by a segmentation-to-image generator using this canvas. This multi-step, retrieval-based approach outperforms existing direct text-to-image generation models on both automatic metrics and human evaluations: overall, its generated images are more photo-realistic and better match descriptions.
Jing Yu Koh, Jason Baldridge, Honglak Lee, Yinfei Yang
WACV2
2020 Mapping Natural Language Instructions to Mobile UI Action Sequences
abstract
We present a new problem: grounding natural language instructions to mobile user interface actions, and create three new datasets for it.For full task evaluation, we create PIX-ELHELP, a corpus that pairs English instructions with actions performed by people on a mobile UI emulator.To scale training, we decouple the language and action data by (a) annotating action phrase spans in HowTo instructions and (b) synthesizing grounded descriptions of actions for mobile user interfaces.We use a Transformer to extract action phrase tuples from long-range natural language instructions.A grounding Transformer then contextually represents UI objects using both their content and screen position and connects them to object descriptions.Given a starting screen and instruction, our model achieves 70.59% accuracy on predicting complete ground-truth action sequences in PIXELHELP.
Yang Li 0058, Jiacong He, Xin Zhou 0018, Yuan Zhang 0001, Jason Baldridge
ACL5
2020 Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding
abstract
We introduce Room-Across-Room (RxR), a new Vision-and-Language Navigation (VLN) dataset.RxR is multilingual (English, Hindi, and Telugu) and larger (more paths and instructions) than other VLN datasets.It emphasizes the role of language in VLN by addressing known biases in paths and eliciting more references to visible entities.Furthermore, each word in an instruction is time-aligned to the virtual poses of instruction creators and validators.We establish baseline scores for monolingual and multilingual settings and multitask learning when including Room-to-Room annotations (Anderson et al., 2018b).We also provide results for a model that learns from synchronized pose traces by focusing only on portions of the panorama attended to in human demonstrations.The size, scope and detail of RxR dramatically expands the frontier for research on embodied language agents in simulated, photo-realistic environments.
Alexander Ku, Roma Patel, Eugene Ie, Jason Baldridge
EMNLP (1)5
2019 A Task in a Suit and a Tie: Paraphrase Generation with Semantic Augmentation
abstract
Paraphrasing is rooted in semantics. We show the effectiveness of transformers (Vaswani et al. 2017) for paraphrase generation and further improvements by incorporating PropBank labels via a multi-encoder. Evaluating on MSCOCO and WikiAnswers, we find that transformers are fast and effective, and that semantic augmentation for both transformers and LSTMs leads to sizable 2-3 point gains in BLEU, METEOR and TER. More importantly, we find surprisingly large gains on human evaluations compared to previous models. Nevertheless, manual inspection of generated paraphrases reveals ample room for improvement: even our best model produces human-acceptable paraphrases for only 28% of captions from the CHIA dataset (Sharma et al. 2018), and it fails spectacularly on sentences from Wikipedia. Overall, these results point to the potential for incorporating semantics in the task while highlighting the need for stronger evaluation.
Su Wang 0001, Nancy Chang, Jason Baldridge
AAAI4
2019 Stay on the Path: Instruction Fidelity in Vision-and-Language Navigation
abstract
Advances in learning and representations have reinvigorated work that connects language to other modalities.A particularly exciting direction is Vision-and-Language Navigation (VLN), in which agents interpret natural language instructions and visual scenes to move through environments and reach goals.Despite recent progress, current research leaves unclear how much of a role language understanding plays in this task, especially because dominant evaluation metrics have focused on goal completion rather than the sequence of actions corresponding to the instructions.Here, we highlight shortcomings of current metrics for the Room-to-Room dataset (Anderson et al., 2018b) and propose a new metric, Coverage weighted by Length Score (CLS).We also show that the existing paths in the dataset are not ideal for evaluating instruction following because they are direct-to-goal shortest paths.We join existing short paths to form more challenging extended paths to create a new data set, Room-for-Room (R4R).Using R4R and CLS, we show that agents that receive rewards for instruction fidelity outperform agents that focus on goal completion.
Vihan Jain, Gabriel Ilharco, Alexander Ku, Ashish Vaswani, Eugene Ie, Jason Baldridge
ACL (1)6
2019 Learning Dense Representations for Entity Retrieval
abstract
Daniel Gillick, Sayali Kulkarni, Larry Lansing, Alessandro Presta, Jason Baldridge, Eugene Ie, Diego Garcia-Olano. Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL). 2019.
Daniel Gillick, Sayali Kulkarni, Larry Lansing, Alessandro Presta, Jason Baldridge, Eugene Ie, Diego Garcia-Olano
CoNLL5
2019 Large-Scale Representation Learning from Visually Grounded Untranscribed Speech
abstract
Systems that can associate images with their spoken audio captions are an important step towards visually grounded language learning.We describe a scalable method to automatically generate diverse audio for image captioning datasets.This supports pretraining deep networks for encoding both audio and images, which we do via a dual encoder that learns to align latent representations from both modalities.We show that a masked margin softmax loss for such models is superior to the standard triplet loss.We fine-tune these models on the Flickr8k Audio Captions Corpus and obtain state-of-the-art results-improving recall in the top 10 from 29.6% to 49.5%.We also obtain human ratings on retrieval outputs to better assess the impact of incidentally matching image-caption pairs that were not associated in the data, finding that automatic evaluation substantially underestimates the quality of the retrieved results. * Work done as a member of the Google AI Residency Program.We address the problem of relating images to audio captions that describe them (Figure 1), building on previous research into learning from visually grounded, untranscribed speech (Harwath and Glass, 2015;Sun et al., 2016;Harwath et al., 2016;Chrupała et al., 2017; Kamper et al., 2017b;Chrupała, 2019;Harwath and Glass, 2019).Such problem settings provide opportunities both to improve our theoretical understanding of language
Gabriel Ilharco, Yuan Zhang 0001, Jason Baldridge
CoNLL3
2019 PAWS-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification
abstract
Yinfei Yang, Yuan Zhang, Chris Tar, Jason Baldridge. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yinfei Yang, Yuan Zhang 0001, Chris Tar, Jason Baldridge
EMNLP/IJCNLP (1)4
2019 Transferable Representation Learning in Vision-and-Language Navigation
abstract
Vision-and-Language Navigation (VLN) tasks such as Room-to-Room (R2R) require machine agents to interpret natural language instructions and learn to act in visually realistic environments to achieve navigation goals. The overall task requires competence in several perception problems: successful agents combine spatio-temporal, vision and language understanding to produce appropriate action sequences. Our approach adapts pre-trained vision and language representations to relevant in-domain tasks making them more effective for VLN. Specifically, the representations are adapted to solve both a cross-modal sequence alignment and sequence coherence task. In the sequence alignment task, the model determines whether an instruction corresponds to a sequence of visual frames. In the sequence coherence task, the model determines whether the perceptual sequences are predictive sequentially in the instruction-conditioned latent space. By transferring the domain-adapted representations, we improve competitive agents in R2R as measured by the success rate weighted by path length (SPL) metric.
Haoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku, Gabriel Ilharco, Jason Baldridge, Eugene Ie
ICCV6
2018 Learning To Split and Rephrase From Wikipedia Edit History
abstract
Split and rephrase is the task of breaking down a sentence into shorter ones that together convey the same meaning.We extract a rich new dataset for this task by mining Wikipedia's edit history: WikiSplit contains one million naturally occurring sentence rewrites, providing sixty times more distinct split examples and a ninety times larger vocabulary than the WebSplit corpus introduced by Narayan et al. (2017) as a benchmark for this task.Incorporating WikiSplit as training data produces a model with qualitatively better predictions that score 32 BLEU points above the prior best result on the WebSplit benchmark.
Jan A. Botha, Manaal Faruqui, John Alex, Jason Baldridge, Dipanjan Das 0001
EMNLP4
2018 A Fast, Compact, Accurate Model for Language Identification of Codemixed Text
abstract
We address fine-grained multilingual language identification: providing a language code for every token in a sentence, including codemixed text containing multiple languages.Such text is prevalent online, in documents, social media, and message boards.We show that a feed-forward network with a simple globally constrained decoder can accurately and rapidly label both codemixed and monolingual text in 100 languages and 100 language pairs.This model outperforms previously published multilingual approaches in terms of both accuracy and speed, yielding an 800x speed-up and a 19.5% averaged absolute gain on three codemixed datasets.It furthermore outperforms several benchmark systems on monolingual language identification.
Yuan Zhang 0001, Jason Riesa, Daniel Gillick, Anton Bakalov, Jason Baldridge, David Weiss 0001
EMNLP5
2018 Mind the GAP: A Balanced Corpus of Gendered Ambiguous Pronouns
abstract
Coreference resolution is an important task for natural language understanding, and the resolution of ambiguous pronouns a longstanding challenge. Nonetheless, existing corpora do not capture ambiguous pronouns in sufficient volume or diversity to accurately indicate the practical utility of models. Furthermore, we find gender bias in existing corpora and systems favoring masculine entities. To address this, we present and release GAP, a gender-balanced labeled corpus of 8,908 ambiguous pronoun–name pairs sampled to provide diverse coverage of challenges posed by real-world text. We explore a range of baselines that demonstrate the complexity of the challenge, the best achieving just 66.9% F1. We show that syntactic structure and continuous neural models provide promising, complementary cues for approaching the challenge.
Kellie Webster, Marta Recasens, Vera Axelrod, Jason Baldridge
Trans. Assoc. Comput. Linguistics4
2015 Gazetteer-Independent Toponym Resolution Using Geographic Word Profiles
abstract
Toponym resolution, or grounding names of places to their actual locations, is an important problem in analysis of both historical corpora and present-day news and web content. Recent approaches have shifted from rule-based spatial minimization methods to machine learned classifiers that use features of the text surrounding a toponym. Such methods have been shown to be highly effective, but they crucially rely on gazetteers and are unable to handle unknown place names or locations. We address this limitation by modeling the geographic distributions of words over the earth's surface: we calculate the geographic profile of each word based on local spatial statistics over a set of geo-referenced language models. These geo-profiles can be further refined by combining in-domain data with background statistics from Wikipedia. Our resolver computes the overlap of all geo-profiles in a given text span; without using a gazetteer, it performs on par with existing classifiers. When combined with a gazetteer, it achieves state-of-the-art performance for two standard toponym resolution corpora (TR-CoNLL and Civil War). Furthermore, it dramatically improves recall when toponyms are identified by named entity recognizers, which often (correctly) find non-standard variants of toponyms.
Grant DeLozier, Jason Baldridge, Loretta London
AAAI2
2015 Weakly-Supervised Grammar-Informed Bayesian CCG Parser Learning
abstract
Combinatory Categorial Grammar (CCG) is a lexicalized grammar formalism in which words are associated with categories that, in combination with a small universal set of rules, specify the syntactic configurations in which they may occur. Categories are selected from a large, recursively-defined set; this leads to high word-to-category ambiguity, which is one of the primary factors that make learning CCG parsers difficult, especially in the face of little data. Previous work has shown that learning sequence models for CCG tagging can be improved by using linguistically-motivated prior probability distributions over potential categories. We extend this approach to the task of learning a CCG parser from weak supervision. We present a Bayesian formulation for CCG parser induction that assumes only supervision in the form of an incomplete tag dictionary mapping some word types to sets of potential categories. Our approach outperforms a baseline model trained with uniform priors by exploiting universal, intrinsic properties of the CCG formalism to bias the model toward simpler, more cross-linguistically common categories.
Dan Garrette, Chris Dyer, Jason Baldridge, Noah A. Smith
AAAI3
2015 Parse Imputation for Dependency Annotations
abstract
Jason Mielens, Liang Sun, Jason Baldridge. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Jason Mielens, Jason Baldridge
ACL (1)3
2015 A Supertag-Context Model for Weakly-Supervised CCG Parser Learning
abstract
Combinatory Categorial Grammar (CCG) is a lexicalized grammar formalism in which words are associated with categories that specify the syntactic configurations in which they may occur.We present a novel parsing model with the capacity to capture the associative adjacent-category relationships intrinsic to CCG by parameterizing the relationships between each constituent label and the preterminal categories directly to its left and right, biasing the model toward constituent categories that can combine with their contexts.This builds on the intuitions of Klein and Manning's (2002) "constituentcontext" model, which demonstrated the value of modeling context, but has the advantage of being able to exploit the properties of CCG.Our experiments show that our model outperforms a baseline in which this context information is not captured.
Dan Garrette, Chris Dyer, Jason Baldridge, Noah A. Smith
CoNLL3
2014 Weakly-Supervised Bayesian Learning of a CCG Supertagger
abstract
We present a Bayesian formulation for weakly-supervised learning of a Combinatory Categorial Grammar (CCG) supertagger with an HMM.We assume supervision in the form of a tag dictionary, and our prior encourages the use of crosslinguistically common category structures as well as transitions between tags that can combine locally according to CCG's combinators.Our prior is theoretically appealing since it is motivated by languageindependent, universal properties of the CCG formalism.Empirically, we show that it yields substantial improvements over previous work that used similar biases to initialize an EM-based learner.Additional gains are obtained by further shaping the prior with corpus-specific information that is extracted automatically from raw text and a tag dictionary.
Dan Garrette, Chris Dyer, Jason Baldridge, Noah A. Smith
CoNLL3
2014 Parsing low-resource languages using Gibbs sampling for PCFGs with latent annotations
abstract
PCFGs with latent annotations have been shown to be a very effective model for phrase structure parsing. We present a Bayesian model and algorithms based on a Gibbs sam-pler for parsing with a grammar with latent an-notations. For PCFG-LA, we present an ad-ditional Gibbs sampler algorithm to learn an-notations from training data, which are parse trees with coarse (unannotated) symbols. We show that a Gibbs sampling technique is ca-pable of parsing sentences in a wide variety of languages and producing results that are on-par with or surpass previous approaches. Our results for Kinyarwanda and Malagasy in particular demonstrate that low-resource lan-guage parsing can benefit substantially from a Bayesian approach. 1
Jason Mielens, Jason Baldridge
EMNLP3
2014 Hierarchical Discriminative Classification for Text-Based Geolocation
abstract
Text-based document geolocation is com-monly rooted in language-based infor-mation retrieval techniques over geodesic grids. These methods ignore the natural hierarchy of cells in such grids and fall afoul of independence assumptions. We demonstrate the effectiveness of using lo-gistic regression models on a hierarchy of nodes in the grid, which improves upon the state of the art accuracy by several percent and reduces mean error distances by hundreds of kilometers on data from Twitter, Wikipedia, and Flickr. We also show that logistic regression performs fea-ture selection effectively, assigning high weights to geocentric terms. 1
Benjamin Wing, Jason Baldridge
EMNLP2
2014 Extracting topics based on authors, recipients and content in microblogs
abstract
Microblogs such as Twitter are important sources for spreading vital information at high speed. They also reflect the general people's reaction and opinion towards major events or stories. With information traveling so quickly, it is helpful to be able to apply unsupervised learning techniques to discover topics for information extraction and analysis. Although graphical models have been traditionally used for topic discovery in microblogs and text streams, previous work may not be as efficient because of the diverse and noisy nature of microblogs.
Nazneen Fatema Rajani, Kate McArdle, Jason Baldridge
SIGIR3
2013 Real-World Semi-Supervised Learning of POS-Taggers for Low-Resource Languages
Dan Garrette, Jason Mielens, Jason Baldridge
ACL (1)3
2013 Text-Driven Toponym Resolution using Indirect Supervision
Michael Speriosu, Jason Baldridge
ACL (1)2
2013 A recursive estimate for the predictive likelihood in a topic model
abstract
We consider the problem of evaluating the predictive log likelihood of a previously un- seen document under a topic model. This task arises when cross-validating for a model hyperparameter, when testing a model on a hold-out set, and when comparing the performance of different fitting strategies. Yet it is known to be very challenging, as it is equivalent to estimating a marginal likelihood in Bayesian model selection. We propose a fast algorithm for approximating this likelihood, one whose computational cost is linear both in document length and in the number of topics. The method is a first-order approximation to the algorithm of Carvalho et al. (2010), and can also be interpreted as a one-particle, Rao-Blackwellized version of the "left-to-right" method of Wallach et al. (2009). On our test examples, the proposed method gives similar answers to these other methods, but at lower computational cost.
James Scott, Jason Baldridge
AISTATS2
2013 Learning a Part-of-Speech Tagger from Two Hours of Annotation
Dan Garrette, Jason Baldridge
HLT-NAACL2
2012 Type-Supervised Hidden Markov Models for Part-of-Speech Tagging with Incomplete Tag Dictionaries
Dan Garrette, Jason Baldridge
EMNLP-CoNLL2
2012 Supervised Text-based Geolocation Using Language Models on an Adaptive Grid
Stephen Roller, Michael Speriosu, Sarat Rallapalli, Benjamin Wing, Jason Baldridge
EMNLP-CoNLL5
2011 Simple Unsupervised Grammar Induction from Raw Text with Cascaded Finite State Models
Elias Ponvert, Jason Baldridge, Katrin Erk
ACL2
2011 Simple supervised document geolocation with geodesic grids
Benjamin Wing, Jason Baldridge
ACL2
2011 Supervised language modeling for temporal resolution of texts
abstract
We investigate temporal resolution of documents, such as determining the date of publication of a story based on its text. We describe and evaluate a model that build histograms encoding the probability of different temporal periods for a document. We construct histograms based on the Kullback-Leibler Divergence between the language model for a test document and supervised language models for each interval. Initial results indicate this language modeling approach is effective for predicting the dates of publication of short stories, which contain few explicit mentions of years.
Abhimanu Kumar, Matthew Lease, Jason Baldridge
CIKM3
2011 Semantic Role Labeling Without Treebanks?
Stephen A. Boxwell, Chris Brew, Jason Baldridge, Dennis Mehay, Sujith Ravi
IJCNLP3
2010 Minimized Models and Grammar-Informed Initialization for Supertagging with Highly Ambiguous Lexicons
Sujith Ravi, Jason Baldridge, Kevin Knight
ACL2
2010 Crouching Dirichlet, Hidden Markov Model: Unsupervised POS Tagging with Context Local Tag Generation
Taesun Moon, Katrin Erk, Jason Baldridge
EMNLP3
2009 How well does active learning
Jason Baldridge, Alexis Palmer
EMNLP1
2009 Unsupervised morphological segmentation and clustering with document boundaries
Taesun Moon, Katrin Erk, Jason Baldridge
EMNLP3
2009 Supertagging with Factorial Hidden Markov Models
Srivatsan Ramanujam, Jason Baldridge
PACLIC2
2008 A Logical Basis for the D Combinator and Normal Form in CCG
Frederick Hoyt, Jason Baldridge
ACL2
2008 Weakly Supervised Supertagging with Grammar-Informed Initialization
Jason Baldridge
COLING1
2008 Specialized Models and Ranking for Coreference Resolution
Pascal Denis, Jason Baldridge
EMNLP2
2008 Active learning and logarithmic opinion pools for HPSG parse selection
abstract
Abstract For complex tasks such as parse selection, the creation of labelled training sets can be extremely costly. Resource-efficient schemes for creating informative labelled material must therefore be considered. We investigate the relationship between two broad strategies for reducing the amount of manual labelling necessary to train accurate parse selection models: ensemble models and active learning. We show that popular active learning methods for reducing annotation costs can be outperformed by instead using a model class which uses the available labelled data more efficiently. For this, we use a simple type of ensemble model called theLogarithmic Opinion Pool(LOP). We furthermore show that LOPs themselves can benefit from active learning. As predicted by a theoretical explanation of the predictive power of LOPs, a detailed analysis of active learning using LOPs shows that component model diversity is a strong predictor of successful LOP performance. Other contributions include a novel active learning method, a justification of our simulation studies using timing information, and cross-domain verification of our main ideas using text classification.
Jason Baldridge, Miles Osborne
Nat. Lang. Eng.1
2007 A Sequencing Model for Situation Entity Classification
Alexis Palmer, Elias Ponvert, Jason Baldridge, Carlota Smith
ACL3
2007 Part-of-Speech Tagging for Middle English through Alignment and Projection of Parallel Diachronic Texts
Taesun Moon, Jason Baldridge
EMNLP-CoNLL2
2007 A Ranking Approach to Pronoun Resolution
Pascal Denis, Jason Baldridge
IJCAI2
2007 Joint Determination of Anaphoricity and Coreference Resolution using Integer Programming
Pascal Denis, Jason Baldridge
HLT-NAACL2
2005 Probabilistic Head-Driven Parsing for Discourse Structure
Jason Baldridge, Alex Lascarides
CoNLL1
2004 Generalizing Dimensionality in Combinatory Categorial Grammar
Geert-Jan M. Kruijff, Jason Baldridge
COLING2
2004 Active Learning and the Total Cost of Annotation
Jason Baldridge, Miles Osborne
EMNLP1
2004 Ensemble-based Active Learning for Parse Selection
Miles Osborne, Jason Baldridge
HLT-NAACL2
2004 Verbmobil: Foundations of Speech-to-Speech Translation, by Wolfgang Wahlster (editor). Springer, 2000. ISBN 3-540-67783-6. Price £44.50 (hardback). xii+679 pages
Jason Baldridge
Nat. Lang. Eng.1
2003 Active learning for HPSG parse selection
Jason Baldridge, Miles Osborne
CoNLL1
2003 Multi-modal combinatory categorial grammar
Geert-Jan M. Kruijff, Jason Baldridge
EACL2
2002 Coupling CCG and Hybrid Logic Dependency Semantics
abstract
Categorial grammar has traditionally used the λ-calculus to represent meaning. We present an alternative, dependency-based perspective on linguistic meaning and situate it in the computational setting. This perspective is formalized in terms of hybrid logic and has a rich yet perspicuous propositional ontology that enables a wide variety of semantic phenomena to be represented in a single meaning formalism. Finally, we show how we can couple this formalization to Combinatory Categorial Grammar to produce interpretations compositionally.
Jason Baldridge, Geert-Jan M. Kruijff
ACL1
2002 Leo: an Architecture for Sharing Resources for Unification-Based Grammars
Jason Baldridge, John Dowding, Susana Early
LREC1