VLDB 2026 Research / reviewers in the wild / expert
Taylor Berg-Kirkpatrick
dblp:22/8160
· DBLP profile ↗
89ranked-venue papers
10as first author
45since 2021 · last 2025
0000-0002-1283-4075ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 79 · 10 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 12 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music ProcessingabstractThe recent explosion of generative AI-Music systems has raised numerous concerns over data copyright, licensing music from musicians, and the conflict between open-source AI and large prestige companies. Such issues highlight the need for publicly available, copyright-free musical data, in which there is a large shortage, particularly for symbolic music data. To alleviate this issue, we present PDMX: a large-scale open-source dataset of over 250K public domain MusicXML scores collected from the score-sharing forum MuseScore, making it the largest available copyright-free symbolic music dataset to our knowledge. PDMX additionally includes a wealth of both tag and user interaction metadata, allowing us to efficiently analyze the dataset and filter for high quality user-generated scores. Given the additional metadata afforded by our data collection process, we conduct multitrack music generation experiments evaluating how different representative subsets of PDMX lead to different behaviors in downstream models, and how user-rating statistics can be used as an effective measure of data quality. Phillip Long, Zachary Novack, Taylor Berg-Kirkpatrick, Julian J. McAuley |
ICASSP | 3 |
| 2025 | ClimaQA: An Automated Evaluation Framework for Climate Question Answering ModelsabstractThe use of Large Language Models (LLMs) in climate science has recently gained significant attention. However, a critical issue remains: the lack of a comprehensive evaluation framework capable of assessing the quality and scientific validity of model outputs. To address this issue, we develop *ClimaGen* (Climate QA Generator), an adaptive learning framework that generates question-answer pairs from graduate textbooks with climate scientists in the loop. As a result, we present *ClimaQA-Gold*, an expert-annotated benchmark dataset alongside *ClimaQA-Silver*, a large-scale, comprehensive synthetic QA dataset for climate science. Finally, we develop evaluation strategies and compare different LLMs on our benchmarks. Our results offer novel insights into various approaches used to enhance knowledge of climate LLMs. ClimaQA’s source code is publicly available at https://github.com/Rose-STL-Lab/genie-climaqa Veeramakali Vignesh Manivannan, Yasaman Jafari, Srikar Eranky, Spencer Ho, Rose Yu, Duncan Watson-Parris, Yi-An Ma, Leon Bergen, Taylor Berg-Kirkpatrick |
ICLR | 9 |
| 2025 | Presto! Distilling Steps and Layers for Accelerating Music GenerationabstractDespite advances in diffusion-based text-to-music (TTM) methods, efficient, high-quality generation remains a challenge. We introduce Presto!, an approach to inference acceleration for score-based diffusion transformers via reducing both sampling steps and cost per step. To reduce steps, we develop a new score-based distribution matching distillation (DMD) method for the EDM-family of diffusion models, the first GAN-based distillation method for TTM. To reduce the cost per step, we develop a simple, but powerful improvement to a recent layer distillation method that improves learning via better preserving hidden state variance. Finally, we combine our step and layer distillation methods together for a dual-faceted approach. We evaluate our step and layer distillation methods independently and show each yield best-in-class performance. Our combined distillation method can generate high-quality outputs with improved diversity, accelerating our base model by 10-18x (230/435ms latency for 32 second mono/stereo 44.1kHz, 15x faster than the comparable SOTA model) — the fastest TTM to our knowledge. Zachary Novack, Jonah Casebeer, Julian J. McAuley, Taylor Berg-Kirkpatrick, Nicholas J. Bryan |
ICLR | 5 |
| 2025 | TeaserGen: Generating Teasers for Long DocumentariesabstractTeasers are an effective tool for promoting content in entertainment, commercial and educational fields. However, creating an effective teaser for long videos is challenging for it requires long-range multimodal modeling capability for the input videos, while necessitating maintaining audiovisual alignments, managing scene transitions and preserving factual accuracy for the output teasers. Due to the lack of a publicly-available dataset, progress along this research direction has been hindered. In this work, we present DocumentaryNet, a collection of 1,269 documentaries paired with their teasers, featuring multimodal data streams of video, speech, music, sound effects and narrations. With DocumentaryNet, we propose a new two-stage system for generating teasers from long documentaries. The proposed TeaserGen system first generates the teaser narration from the transcribed narration from the documentary using a pretrained large language model, and then selects the most relevant visual content to accompany the generated narration through language-vision models. For narration-video matching, we explore two approaches: a pretraining-based model using pretrained contrastive language-vision models and a deep sequential model that learns the mapping between the narrations and visuals. Our experimental results show that the pretraining-based approach is more effective at identifying relevant visual content than directly trained deep autoregressive models. Weihan Xu, Paul Pu Liang, Haven Kim, Julian J. McAuley, Taylor Berg-Kirkpatrick, Hao-Wen Dong |
ICLR | 5 |
| 2025 | Adapting While Learning: Grounding LLMs for Scientific Problems with Tool Usage AdaptationabstractLarge Language Models (LLMs) demonstrate promising capabilities in solving scientific problems but often suffer from the issue of hallucination. While integrating LLMs with tools can mitigate this issue, models fine-tuned on tool usage become overreliant on them and incur unnecessary costs. Inspired by how human experts assess problem complexity before selecting solutions, we propose a novel two-component fine-tuning method, Adapting while Learning (AWL). In the first component World Knowledge Learning (WKL), LLMs internalize scientific knowledge by learning from tool-generated solutions. In the second component Tool Usage Adaptation (TUA), we categorize problems as easy or hard based on the model’s accuracy, and train it to maintain direct reasoning for easy problems while switching to tools for hard ones. We validate our method on 6 scientific benchmark datasets across climate science, epidemiology, physics, and other domains. Compared to the original instruct model (8B), models post-trained with AWL achieve 29.11% higher answer accuracy and 12.72% better tool usage accuracy, even surpassing state-of-the-art models including GPT-4o and Claude-3.5 on 4 custom-created datasets. Our code is open-source at https://github.com/Rose-STL-Lab/Adapting-While-Learning. Bohan Lyu 0001, Yadi Cao, Duncan Watson-Parris, Leon Bergen, Taylor Berg-Kirkpatrick, Rose Yu |
ICML | 5 |
| 2025 | Synthesizing Composite Hierarchical Structure from Symbolic Music CorporaabstractWestern music is an innately hierarchical system of interacting levels of structure, from fine-grained melody to high-level form. In order to analyze music compositions holistically and at multiple granularities, we propose a unified, hierarchical meta-representation of musical structure called the structural temporal graph (STG). For a single piece, the STG is a data structure that defines a hierarchy of progressively finer structural musical features and the temporal relationships between them. We use the STG to enable a novel approach for deriving a representative structural summary of a music corpus, which we formalize as a dually NP-hard combinatorial optimization problem. Our approach first applies simulated annealing to develop a measure of structural distance between two music pieces rooted in graph isomorphism. Our approach then combines the formal guarantees of SMT solvers with nested simulated annealing over structural distances to produce a structurally sound, representative centroid STG for an entire corpus of STGs from individual pieces. To evaluate our approach, we conduct experiments verifying that structural distance accurately differentiates between music pieces, and that derived centroids accurately structurally characterize their corpora. Ilana Shapiro, Ruanqianqian (Lisa) Huang, Zachary Novack, Cheng-i Wang, Hao-Wen Dong, Taylor Berg-Kirkpatrick, Shlomo Dubnov, Sorin Lerner |
IJCAI | 6 |
| 2025 | Constrained Sampling for Language Models Should Be Easy: An MCMC PerspectiveabstractConstrained decoding enables Language Models (LMs) to produce samples that provably satisfy hard constraints.
However, existing constrained-decoding approaches often distort the underlying model distribution, a limitation that is especially problematic in applications like program fuzzing, where one wants to generate diverse and valid program inputs for testing purposes.
We propose a new constrained sampling framework based on Markov Chain Monte Carlo (MCMC) that simultaneously satisfies three core desiderata: constraint satisfying (every sample satisfies the constraint), monotonically converging (the sampling process converges to the true conditional distribution), and efficient (high-quality samples emerge in few steps). Our method constructs a proposal distribution over valid outputs and applies a Metropolis-Hastings acceptance criterion based on the LM’s likelihood, ensuring principled and efficient exploration of the constrained space. Empirically, our sampler outperforms existing methods on both synthetic benchmarks and real-world program fuzzing tasks. Emmanuel Anaya Gonzalez, Sairam Vaidya, Kanghee Park, Ruyi Ji, Taylor Berg-Kirkpatrick, Loris D'Antoni |
NeurIPS | 5 |
| 2025 | REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video EditingabstractShort videos are an effective tool for promoting contents and improving knowledge accessibility. While existing extractive video summarization methods struggle to produce a coherent narrative, existing abstractive methods cannot `quote' from the input videos, i.e., inserting short video clips in their outputs. In this work, we explore novel video editing models for generating shorts that feature a coherent narrative with embedded video insertions extracted from a long input video. We propose a novel retrieval-embedded generation framework that allows a large language model to quote multimodal resources while maintaining a coherent narrative. Our proposed REGen system first generates the output story script with quote placeholders using a finetuned large language model, and then uses a novel retrieval model to replace the quote placeholders by selecting a video clip that best supports the narrative from a pool of candidate quotable video clips. We examine the proposed method on the task of documentary teaser generation, where short interview insertions are commonly used to support the narrative of a documentary. Our objective evaluations show that the proposed method can effectively insert short video clips while maintaining a coherent narrative. In a subjective survey, we show that our proposed method outperforms existing abstractive and extractive approaches in terms of coherence, alignment, and realism in teaser generation. Weihan Xu, Yimeng Ma, Jingyue Huang, Wenye Ma, Taylor Berg-Kirkpatrick, Julian J. McAuley, Paul Pu Liang, Hao-Wen Dong |
NeurIPS | 6 |
| 2024 | LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLPabstractStandard natural language processing (NLP) pipelines operate on symbolic representations of language, which typically consist of sequences of discrete tokens.However, creating an analogous representation for ancient logographic writing systems is an extremely laborintensive process that requires expert knowledge.At present, a large portion of logographic data persists in a purely visual form due to the absence of transcription-this issue poses a bottleneck for researchers seeking to apply NLP toolkits to study ancient logographic languages: most of the relevant data are images of writing.This paper investigates whether direct processing of visual representations of language offers a potential solution.We introduce LogogramNLP, the first benchmark enabling NLP analysis of ancient logographic languages, featuring both transcribed and visual datasets for four writing systems along with annotations for tasks like classification, translation, and parsing.Our experiments compare systems that employ recent visual and text encoding strategies as backbones.The results demonstrate that visual representations outperform textual representations for some investigated tasks, suggesting that visual processing pipelines may unlock a large amount of cultural heritage data of logographic languages for NLP-based analyses. Danlu Chen, Freda Shi, Aditi Agarwal, Jacobo Myerston, Taylor Berg-Kirkpatrick |
ACL (1) | 5 |
| 2024 | MusicLDM: Enhancing Novelty in text-to-music Generation Using Beat-Synchronous mixup StrategiesabstractDiffusion models have shown promising results in cross-modal generation tasks, including text-to-image and text-to-audio generation. However, generating music, as a special type of audio, presents unique challenges due to limited availability of music data and sensitive issues related to copyright and plagiarism. In this paper, to tackle these challenges, we first construct a state-of-the-art text-to-music model, MusicLDM, that adapts Stable Diffusion and AudioLDM architectures to the music domain. Then, to address the limitations of training data and to avoid plagiarism, we leverage a beat tracking model and propose two different mixup strategies for data augmentation: beat-synchronous audio mixup and beat-synchronous latent mixup, which recombine training audio directly or via a latent embeddings space, respectively. Such mixup strategies encourage the model to interpolate between musical training samples and generate new music within the convex hull of the training data, making the generated music more diverse while still staying faithful to the corresponding style. In addition to popular evaluation metrics, we design several new evaluation metrics based on CLAP score to demonstrate that our proposed MusicLDM and beat-synchronous mixup strategies improve both the quality and novelty of generated music, as well as the correspondence between input text and generated music. Ke Chen 0021, Yusong Wu, Haohe Liu, Marianna Nezhurina, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
ICASSP | 5 |
| 2024 | Clustering Running Titles to Understand the Printing of Early Modern Books
Nikolai Vogler, Kartik Goyal, Samuel V. Lemley, D. J. Schuldt, Christopher N. Warren, Max G'Sell, Taylor Berg-Kirkpatrick |
ICDAR (3) | 7 |
| 2024 | Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning AbilityabstractWhat is the relationship between model architecture and the ability to perform in-context learning? In this empirical study, we take the first steps toward answering this question. We evaluate thirteen model architectures capable of causal language modeling across a suite of synthetic in-context learning tasks. These selected architectures represent a broad range of paradigms, including recurrent and convolution-based neural networks, transformers, state space model inspired, and other emerging attention alternatives. We discover that all the considered architectures can perform in-context learning under a wider range of conditions than previously documented. Additionally, we observe stark differences in statistical efficiency and consistency by varying the number of in-context examples and task difficulty. We also measure each architecture's predisposition towards in-context learning when presented with the option to memorize rather than leverage in-context examples. Finally, and somewhat surprisingly, we find that several attention alternatives are sometimes competitive with or better in-context learners than transformers. However, no single architecture demonstrates consistency across all tasks, with performance either plateauing or declining when confronted with a significantly larger number of in-context examples than those encountered during gradient-based training. Taylor Berg-Kirkpatrick |
ICLR | 3 |
| 2024 | Alt-Text with Context: Improving Accessibility for Images on TwitterabstractIn this work we present an approach for generating alternative text (or alt-text) descriptions for images shared on social media, specifically Twitter. More than just a special case of image captioning, alt-text is both more literally descriptive and context-specific. Also critically, images posted to Twitter are often accompanied by user-written text that despite not necessarily describing the image may provide useful context that if properly leveraged can be informative. We address this task with a multimodal model that conditions on both textual information from the associated social media post as well as visual signal from the image, and demonstrate that the utility of these two information sources stacks. We put forward a new dataset of 371k images paired with alt-text and tweets scraped from Twitter and evaluate on it across a variety of automated metrics as well as human evaluation. We show that our approach of conditioning on both tweet text and visual information significantly outperforms prior work, by more than 2x on BLEU@4. Nikita Srivatsan, Sofía Samaniego, Omar Florez, Taylor Berg-Kirkpatrick |
ICLR | 4 |
| 2024 | DITTO: Diffusion Inference-Time T-Optimization for Music GenerationabstractWe propose Diffusion Inference-Time T-Optimization (DITTO), a general-purpose framework for controlling pre-trained text-to-music diffusion models at inference-time via optimizing initial noise latents. Our method can be used to optimize through any differentiable feature matching loss to achieve a target (stylized) output and leverages gradient checkpointing for memory efficiency. We demonstrate a surprisingly wide-range of applications for music generation including inpainting, outpainting, and looping as well as intensity, melody, and musical structure control – all without ever fine-tuning the underlying model. When we compare our approach against related training, guidance, and optimization-based methods, we find DITTO achieves state-of-the-art performance on nearly all tasks, including outperforming comparable approaches on controllability, audio quality, and computational efficiency, thus opening the door for high-quality, flexible, training-free control of diffusion models. Sound examples can be found at https://ditto-music.github.io/web/. Zachary Novack, Julian J. McAuley, Taylor Berg-Kirkpatrick, Nicholas J. Bryan |
ICML | 3 |
| 2024 | Retrieval Guided Music Captioning via Multimodal Prefixes
Nikita Srivatsan, Ke Chen 0021, Shlomo Dubnov, Taylor Berg-Kirkpatrick |
IJCAI | 4 |
| 2024 | Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation
Ke Chen 0021, Jiaqi Su, Taylor Berg-Kirkpatrick, Shlomo Dubnov, Zeyu Jin |
INTERSPEECH | 3 |
| 2024 | HYSYNTH: Context-Free LLM Approximation for Guiding Program SynthesisabstractMany structured prediction and reasoning tasks can be framed as program synthesis problems, where the goal is to generate a program in a \emph{domain-specific language} (DSL) that transforms input data into the desired output. Unfortunately, purely neural approaches, such as large language models (LLMs), often fail to produce fully correct programs in unfamiliar DSLs, while purely symbolic methods based on combinatorial search scale poorly to complex problems. Motivated by these limitations, we introduce a hybrid approach, where LLM completions for a given task are used to learn a task-specific, context-free surrogate model, which is then used to guide program synthesis. We evaluate this hybrid approach on three domains, and show that it outperforms both unguided search and direct sampling from LLMs, as well as existing program synthesizers. Shraddha Barke, Emmanuel Anaya Gonzalez, Saketh Ram Kasibatla, Taylor Berg-Kirkpatrick, Nadia Polikarpova |
NeurIPS | 4 |
| 2024 | Grammar-Aligned DecodingabstractLarge Language Models (LLMs) struggle with reliably generating highly structured outputs, such as program code, mathematical formulas, or well-formed markup. Constrained decoding approaches mitigate this problem by greedily restricting what tokens an LLM can output at each step to guarantee that the output matches a given constraint. Specifically, in grammar-constrained decoding (GCD), the LLM's output must follow a given grammar. In this paper we demonstrate that GCD techniques (and in general constrained decoding techniques) can distort the LLM's distribution, leading to outputs that are grammatical but appear with likelihoods that are not proportional to the ones given by the LLM, and so ultimately are low-quality. We call the problem of aligning sampling with a grammar constraint, grammar-aligned decoding (GAD), and propose adaptive sampling with approximate expected futures (ASAp), a decoding algorithm that guarantees the output to be grammatical while provably producing outputs that match the conditional probability of the LLM's distribution conditioned on the given grammar constraint. Our algorithm uses prior sample outputs to soundly overapproximate the future grammaticality of different output prefixes. Our evaluation on code generation and structured NLP tasks shows how ASAp often produces outputs with higher likelihood (according to the LLM's distribution) than existing GCD techniques, while still enforcing the desired grammatical constraints. Kanghee Park, Taylor Berg-Kirkpatrick, Nadia Polikarpova, Loris D'Antoni |
NeurIPS | 3 |
| 2024 | MAWI Rec: Leveraging Severe Weather Data in RecommendationabstractInferring user intent in recommender systems can help performance but is difficult because intent is personal and not directly observable. Previous work has leveraged signals to stand as a proxy for intent (e.g. user interactions with resource pages), but such signals are not always available. In this paper, we instead recognize that certain events, which are observable, directly influence user intent. For example, after a flood, home improvement customers are more likely to undertake a renovation project to dry out their basement. We introduce MAWI Rec, a recommender system that leverages severe weather data to improve recommendation. Our weather-aware system achieves a significant improvement over a state-of-the-art baseline for online and in-store datasets of home improvement customers. This gain is most significant for weather-related product categories such as roof panels and flashings. Brendan Andrew Duncan, Surya Kallumadi, Taylor Berg-Kirkpatrick, Julian J. McAuley |
RecSys | 3 |
| 2023 | Contrastive Attention Networks for Attribution of Early Modern PrintabstractIn this paper, we develop machine learning techniques to identify unknown printers in early modern (c.~1500--1800) English printed books. Specifically, we focus on matching uniquely damaged character type-imprints in anonymously printed books to works with known printers in order to provide evidence of their origins. Until now, this work has been limited to manual investigations by analytical bibliographers. We present a Contrastive Attention-based Metric Learning approach to identify similar damage across character image pairs, which is sensitive to very subtle differences in glyph shapes, yet robust to various confounding sources of noise associated with digitized historical books. To overcome the scarce amount of supervised data, we design a random data synthesis procedure that aims to simulate bends, fractures, and inking variations induced by the early printing process. Our method successfully improves downstream damaged type-imprint matching among printed works from this period, as validated by in-domain human experts. The results of our approach on two important philosophical works from the Early Modern period demonstrate potential to extend the extant historical research about the origins and content of these books. Nikolai Vogler, Kartik Goyal, Kishore PV Reddy, Elizaveta Pertseva, Samuel V. Lemley, Christopher N. Warren, Max G'Sell, Taylor Berg-Kirkpatrick |
AAAI | 8 |
| 2023 | Beyond Contrastive Learning: A Variational Generative Model for Multilingual RetrievalabstractJohn Wieting, Jonathan Clark, William Cohen, Graham Neubig, Taylor Berg-Kirkpatrick. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. John Wieting, Jonathan H. Clark, William W. Cohen, Graham Neubig, Taylor Berg-Kirkpatrick |
ACL (1) | 5 |
| 2023 | A Block Metropolis-Hastings Sampler for Controllable Energy-based Text GenerationabstractRecent work has shown that energy-based language modeling is an effective framework for controllable text generation because it enables flexible integration of arbitrary discriminators.However, because energy-based LMs are globally normalized, approximate techniques like Metropolis-Hastings (MH) are required for inference.Past work has largely explored simple proposal distributions that modify a single token at a time, like in Gibbs sampling.In this paper, we develop a novel MH sampler that, in contrast, proposes re-writes of the entire sequence in each step via iterative prompting of a large language model.Our new sampler (a) allows for more efficient and accurate sampling from a target distribution and (b) allows generation length to be determined through the sampling procedure rather than fixed in advance, as past work has required.We perform experiments on two controlled generation tasks, showing both downstream performance gains and more accurate target distribution sampling in comparison with single-token proposal techniques. Token-level sampling (Prior work)How are you?Utterance-level block sampling (Ours) Iteration i: How are you?Iteration i+1: How are you?Proposal: How art you?BERT single-token proposal accept / reject Iteration i: How are you?Iteration i+1: How art thou? Jarad Forristal, Niloofar Mireshghallah, Greg Durrett, Taylor Berg-Kirkpatrick |
CoNLL | 4 |
| 2023 | Simple Temporal Adaptation to Changing Label Sets: Hashtag Prediction via Dense KNNabstractUser-generated social media data is constantly changing as new trends influence online discussion and personal information is deleted due to privacy concerns.However, traditional NLP models rely on fixed training datasets, which means they are unable to adapt to temporal change-both test distribution shift and deleted training data-without frequent, costly re-training.In this paper, we study temporal adaptation through the task of longitudinal hashtag prediction and propose a nonparametric dense retrieval technique, which does not require re-training, as a simple but effective solution.In experiments on a newly collected, publicly available, year-long Twitter dataset exhibiting temporal distribution shift, our method improves by 64% over the best static parametric baseline while avoiding costly gradient-based re-training.Our approach is also particularly well-suited to dynamically deleted user data in line with data privacy laws, with negligible computational cost/performance loss. Niloofar Mireshghallah, Nikolai Vogler, Junxian He, Omar Florez, Ahmed El-Kishky, Taylor Berg-Kirkpatrick |
EMNLP | 6 |
| 2023 | Multitrack Music TransformerabstractExisting approaches for generating multitrack music with transformer models have been limited in terms of the number of instruments, the length of the music segments and slow inference. This is partly due to the memory requirements of the lengthy input sequences necessitated by existing representations. In this work, we propose a new multitrack music representation that allows a diverse set of instruments while keeping a short sequence length. Our proposed Multitrack Music Transformer (MMT) achieves comparable performance with state-of-the-art systems, landing in between two recently proposed models in a subjective listening test, while achieving substantial speedups and memory reductions over both, making the method attractive for real time improvisation or near real time creative applications. Further, we propose a new measure for analyzing musical self-attention and show that the trained model attends more to notes that form a consonant interval with the current note and to notes that are 4N beats away from the current step. Hao-Wen Dong, Ke Chen 0021, Shlomo Dubnov, Julian J. McAuley, Taylor Berg-Kirkpatrick |
ICASSP | 5 |
| 2023 | Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption AugmentationabstractContrastive learning has shown remarkable success in the field of multimodal representation learning. In this paper, we propose a pipeline of contrastive language-audio pretraining to develop an audio representation by combining audio data with natural language descriptions. To accomplish this target, we first release LAION-Audio-630K, a large collection of 633,526 audio-text pairs from different data sources. Second, we construct a contrastive language-audio pretraining model by considering different audio encoders and text encoders. We incorporate the feature fusion mechanism and keyword-to-caption augmentation into the model design to further enable the model to process audio inputs of variable lengths and enhance the performance. Third, we perform comprehensive experiments to evaluate our model across three tasks: text-to-audio retrieval, zero-shot audio classification, and supervised audio classification. The results demonstrate that our model achieves superior performance in text-to-audio retrieval task. In audio classification tasks, the model achieves state-of-the-art performance in the zero-shot setting and is able to obtain performance comparable to models’ results in the non-zero-shot setting. LAION-Audio-630K1and the proposed model2are both available to the public. Yusong Wu, Ke Chen 0021, Yuchen Hui, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
ICASSP | 5 |
| 2023 | EEBO-Verse: Sifting for Poetry in Large Early Modern Corpora Using Visual Features
Danlu Chen, Taylor Berg-Kirkpatrick |
ICDAR (5) | 3 |
| 2023 | CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos
Hao-Wen Dong, Naoya Takahashi, Yuki Mitsufuji, Julian J. McAuley, Taylor Berg-Kirkpatrick |
ICLR | 5 |
| 2022 | Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled DataabstractDeep learning techniques for separating audio into different sound sources face several challenges. Standard architectures require training separate models for different types of audio sources. Although some universal separators employ a single model to target multiple sources, they have difficulty generalizing to unseen sources. In this paper, we propose a three-component pipeline to train a universal audio source separator from a large, but weakly-labeled dataset: AudioSet. First, we propose a transformer-based sound event detection system for processing weakly-labeled training data. Second, we devise a query-based audio separation model that leverages this data for model training. Third, we design a latent embedding processor to encode queries that specify audio targets for separation, allowing for zero-shot generalization. Our approach uses a single model for source separation of multiple sound types, and relies solely on weakly-labeled data for training. In addition, the proposed audio separator can be used in a zero-shot setting, learning to separate types of audio sources that were never seen in training. To evaluate the separation performance, we test our model on MUSDB18, while training on the disjoint AudioSet. We further verify the zero-shot performance by conducting another experiment on audio source types that are held-out from training. The model achieves comparable Source-to-Distortion Ratio (SDR) performance to current supervised models in both cases. Ke Chen 0021, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
AAAI | 5 |
| 2022 | HOLM: Hallucinating Objects with Language Models for Referring Expression Recognition in Partially-Observed ScenesabstractAI systems embodied in the physical world face a fundamental challenge of partial observability; operating with only a limited view and knowledge of the environment.This creates challenges when AI systems try to reason about language and its relationship with the environment: objects referred to through language (e.g.giving many instructions) are not immediately visible.Actions by the AI system may be required to bring these objects in view.A good benchmark to study this challenge is Dynamic Referring Expression Recognition (dRER) task where the goal is to find a target location by dynamically adjusting the field of view (FoV) in a partially observed 360 • scenes.In this paper, we introduce HOLM, Hallucinating Objects with Language Models, to address the challenge of partial observability.HOLM uses large pre-trained language models (LMs) to infer object hallucinations for the unobserved part of the environment.Our core intuition is that if a pair of objects coappear in an environment frequently, our usage of language should reflect this fact about the world.Based on this intuition, we prompt language models to extract knowledge about object affinities which gives us a proxy for spatial relationships of objects.Our experiments show that HOLM performs better than the state-of-the-art approaches on two datasets for dRER; allowing to study generalization for both indoor and outdoor settings. Volkan Cirik, Louis-Philippe Morency, Taylor Berg-Kirkpatrick |
ACL (1) | 3 |
| 2022 | Achieving Conversational Goals with Unsupervised Post-hoc Knowledge InjectionabstractBodhisattwa Prasad Majumder, Harsh Jhamtani, Taylor Berg-Kirkpatrick, Julian McAuley. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Bodhisattwa Prasad Majumder, Harsh Jhamtani, Taylor Berg-Kirkpatrick, Julian J. McAuley |
ACL (1) | 3 |
| 2022 | Mix and Match: Learning-free Controllable Text Generationusing Energy Language ModelsabstractRecent work on controlled text generation has either required attribute-based fine-tuning of the base language model (LM), or has restricted the parameterization of the attribute discriminator to be compatible with the base autoregressive LM.In this work, we propose Mix and Match LM, a global score-based alternative for controllable text generation that combines arbitrary pre-trained black-box models for achieving the desired attributes in the generated text without involving any fine-tuning or structural assumptions about the black-box models.We interpret the task of controllable generation as drawing samples from an energy-based model whose energy values are a linear combination of scores from black-box models that are separately responsible for fluency, the control attribute, and faithfulness to any conditioning context.We use a Metropolis-Hastings sampling scheme to sample from this energy-based model using bidirectional context and global attribute features.We validate the effectiveness of our approach on various controlled generation and style-based text revision tasks by outperforming recently proposed methods that involve extra training, fine-tuning, or restrictive assumptions over the form of models. Niloofar Mireshghallah, Kartik Goyal, Taylor Berg-Kirkpatrick |
ACL (1) | 3 |
| 2022 | Quantifying Privacy Risks of Masked Language Models Using Membership Inference AttacksabstractThe wide adoption and application of Masked language models (MLMs) on sensitive data (from legal to medical) necessitates a thorough quantitative investigation into their privacy vulnerabilities.Prior attempts at measuring leakage of MLMs via membership inference attacks have been inconclusive, implying potential robustness of MLMs to privacy attacks.In this work, we posit that prior attempts were inconclusive because they based their attack solely on the MLM's model score.We devise a stronger membership inference attack based on likelihood ratio hypothesis testing that involves an additional reference MLM to more accurately quantify the privacy risks of memorization in MLMs.We show that masked language models are indeed susceptible to likelihood ratio membership inference attacks: Our empirical results, on models trained on medical notes, show that our attack improves the AUC of prior membership inference attacks from 0.66 to an alarmingly high 0.90 level. Niloofar Mireshghallah, Kartik Goyal, Archit Uniyal, Taylor Berg-Kirkpatrick, Reza Shokri |
EMNLP | 4 |
| 2022 | An Empirical Analysis of Memorization in Fine-tuned Autoregressive Language ModelsabstractSeveral recent works have shown that large language models present privacy risks through memorization of training data.Little attention, however, has been given to the fine-tuning phase and it is not well understood how memorization risk varies across different fine-tuning methods (such as fine-tuning the full model, the model head, and adapter).This presents increasing concern as the "pre-train and fine-tune" paradigm proliferates.We empirically study memorization of fine-tuning methods using membership inference and extraction attacks, and show that their susceptibility to attacks is very different.We observe that fine-tuning the head of the model has the highest susceptibility to attacks, whereas fine-tuning smaller adapters appears to be less vulnerable to known extraction attacks. Niloofar Mireshghallah, Archit Uniyal, Tianhao Wang 0001, David Evans 0001, Taylor Berg-Kirkpatrick |
EMNLP | 5 |
| 2022 | HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and DetectionabstractAudio classification is an important task of mapping audio samples into their corresponding labels. Recently, the transformer model with self-attention mechanisms has been adopted in this field. However, existing audio transformers require large GPU memories and long training time, meanwhile relying on pretrained vision models to achieve high performance, which limits the model’s scalability in audio tasks. To combat these problems, we introduce HTS-AT: an audio transformer with a hierarchical structure to reduce the model size and training time. It is further combined with a token-semantic module to map final outputs into class featuremaps, thus enabling the model for the audio event detection (i.e. localization in time). We evaluate HTS-AT on three datasets of audio classification where it achieves new state-of-the-art (SOTA) results on AudioSet and ESC50, and equals the SOTA on Speech Command V2. It also achieves better performance in event localization than the previous CNN-based models. Moreover, HTS-AT requires only 35% model parameters and 15% training time of the previous audio transformer. These results demonstrate the high performance and high efficiency of HTS-AT. Ke Chen 0021, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
ICASSP | 5 |
| 2022 | Tonet: Tone-Octave Network for Singing Melody Extraction from Polyphonic MusicabstractSinging melody extraction is an important problem in the field of music information retrieval. Existing methods typically rely on frequency-domain representations to estimate the sung frequencies. However, this design does not lead to human-level performance in the perception of melody information for both tone (pitch-class) and octave. In this paper, we propose TONet1, a plug-and-play model that improves both tone and octave perceptions by leveraging a novel input representation and a novel network architecture. First, we present an improved input representation, the Tone-CFP, that explicitly groups harmonics via a rearrangement of frequency-bins. Second, we introduce an encoder-decoder architecture that is designed to obtain a salience feature map, a tone feature map, and an octave feature map. Third, we propose a tone-octave fusion mechanism to improve the final salience feature map. Experiments are done to verify the capability of TONet with various baseline backbone models. Our results show that tone-octave fusion with Tone-CFP can significantly improve the singing voice extraction performance across various datasets – with substantial gains in octave and tone accuracy. Ke Chen 0021, Shuai Yu 0002, Cheng-i Wang, Wei Li 0012, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
ICASSP | 5 |
| 2022 | Deep Performer: Score-to-Audio Music Performance SynthesisabstractMusic performance synthesis aims to synthesize a musical score into a natural performance. In this paper, we borrow recent advances in text-to-speech synthesis and present the Deep Performer—a novel system for score-to-audio music performance synthesis. Unlike speech, music often contains polyphony and long notes. Hence, we propose two new techniques for handling polyphonic inputs and providing a fine-grained conditioning in a transformer encoder-decoder model. To train our proposed system, we present a new violin dataset consisting of paired recordings and scores along with estimated alignments between them. We show that our proposed model can synthesize music with clear polyphony and harmonic structures. In a listening test, we achieve competitive quality against the baseline model, a conditional generative audio model, in terms of pitch accuracy, timbre and noise level. Moreover, our proposed model significantly outperforms the baseline on an existing piano dataset in overall quality. Hao-Wen Dong, Taylor Berg-Kirkpatrick, Julian J. McAuley |
ICASSP | 3 |
| 2022 | Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--Hastings
Kartik Goyal, Chris Dyer, Taylor Berg-Kirkpatrick |
ICLR | 3 |
| 2022 | Towards a Unified View of Parameter-Efficient Transfer Learning
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, Graham Neubig |
ICLR | 4 |
| 2022 | UserIdentifier: Implicit User Representations for Simple and Effective Personalized Sentiment AnalysisabstractFatemehsadat Mireshghallah, Vaishnavi Shrivastava, Milad Shokouhi, Taylor Berg-Kirkpatrick, Robert Sim, Dimitrios Dimitriadis. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Niloofar Mireshghallah, Vaishnavi Shrivastava, Milad Shokouhi, Taylor Berg-Kirkpatrick, Robert Sim, Dimitrios Dimitriadis |
NAACL-HLT | 4 |
| 2021 | Efficient Nearest Neighbor Language ModelsabstractNon-parametric neural language models (NLMs) learn predictive distributions of text utilizing an external datastore, which allows them to learn through explicitly memorizing the training datapoints.While effective, these models often require retrieval from a large datastore at test time, significantly increasing the inference overhead and thus limiting the deployment of non-parametric NLMs in practical applications.In this paper, we take the recently proposed k-nearest neighbors language model (Khandelwal et al., 2019) as an example, exploring methods to improve its efficiency along various dimensions.Experiments on the standard WikiText-103 benchmark and domain-adaptation datasets show that our methods are able to achieve up to a 6x speed-up in inference speed while retaining comparable performance.The empirical analysis we present may provide guidelines for future research seeking to develop or deploy more efficient non-parametric NLMs. 1 Junxian He, Graham Neubig, Taylor Berg-Kirkpatrick |
EMNLP (1) | 3 |
| 2021 | Truth-Conditional Captions for Time Series DataabstractIn this paper, we explore the task of automatically generating natural language descriptions of salient patterns in a time series, such as stock prices of a company over a week.A model for this task should be able to extract high-level patterns such as presence of a peak or a dip.While typical contemporary neural models with attention mechanisms can generate fluent output descriptions for this task, they often generate factually incorrect descriptions.We propose a computational model with a truth-conditional architecture which first runs small learned programs on the input time series, then identifies the programs/patterns which hold true for the given input, and finally conditions on only the chosen valid program (rather than the input time series) to generate the output text description.A program in our model is constructed from modules, which are small neural networks that are designed to capture numerical patterns and temporal information.The modules are shared across multiple programs, enabling compositionality as well as efficient learning of module parameters.The modules, as well as the composition of the modules, are unobserved in data, and we learn them in an end-to-end fashion with the only training signal coming from the accompanying natural language text descriptions.We find that the proposed model is able to generate high-precision captions even though we consider a small and simple space of module types. Harsh Jhamtani, Taylor Berg-Kirkpatrick |
EMNLP (1) | 2 |
| 2021 | Investigating Robustness of Dialog Models to Popular Figurative Language ConstructsabstractHumans often employ figurative language use in communication, including during interactions with dialog systems.Thus, it is important for real-world dialog systems to be able to handle popular figurative language constructs like metaphor and simile.In this work, we analyze the performance of existing dialog models in situations where the input dialog context exhibits use of figurative language.We observe large gaps in handling of figurative language when evaluating the models on two open domain dialog datasets.When faced with dialog contexts consisting of figurative language, some models show very large drops in performance compared to contexts without figurative language.We encourage future research in dialog modeling to separately analyze and report results on figurative language in order to better test model capabilities relevant to real-world use.Finally, we propose lightweight solutions to help existing models become more robust to figurative language by simply using an external resource to translate figurative language to literal (non-figurative) forms while preserving the meaning to the best extent possible. Harsh Jhamtani, Varun Gangal, Eduard H. Hovy, Taylor Berg-Kirkpatrick |
EMNLP (1) | 4 |
| 2021 | Style Pooling: Automatic Text Style Obfuscation for Improved Classification FairnessabstractText style can reveal sensitive attributes of the author (e.g.race or age) to the reader, which can, in turn, lead to privacy violations and bias in both human and algorithmic decisions based on text.For example, the style of writing in job applications might reveal protected attributes of the candidate which could lead to bias in hiring decisions, regardless of whether hiring decisions are made algorithmically or by humans.We propose a VAE-based framework that obfuscates stylistic features of human-generated text through style transfer by automatically re-writing the text itself.Our framework operationalizes the notion of obfuscated style in a flexible way that enables two distinct notions of obfuscated style: (1) a minimal notion that effectively intersects the various styles seen in training, and (2) a maximal notion that seeks to obfuscate by adding stylistic features of all sensitive attributes to text, in effect, computing a union of styles.Our style-obfuscation framework can be used for multiple purposes, however, we demonstrate its effectiveness in improving the fairness of downstream classifiers.We also conduct a comprehensive study on style pooling's effect on fluency, semantic consistency, and attribute removal from text, in two and three domain style obfuscation. 1 Niloofar Mireshghallah, Taylor Berg-Kirkpatrick |
EMNLP (1) | 2 |
| 2021 | Scalable Font Reconstruction with Dual Latent ManifoldsabstractWe propose a deep generative model that performs typography analysis and font reconstruction by learning disentangled manifolds of both font style and character shape.Our approach enables us to massively scale up the number of character types we can effectively model compared to previous methods.Specifically, we infer separate latent variables representing character and font via a pair of inference networks which take as input sets of glyphs that either all share a character type, or belong to the same font.This design allows our model to generalize to characters that were not observed during training time, an important task in light of the relative sparsity of most fonts.We also put forward a new loss, adapted from prior work that measures likelihood using an adaptive distribution in a projected space, resulting in more natural images without requiring a discriminator.We evaluate on the task of font reconstruction over various datasets representing character types of many languages, and compare favorably to modern style transfer systems according to both automatic and manually-evaluated metrics. Nikita Srivatsan, Jonathan T. Barron, Taylor Berg-Kirkpatrick |
EMNLP (1) | 4 |
| 2021 | Privacy Regularization: Joint Privacy-Utility Optimization in LanguageModelsabstractFatemehsadat Mireshghallah, Huseyin Inan, Marcello Hasegawa, Victor Rühle, Taylor Berg-Kirkpatrick, Robert Sim. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Niloofar Mireshghallah, Huseyin A. Inan, Marcello Hasegawa, Victor Rühle, Taylor Berg-Kirkpatrick, Robert Sim |
NAACL-HLT | 5 |
| 2020 | Refer360$^\circ$: A Referring Expression Recognition Dataset in 360$^\circ$ ImagesabstractWe propose a novel large-scale referring expression recognition dataset, Refer360°, consisting of 17,137 instruction sequences and ground-truth actions for completing these instructions in 360° scenes.Refer360° differs from existing related datasets in three ways.First, we propose a more realistic scenario where instructors and the followers have partial, yet dynamic, views of the scene -followers continuously modify their field-of-view (FoV) while interpreting instructions that specify a final target location.Second, instructions to find the target location consist of multiple steps for followers who will start at random FoVs.As a result, intermediate instructions are strongly grounded in object references and followers must identify intermediate FoVs to find the final target location correctly.Third, the target locations are neither restricted to predefined objects nor chosen by annotators; instead, they are distributed randomly across scenes.This "point anywhere" approach leads to more linguistically complex instructions, as shown in our analyses.Our examination of the dataset shows that Refer360° manifests linguistically rich phenomena in a language grounding task that poses novel challenges for computational modeling of language, vision, and navigation. Volkan Cirik, Taylor Berg-Kirkpatrick, Louis-Philippe Morency |
ACL | 2 |
| 2020 | A Probabilistic Generative Model for Typographical Analysis of Early Modern PrintingabstractWe propose a deep and interpretable probabilistic generative model to analyze glyph shapes in printed Early Modern documents.We focus on clustering extracted glyph images into underlying templates in the presence of multiple confounding sources of variance.Our approach introduces a neural editor model that first generates well-understood printing phenomena like spatial perturbations from template parameters via interpertable latent variables, and then modifies the result by generating a non-interpretable latent vector responsible for inking variations, jitter, noise from the archiving process, and other unforeseen phenomena associated with Early Modern printing.Critically, by introducing an inference network whose input is restricted to the visual residual between the observation and the interpretably-modified template, we are able to control and isolate what the vector-valued latent variable captures.We show that our approach outperforms rigid interpretable clustering baselines (Ocular) and overly-flexible deep generative models (VAE) alike on the task of completely unsupervised discovery of typefaces in mixed-font documents. Kartik Goyal, Chris Dyer, Christopher N. Warren, Max G'Sell, Taylor Berg-Kirkpatrick |
ACL | 5 |
| 2020 | Phonetic and Visual Priors for Decipherment of Informal RomanizationabstractInformal romanization is an idiosyncratic process used by humans in informal digital communication to encode non-Latin script languages into Latin character sets found on common keyboards. Character substitution choices differ between users but have been shown to be governed by the same main principles observed across a variety of languages—namely, character pairs are often associated through phonetic or visual similarity. We propose a noisy-channel WFST cascade model for deciphering the original non-Latin script from observed romanized text in an unsupervised fashion. We train our model directly on romanized data from two languages: Egyptian Arabic and Russian. We demonstrate that adding inductive bias through phonetic and visual priors on character mappings substantially improves the model’s performance on both languages, yielding results much closer to the supervised skyline. Finally, we introduce a new dataset of romanized Russian, collected from a Russian social network website and partially annotated for our experiments. Maria Ryskina, Matthew R. Gormley, Taylor Berg-Kirkpatrick |
ACL | 3 |
| 2020 | Domain Adaptation via Context Prediction for Engineering Diagram Search
Harsh Jhamtani, Taylor Berg-Kirkpatrick |
ECIR (2) | 2 |
| 2020 | An Empirical Investigation of Contextualized Number PredictionabstractWe conduct a large scale empirical investigation of contextualized number prediction in running text.Specifically, we consider two tasks: (1) masked number prediction -predicting a missing numerical value within a sentence, and (2) numerical anomaly detectiondetecting an errorful numeric value within a sentence.We experiment with novel combinations of contextual encoders and output distributions over the real number line.Specifically, we introduce a suite of output distribution parameterizations that incorporate latent variables to add expressivity and better fit the natural distribution of numeric values in running text, and combine them with both recurrent and transformer-based encoder architectures.We evaluate these models on two numeric datasets in the financial and scientific domain.Our findings show that output distributions that incorporate discrete latent variables and allow for multiple modes outperform simple flow-based counterparts on all datasets, yielding more accurate numerical prediction and anomaly detection.We also show that our models effectively utilize textual context and benefit from general-purpose unsupervised pretraining. 1 Taylor Berg-Kirkpatrick, Daniel Spokoyny |
EMNLP (1) | 1 |
| 2020 | Like hiking? You probably enjoy nature: Persona-grounded Dialog with Commonsense ExpansionsabstractExisting persona-grounded dialog models often fail to capture simple implications of given persona descriptions, something which humans are able to do seamlessly.For example, state-of-the-art models cannot infer that interest in hiking might imply love for nature or longing for a break.In this paper, we propose to expand available persona sentences using existing commonsense knowledge bases and paraphrasing resources to imbue dialog models with access to an expanded and richer set of persona descriptions.Additionally, we introduce fine-grained grounding on personas by encouraging the model to make a discrete choice among persona sentences while synthesizing a dialog response.Since such a choice is not observed in the data, we model it using a discrete latent random variable and use variational learning to sample from hundreds of persona expansions.Our model outperforms competitive baselines on the PERSONA-CHAT dataset in terms of dialog quality and diversity while achieving persona-consistent and controllable dialog generation. Bodhisattwa Prasad Majumder, Harsh Jhamtani, Taylor Berg-Kirkpatrick, Julian J. McAuley |
EMNLP (1) | 3 |
| 2020 | A Bilingual Generative Transformer for Semantic Sentence EmbeddingabstractSemantic sentence embedding models encode natural language sentences into vectors, such that closeness in embedding space indicates closeness in the semantics between the sentences.Bilingual data offers a useful signal for learning such embeddings: properties shared by both sentences in a translation pair are likely semantic, while divergent properties are likely stylistic or language-specific.We propose a deep latent variable model that attempts to perform source separation on parallel sentences, isolating what they have in common in a latent semantic vector, and explaining what is left over with language-specific latent vectors.Our proposed approach differs from past work on semantic sentence encoding in two ways.First, by using a variational probabilistic framework, we introduce priors that encourage source separation, and can use our model's posterior to predict sentence embeddings for monolingual data at test time.Second, we use high-capacity transformers as both data generating distributions and inference networkscontrasting with most past work on sentence embeddings.In experiments, our approach substantially outperforms the state-of-the-art on a standard suite of unsupervised semantic similarity evaluations.Further, we demonstrate that our approach yields the largest gains on more difficult subsets of these evaluations where simple word overlap is not a good indicator of similarity. 1 John Wieting, Graham Neubig, Taylor Berg-Kirkpatrick |
EMNLP (1) | 3 |
| 2020 | A Probabilistic Formulation of Unsupervised Text Style Transfer
Junxian He, Xinyi Wang 0001, Graham Neubig, Taylor Berg-Kirkpatrick |
ICLR | 4 |
| 2020 | Learning Sparse Prototypes for Text GenerationabstractPrototype-driven text generation uses non-parametric models that first choose from a library of sentence "prototypes" and then modify the prototype to generate the output text. While effective, these methods are inefficient at test time as a result of needing to store and index the entire training corpus. Further, existing methods often require heuristics to identify which prototypes to reference at training time. In this paper, we propose a novel generative model that automatically learns a sparse prototype support set that, nonetheless, achieves strong language modeling performance. This is achieved by (1) imposing a sparsity-inducing prior on the prototype selection distribution, and (2) utilizing amortized variational inference to learn a prototype retrieval function. In experiments, our model outperforms previous prototype-driven language models while achieving up to a 1000x memory reduction, as well as a 1000x speed-up at test time. More interestingly, we show that the learned prototypes are able to capture semantics and syntax at different granularity as we vary the sparsity of prototype selection, and that certain sentence attributes can be controlled by specifying the prototype for generation. Junxian He, Taylor Berg-Kirkpatrick, Graham Neubig |
NeurIPS | 2 |
| 2019 | Cross-Lingual Syntactic Transfer through Unsupervised Adaptation of Invertible ProjectionsabstractCross-lingual transfer is an effective way to build syntactic analysis tools in low-resource languages.However, transfer is difficult when transferring to typologically distant languages, especially when neither annotated target data nor parallel corpora are available.In this paper, we focus on methods for cross-lingual transfer to distant languages and propose to learn a generative model with a structured prior that utilizes labeled source data and unlabeled target data jointly.The parameters of source model and target model are softly shared through a regularized log likelihood objective.An invertible projection is employed to learn a new interlingual latent embedding space that compensates for imperfect crosslingual word embedding input.We evaluate our method on two syntactic tasks: part-ofspeech (POS) tagging and dependency parsing.On the Universal Dependency Treebanks, we use English as the only source corpus and transfer to a wide range of target languages.On the 10 languages in this dataset that are distant from English, our method yields an average of 5.2% absolute improvement on POS tagging and 8.3% absolute improvement on dependency parsing over a direct transfer method using state-of-the-art discriminative models. 1 3 Following Ahmad et al. (2019), we use the offline pre-trained alignment matrix present in https://github.com/Babylonpartners/ fastText_multilingual, which contains alignment matrices for 78 languages, which also allows comparison with their numbers in Section 4.3. Junxian He, Zhisong Zhang, Taylor Berg-Kirkpatrick, Graham Neubig |
ACL (1) | 3 |
| 2019 | Beyond BLEU: Training Neural Machine Translation with Semantic SimilarityabstractWhile most neural machine translation (NMT) systems are still trained using maximum likelihood estimation, recent work has demonstrated that optimizing systems to directly improve evaluation metrics such as BLEU can substantially improve final translation accuracy.However, training with BLEU has some limitations: it doesn't assign partial credit, it has a limited range of output values, and it can penalize semantically correct hypotheses if they differ lexically from the reference.In this paper, we introduce an alternative reward function for optimizing NMT systems that is based on recent work in semantic similarity.We evaluate on four disparate languages translated to English, and find that training with our proposed metric results in better translations as evaluated by BLEU, semantic similarity, and human evaluation, and also that the optimization procedure converges faster.Analysis suggests that this is because the proposed metric is more conducive to optimization, assigning partial credit and providing more diversity in scores than BLEU. 1 John Wieting, Taylor Berg-Kirkpatrick, Kevin Gimpel, Graham Neubig |
ACL (1) | 2 |
| 2019 | Simple and Effective Paraphrastic Similarity from Parallel TranslationsabstractWe present a model and methodology for learning paraphrastic sentence embeddings directly from bitext, removing the timeconsuming intermediate step of creating paraphrase corpora.Further, we show that the resulting model can be applied to cross-lingual tasks where it both outperforms and is orders of magnitude faster than more complex stateof-the-art baselines.1 John Wieting, Kevin Gimpel, Graham Neubig, Taylor Berg-Kirkpatrick |
ACL (1) | 4 |
| 2019 | Learning Rhyming Constraints using Structured AdversariesabstractHarsh Jhamtani, Sanket Vaibhav Mehta, Jaime Carbonell, Taylor Berg-Kirkpatrick. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Harsh Jhamtani, Sanket Vaibhav Mehta, Jaime G. Carbonell, Taylor Berg-Kirkpatrick |
EMNLP/IJCNLP (1) | 4 |
| 2019 | A Surprisingly Effective Fix for Deep Latent Variable Modeling of TextabstractBohan Li, Junxian He, Graham Neubig, Taylor Berg-Kirkpatrick, Yiming Yang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Junxian He, Graham Neubig, Taylor Berg-Kirkpatrick, Yiming Yang 0002 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | A Deep Factorization of Style and Structure in FontsabstractNikita Srivatsan, Jonathan Barron, Dan Klein, Taylor Berg-Kirkpatrick. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Nikita Srivatsan, Jonathan T. Barron, Daniel Klein 0001, Taylor Berg-Kirkpatrick |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Lagging Inference Networks and Posterior Collapse in Variational Autoencoders
Junxian He, Daniel Spokoyny, Graham Neubig, Taylor Berg-Kirkpatrick |
ICLR (Poster) | 4 |
| 2018 | Using Syntax to Ground Referring Expressions in Natural ImagesabstractWe introduce GroundNet, a neural network for referring expression recognition---the task of localizing (or grounding) in an image the object referred to by a natural language expression. Our approach to this task is the first to rely on a syntactic analysis of the input referring expression in order to inform the structure of the computation graph. Given a parse tree for an input expression, we explicitly map the syntactic constituents and relationships present in the tree to a composed graph of neural modules that defines our architecture for performing localization. This syntax-based approach aids localization of both the target object and auxiliary supporting objects mentioned in the expression. As a result, GroundNet is more interpretable than previous methods: we can (1) determine which phrase of the referring expression points to which object in the image and (2) track how the localization of the target object is determined by the network. We study this property empirically by introducing a new set of annotations on the GoogleRef dataset to evaluate localization of supporting objects. Our experiments show that GroundNet achieves state-of-the-art accuracy in identifying supporting objects, while maintaining comparable performance in the localization of target objects. Volkan Cirik, Taylor Berg-Kirkpatrick, Louis-Philippe Morency |
AAAI | 2 |
| 2018 | A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence ModelsabstractBeam search is a desirable choice of test-time decoding algorithm for neural sequence models because it potentially avoids search errors made by simpler greedy methods. However, typical cross entropy training procedures for these models do not directly consider the behaviour of the final decoding method. As a result, for cross-entropy trained models, beam decoding can sometimes yield reduced test performance when compared with greedy decoding. In order to train models that can more effectively make use of beam search, we propose a new training procedure that focuses on the final loss metric (e.g. Hamming loss) evaluated on the output of beam search. While well-defined, this "direct loss" objective is itself discontinuous and thus difficult to optimize. Hence, in our approach, we form a sub-differentiable surrogate objective by introducing a novel continuous approximation of the beam search decoding procedure.In experiments, we show that optimizing this new training objective yields substantially better results on two sequence tasks (Named Entity Recognition and CCG Supertagging) when compared with both cross entropy trained greedy decoding and cross entropy trained beam decoding baselines. Kartik Goyal, Graham Neubig, Chris Dyer, Taylor Berg-Kirkpatrick |
AAAI | 4 |
| 2018 | SPINE: SParse Interpretable Neural EmbeddingsabstractPrediction without justification has limited utility. Much of the success of neural models can be attributed to their ability to learn rich, dense and expressive representations. While these representations capture the underlying complexity and latent trends in the data, they are far from being interpretable. We propose a novel variant of denoising k-sparse autoencoders that generates highly efficient and interpretable distributed word representations (word embeddings), beginning with existing word representations from state-of-the-art methods like GloVe and word2vec. Through large scale human evaluation, we report that our resulting word embedddings are much more interpretable than the original GloVe and word2vec embeddings. Moreover, our embeddings outperform existing popular word embeddings on a diverse suite of benchmark downstream tasks. Anant Subramanian, Danish Pruthi, Harsh Jhamtani, Taylor Berg-Kirkpatrick, Eduard H. Hovy |
AAAI | 4 |
| 2018 | Learning to Generate Move-by-Move Commentary for Chess Games from Large-Scale Social Forum DataabstractHarsh Jhamtani, Varun Gangal, Eduard Hovy, Graham Neubig, Taylor Berg-Kirkpatrick. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Harsh Jhamtani, Varun Gangal, Eduard H. Hovy, Graham Neubig, Taylor Berg-Kirkpatrick |
ACL (1) | 5 |
| 2018 | Unsupervised Learning of Syntactic Structure with Invertible Neural ProjectionsabstractUnsupervised learning of syntactic structure is typically performed using generative models with discrete latent variables and multinomial parameters.In most cases, these models have not leveraged continuous word representations.In this work, we propose a novel generative model that jointly learns discrete syntactic structure and continuous word representations in an unsupervised fashion by cascading an invertible neural network with a structured generative prior.We show that the invertibility condition allows for efficient exact inference and marginal likelihood computation in our model so long as the prior is well-behaved.In experiments we instantiate our approach with both Markov and tree-structured priors, evaluating on two tasks: part-of-speech (POS) induction, and unsupervised dependency parsing without gold POS annotation.On the Penn Treebank, our Markov-structured model surpasses state-of-the-art results on POS induction.Similarly, we find that our tree-structured model achieves state-of-the-art performance on unsupervised dependency parsing for the difficult training condition where neither gold POS annotation nor punctuation-based constraints are available. Junxian He, Graham Neubig, Taylor Berg-Kirkpatrick |
EMNLP | 3 |
| 2018 | Learning to Describe Differences Between Pairs of Similar ImagesabstractIn this paper, we introduce the task of automatically generating text to describe the differences between two similar images.We collect a new dataset by crowd-sourcing difference descriptions for pairs of image frames extracted from video-surveillance footage.Annotators were asked to succinctly describe all the differences in a short paragraph.As a result, our novel dataset provides an opportunity to explore models that align language and vision, and capture visual salience.The dataset may also be a useful benchmark for coherent multi-sentence generation.We perform a firstpass visual analysis that exposes clusters of differing pixels as a proxy for object-level differences.We propose a model that captures visual salience by using a latent variable to align clusters of differing pixels with output sentences.We find that, for both single-sentence generation and as well as multi-sentence generation, the proposed model outperforms the models that use attention alone. Harsh Jhamtani, Taylor Berg-Kirkpatrick |
EMNLP | 2 |
| 2018 | Modeling Online Discourse with Coupled Distributed TopicsabstractIn this paper, we propose a deep, globally normalized topic model that incorporates structural relationships connecting documents in socially generated corpora, such as online forums.Our model (1) captures discursive interactions along observed reply links in addition to traditional topic information, and (2) incorporates latent distributed representations arranged in a deep architecture, which enables a GPU-based mean-field inference procedure that scales efficiently to large data.We apply our model to a new social media dataset consisting of 13M comments mined from the popular internet forum Reddit, a domain that poses significant challenges to models that do not account for relationships connecting user comments.We evaluate against existing methods across multiple metrics including perplexity and metadata prediction, and qualitatively analyze the learned interaction patterns. Nikita Srivatsan, Zachary Wojtowicz, Taylor Berg-Kirkpatrick |
EMNLP | 3 |
| 2018 | Speaker-Follower Models for Vision-and-Language NavigationabstractNavigation guided by natural language instructions presents a challenging reasoning problem for instruction followers. Natural language instructions typically identify only a few high-level decisions and landmarks rather than complete low-level motor behaviors; much of the missing information must be inferred based on perceptual context. In machine learning settings, this is doubly challenging: it is difficult to collect enough annotated data to enable learning of this reasoning process from scratch, and also difficult to implement the reasoning process using generic sequence models. Here we describe an approach to vision-and-language navigation that addresses both these issues with an embedded speaker model. We use this speaker model to (1) synthesize new instructions for data augmentation and to (2) implement pragmatic reasoning, which evaluates how well candidate action sequences explain an instruction. Both steps are supported by a panoramic action space that reflects the granularity of human-generated instructions. Experiments show that all three components of this approach---speaker-driven data augmentation, pragmatic reasoning and panoramic action space---dramatically improve the performance of a baseline instruction follower, more than doubling the success rate over the best existing approach on a standard benchmark. Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Daniel Klein 0001, Trevor Darrell |
NeurIPS | 7 |
| 2018 | Unsupervised Text Style Transfer using Language Models as DiscriminatorsabstractBinary classifiers are employed as discriminators in GAN-based unsupervised style transfer models to ensure that transferred sentences are similar to sentences in the target domain. One difficulty with the binary discriminator is that error signal is sometimes insufficient to train the model to produce rich-structured language. In this paper, we propose a technique of using a target domain language model as the discriminator to provide richer, token-level feedback during the learning process. Because our language model scores sentences directly using a product of locally normalized probabilities, it offers more stable and more useful training signal to the generator. We train the generator to minimize the negative log likelihood (NLL) of generated sentences evaluated by a language model. By using continuous approximation of the discrete samples, our model can be trained using back-propagation in an end-to-end way. Moreover, we find empirically with a language model as a structured discriminator, it is possible to eliminate the adversarial training steps using negative samples, thus making training more stable. We compare our model with previous work using convolutional neural networks (CNNs) as discriminators and show our model outperforms them significantly in three tasks including word substitution decipherment, sentiment modification and related language translation. Zhiting Hu, Chris Dyer, Eric P. Xing, Taylor Berg-Kirkpatrick |
NeurIPS | 5 |
| 2017 | Identifying Products in Online Cybercrime Marketplaces: A Dataset for Fine-grained Domain AdaptationabstractGreg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Rebecca Portnoff, Sadia Afroz, Damon McCoy, Kirill Levchenko, Vern Paxson. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2017. Greg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Rebecca S. Portnoff, Sadia Afroz 0001, Damon McCoy, Kirill Levchenko, Vern Paxson |
EMNLP | 3 |
| 2017 | Improved Variational Autoencoders for Text Modeling using Dilated ConvolutionsabstractRecent work on generative text modeling has found that variational autoencoders (VAE) with LSTM decoders perform worse than simpler LSTM language models (Bowman et al., 2015). This negative result is so far poorly understood, but has been attributed to the propensity of LSTM decoders to ignore conditioning information from the encoder. In this paper, we experiment with a new type of decoder for VAE: a dilated CNN. By changing the decoder’s dilation architecture, we control the size of context from previously generated words. In experiments, we find that there is a trade-off between contextual capacity of the decoder and effective use of encoding information. We show that when carefully managed, VAEs can outperform LSTM language models. We demonstrate perplexity gains on two datasets, representing the first positive language modeling result with VAE. Further, we conduct an in-depth investigation of the use of VAE (with our new decoding architecture) for semi-supervised and unsupervised labeling tasks, demonstrating gains over several strong baselines. Zhiting Hu, Ruslan Salakhutdinov, Taylor Berg-Kirkpatrick |
ICML | 4 |
| 2017 | Efficient Correlated Topic Modeling with Topic EmbeddingabstractCorrelated topic modeling has been limited to small model and problem sizes due to their high computational cost and poor scaling. In this paper, we propose a new model which learns compact topic embeddings and captures topic correlations through the closeness between the topic vectors. Our method enables efficient inference in the low-dimensional embedding space, reducing previous cubic or quadratic time complexity to linear w.r.t the topic size. We further speedup variational inference with a fast sampler to exploit sparsity of topic occurrence. Extensive experiments show that our approach is capable of handling model and data scales which are several orders of magnitude larger than existing correlation results, without sacrificing modeling quality by providing competitive or superior performance in document classification and retrieval. Junxian He, Zhiting Hu, Taylor Berg-Kirkpatrick, Eric P. Xing |
KDD | 3 |
| 2017 | Tools for Automated Analysis of Cybercriminal MarketsabstractUnderground forums are widely used by criminals to buy and sell a host of stolen items, datasets, resources, and criminal services. These forums contain important resources for understanding cybercrime. However, the number of forums, their size, and the domain expertise required to understand the markets makes manual exploration of these forums unscalable. In this work, we propose an automated, top-down approach for analyzing underground forums. Our approach uses natural language processing and machine learning to automatically generate high-level information about underground forums, first identifying posts related to transactions, and then extracting products and prices. We also demonstrate, via a pair of case studies, how an analyst can use these automated approaches to investigate other categories of products and transactions. We use eight distinct forums to assess our tools: Antichat, Blackhat World, Carders, Darkode, Hack Forums, Hell, L33tCrew and Nulled. Our automated approach is fast and accurate, achieving over 80% accuracy in detecting post category, product, and prices. Rebecca S. Portnoff, Sadia Afroz 0001, Greg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Damon McCoy, Kirill Levchenko, Vern Paxson |
WWW | 5 |
| 2016 | Learning-Based Single-Document Summarization with Compression and Anaphoricity ConstraintsabstractWe present a discriminative model for single-document summarization that integrally combines compression and anaphoricity constraints.Our model selects textual units to include in the summary based on a rich set of sparse features whose weights are learned on a large corpus.We allow for the deletion of content within a sentence when that deletion is licensed by compression rules; in our framework, these are implemented as dependencies between subsentential units of text.Anaphoricity constraints then improve cross-sentence coherence by guaranteeing that, for each pronoun included in the summary, the pronoun's antecedent is included as well or the pronoun is rewritten as a full mention.When trained end-to-end, our final system 1 outperforms prior work on both ROUGE as well as on human judgments of linguistic quality. Greg Durrett, Taylor Berg-Kirkpatrick, Daniel Klein 0001 |
ACL (1) | 2 |
| 2015 | An Empirical Analysis of Optimization for Max-Margin NLPabstractDespite the convexity of structured maxmargin objectives (Taskar et al., 2004;Tsochantaridis et al., 2004), the many ways to optimize them are not equally effective in practice.We compare a range of online optimization methods over a variety of structured NLP tasks (coreference, summarization, parsing, etc) and find several broad trends.First, margin methods do tend to outperform both likelihood and the perceptron.Second, for max-margin objectives, primal optimization methods are often more robust and progress faster than dual methods.This advantage is most pronounced for tasks with dense or continuous-valued features.Overall, we argue for a particularly simple online primal subgradient descent method that, despite being rarely mentioned in the literature, is surprisingly effective in relation to its alternatives. Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Daniel Klein 0001 |
EMNLP | 2 |
| 2015 | GPU-Friendly Local Regression for Voice ConversionabstractVoice conversion is the task of transforming a source speaker's voice so that it sounds like a target speaker's voice.We present a GPUfriendly local regression model for voice conversion that is capable of converting speech in real-time and achieves state-of-the-art accuracy on this task.Our model uses a new approximation for computing local regression coefficients that is explicitly designed to preserve memory locality.As a result, our inference procedure is amenable to efficient implementation on the GPU.Our approach is more than 10X faster than a highly optimized CPUbased implementation, and is able to convert speech 2.7X faster than real-time. Taylor Berg-Kirkpatrick, Daniel Klein 0001 |
HLT-NAACL | 1 |
| 2015 | Unsupervised Code-Switching for Multilingual Historical Document TranscriptionabstractDan Garrette, Hannah Alpert-Abrams, Taylor Berg-Kirkpatrick, Dan Klein. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Dan Garrette, Hannah Alpert-Abrams, Taylor Berg-Kirkpatrick, Daniel Klein 0001 |
HLT-NAACL | 3 |
| 2014 | Sparser, Better, Faster GPU ParsingabstractDue to their origin in computer graphics, graphics processing units (GPUs) are highly optimized for dense problems, where the exact same operation is applied repeatedly to all data points.Natural language processing algorithms, on the other hand, are traditionally constructed in ways that exploit structural sparsity.Recently, Canny et al. (2013) presented an approach to GPU parsing that sacrifices traditional sparsity in exchange for raw computational power, obtaining a system that can compute Viterbi parses for a high-quality grammar at about 164 sentences per second on a mid-range GPU.In this work, we reintroduce sparsity to GPU parsing by adapting a coarse-to-fine pruning approach to the constraints of a GPU.The resulting system is capable of computing over 404 Viterbi parses per second-more than a 2x speedup-on the same hardware.Moreover, our approach allows us to efficiently implement less GPU-friendly minimum Bayes risk inference, improving throughput for this more accurate algorithm from only 32 sentences per second unpruned to over 190 sentences per second using pruning-nearly a 6x speedup. David Hall 0006, Taylor Berg-Kirkpatrick, Daniel Klein 0001 |
ACL (1) | 2 |
| 2014 | Unsupervised Transcription of Piano Music
Taylor Berg-Kirkpatrick, Jacob Andreas, Daniel Klein 0001 |
NIPS | 1 |
| 2013 | Unsupervised Transcription of Historical Documents
Taylor Berg-Kirkpatrick, Greg Durrett, Daniel Klein 0001 |
ACL (1) | 1 |
| 2013 | Decipherment with a Million Random RestartsabstractThis paper investigates the utility and effect of running numerous random restarts when using EM to attack decipherment problems.We find that simple decipherment models are able to crack homophonic substitution ciphers with high accuracy if a large number of random restarts are used but almost completely fail with only a few random restarts.For particularly difficult homophonic ciphers, we find that big gains in accuracy are to be had by running upwards of 100K random restarts, which we accomplish efficiently using a GPU-based parallel implementation.We run a series of experiments using millions of random restarts in order to investigate other empirical properties of decipherment problems, including the famously uncracked Zodiac 340. Taylor Berg-Kirkpatrick, Daniel Klein 0001 |
EMNLP | 1 |
| 2013 | Learning Whom to Trust with MACE
Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, Eduard H. Hovy |
HLT-NAACL | 2 |
| 2012 | An Empirical Investigation of Statistical Significance in NLP
Taylor Berg-Kirkpatrick, David Burkett, Daniel Klein 0001 |
EMNLP-CoNLL | 1 |
| 2011 | Jointly Learning to Extract and Compress
Taylor Berg-Kirkpatrick, Daniel Gillick, Daniel Klein 0001 |
ACL | 1 |
| 2011 | Simple Effective Decipherment via Combinatorial Optimization
Taylor Berg-Kirkpatrick, Daniel Klein 0001 |
EMNLP | 1 |
| 2010 | Phylogenetic Grammar Induction
Taylor Berg-Kirkpatrick, Daniel Klein 0001 |
ACL | 1 |
| 2010 | Painless Unsupervised Learning with Features
Taylor Berg-Kirkpatrick, Alexandre Bouchard-Côté, John DeNero, Daniel Klein 0001 |
HLT-NAACL | 1 |
| 2008 | Learning Bilingual Lexicons from Monolingual Corpora
Aria Haghighi, Percy Liang, Taylor Berg-Kirkpatrick, Daniel Klein 0001 |
ACL | 3 |