VLDB 2026 Research / reviewers in the wild / expert
Robin Jia
dblp:182/2556
· DBLP profile ↗
36ranked-venue papers
4as first author
28since 2021 · last 2026
0009-0002-8123-7132ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 3 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language ModelsabstractWoody Haosheng Gan, Deqing Fu, Julian Asilis, Ollie Liu, Vatsal Sharan, Robin Jia, Willie Neiswanger. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Woody Haosheng Gan, Deqing Fu, Julian Asilis, Ollie Liu, Vatsal Sharan, Robin Jia, Willie Neiswanger |
ACL (1) | 6 |
| 2025 | Why Do Some Inputs Break Low-Bit LLM Quantization?abstractLow-bit weight-only quantization significantly reduces the memory footprint of large language models (LLMs), but disproportionately affects certain examples.We analyze diverse 3-4 bit methods on LLMs ranging from 7B-70B in size and find that the quantization errors of 50 pairs of methods are strongly correlated (avg.ρ = 0.82) on FineWeb examples.Moreover, the residual stream magnitudes of full-precision models are indicative of future quantization errors.We further establish a hypothesis that relates the residual stream magnitudes to error amplification and accumulation over layers.Using LLM localization techniques, early exiting, and activation patching, we show that examples with large errors rely on precise residual activations in the later layers, and that the outputs of MLP gates play a crucial role in maintaining the perplexity.Our work reveals why certain examples result in large quantization errors and which model components are most critical for performance preservation. Ting-Yun Chang, Muru Zhang, Jesse Thomason, Robin Jia |
EMNLP | 4 |
| 2025 | Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual QueriesabstractTo answer one-to-many factual queries (e.g., listing cities of a country), a language model (LM) must simultaneously recall knowledge and avoid repeating previous answers. How are these two subtasks implemented and integrated internally? Across multiple datasets, models, and prompt templates, we identify a promote-then-suppress mechanism: the model first recalls all answers, and then suppresses previously generated ones. Specifically, LMs use both the subject and previous answer tokens to perform knowledge recall, with attention propagating subject information and MLPs promoting the answers. Then, attention attends to and suppresses previous answer tokens, while MLPs amplify the suppression signal. Our mechanism is corroborated by extensive experimental evidence: in addition to using early decoding and causal tracing, we analyze how components use different tokens by introducing both Token Lens, which decodes aggregated attention updates from specified tokens, and a knockout method that analyzes changes in MLP outputs after removing attention to specified tokens. Overall, we provide new insights into how LMs’ internal components interact with different input tokens to support complex factual recall. Tianyi Lorena Yan, Robin Jia |
EMNLP | 2 |
| 2025 | Rethinking Backdoor Detection Evaluation for Language ModelsabstractBackdoor attacks, in which a model behaves maliciously when given an attacker-specified trigger, pose a major security risk for practitioners who depend on publicly released language models.As a countermeasure, backdoor detection methods aim to detect whether a released model contains a backdoor.While existing backdoor detection methods have high accuracy in detecting backdoored models on standard benchmarks, it is unclear whether they can robustly identify backdoors in the wild.In this paper, we examine the robustness of backdoor detectors by manipulating different factors during backdoor planting.We find that the success of existing methods based on trigger inversion or meta classifiers highly depends on how intensely the model is trained on poisoned data.Specifically, backdoors planted with more aggressive or more conservative training are significantly more difficult to detect than the default ones.Our results highlight a lack of robustness of existing backdoor detectors and the limitations in current benchmark construction. Jun Yan 0012, Wenjie Mo 0001, Xiang Ren 0001, Robin Jia |
EMNLP | 4 |
| 2025 | TLDR: Token-Level Detective Reward Model for Large Vision Language ModelsabstractAlthough reward models have been successful in improving multimodal large language models, the reward models themselves remain brutal and contain minimal information. Notably, existing reward models only mimic human annotations by assigning only one feedback to any text, no matter how long the text is. In the realm of multimodal language models, where models are required to process both images and texts, a naive reward model may learn implicit biases toward texts and become less grounded in images. In this paper, we propose a **T**oken-**L**evel **D**etective **R**eward Model (**TLDR**) to provide fine-grained annotations to each text token. We first introduce a perturbation-based method to generate synthetic hard negatives and their token-level labels to train TLDR models. Then we show the rich usefulness of TLDR models both in assisting off-the-shelf models to self-correct their generations, and in serving as a hallucination evaluation tool. We show that TLDR automatically trains a token-level likelihood optimization, and can improve the base model's performance significantly. Finally, we show that TLDR models can significantly speed up human annotation by 3 times to acquire a broader range of high-quality vision language data. Deqing Fu, Tong Xiao 0003, Wang Zhu 0001, Pengchuan Zhang, Guan Pang, Robin Jia, Lawrence Chen 0002 |
ICLR | 7 |
| 2025 | Language Models Can Infer Action Semantics for Symbolic Planners from Environment FeedbackabstractWang Bill Zhu, Ishika Singh, Robin Jia, Jesse Thomason. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Wang Zhu 0001, Ishika Singh, Robin Jia, Jesse Thomason |
NAACL (Long Papers) | 3 |
| 2024 | Operationalizing Content Moderation "Accuracy" in the Digital Services ActabstractThe Digital Services Act, recently adopted by the EU, requires social media platforms to report the ``accuracy'' of their automated content moderation systems. The colloquial term is vague, or open-textured---the literal accuracy (number of correct predictions divided by the total) is not suitable for problems with large class imbalance, and the ground truth and dataset to measure accuracy against is unspecified. Without further specification, the regulatory requirement allows for deficient reporting. In this interdisciplinary work, we operationalize ``accuracy'' reporting by refining legal concepts and relating them to technical implementation. We start by elucidating the legislative purpose of the Act to legally justify an interpretation of ``accuracy'' as precision and recall. These metrics remain informative in class imbalanced settings, and reflect the proportional balancing of Fundamental Rights of the EU Charter. We then focus on the estimation of recall, as its naive estimation can incur extremely high annotation costs and disproportionately interfere with the platform's right to conduct business. Through a simulation study, we show that recall can be efficiently estimated using stratified sampling with trained classifiers, and provide concrete recommendations for its application. Finally, we present a case study of recall reporting for a subset of Reddit under the Act. Based on the language in the Act, we identify a number of ways recall could be reported due to underspecification. We report on one possibility using our improved estimator, and discuss the implications and areas for further legal clarification. Johnny Tian-Zheng Wei, Frederike Zufall, Robin Jia |
AIES (1) | 3 |
| 2024 | When Parts Are Greater Than Sums: Individual LLM Components Can Outperform Full ModelsabstractThis paper studies in-context learning by decomposing the output of large language models into the individual contributions of attention heads and MLPs (components).We observe curious components: good-performing ones that individually do well on a classification task, even when the full model performs poorly; bad-performing ones that do much worse than chance; and label-biased components that always predict the same label.We find that component accuracies are well-correlated across different demonstration sets and perturbations of prompt templates.Based on our findings, we propose component reweighting, which learns to linearly re-scale the component activations from a few labeled examples.Given 24 labeled examples, our method improves by an average of 6.0% accuracy points over 24-shot ICL across 8 tasks on Llama-2-7B.Overall, this paper both enriches our understanding of ICL and provides a practical method for improvement by examining model internals.Template 1 {text}\nIs this a piece of news regarding World, Sports, Business, or Technology?Template 2 {text} Is this a piece of news regarding World, Sports, Business, or Technology?Correlation = 0.81 Ting-Yun Chang, Jesse Thomason, Robin Jia |
EMNLP | 3 |
| 2024 | Efficient End-to-End Visual Document Understanding with Rationale DistillationabstractWang Zhu, Alekh Agarwal, Mandar Joshi, Robin Jia, Jesse Thomason, Kristina Toutanova. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Wang Zhu 0001, Alekh Agarwal, Mandar Joshi, Robin Jia, Jesse Thomason, Kristina Toutanova |
NAACL-HLT | 4 |
| 2024 | Do Localization Methods Actually Localize Memorized Data in LLMs? A Tale of Two BenchmarksabstractTing-Yun Chang, Jesse Thomason, Robin Jia. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Ting-Yun Chang, Jesse Thomason, Robin Jia |
NAACL-HLT | 3 |
| 2024 | Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear RegressionabstractTransformers excel at *in-context learning* (ICL)---learning from demonstrations without parameter updates---but how they do so remains a mystery. Recent work suggests that Transformers may internally run Gradient Descent (GD), a first-order optimization method, to perform ICL. In this paper, we instead demonstrate that Transformers learn to approximate second-order optimization methods for ICL. For in-context linear regression, Transformers share a similar convergence rate as *Iterative Newton's Method*, both *exponentially* faster than GD. Empirically, predictions from successive Transformer layers closely match different iterations of Newton’s Method linearly, with each middle layer roughly computing 3 iterations; thus, Transformers and Newton’s method converge at roughly the same rate. In contrast, Gradient Descent converges exponentially more slowly. We also show that Transformers can learn in-context on ill-conditioned data, a setting where Gradient Descent struggles but Iterative Newton succeeds. Finally, to corroborate our empirical findings, we prove that Transformers can implement $k$ iterations of Newton's method with $k + \mathcal O(1)$ layers. Deqing Fu, Robin Jia, Vatsal Sharan |
NeurIPS | 3 |
| 2024 | Pre-trained Large Language Models Use Fourier Features to Compute AdditionabstractPre-trained large language models (LLMs) exhibit impressive mathematical reasoning capabilities, yet how they compute basic arithmetic, such as addition, remains unclear.
This paper shows that pre-trained LLMs add numbers using Fourier features---dimensions in the hidden state that represent numbers via a set of features sparse in the frequency domain.
Within the model, MLP and attention layers use Fourier features in complementary ways: MLP layers primarily approximate the magnitude of the answer using low-frequency features, while attention layers primarily perform modular addition (e.g., computing whether the answer is even or odd) using high-frequency features.
Pre-training is crucial for this mechanism: models trained from scratch to add numbers only exploit low-frequency features, leading to lower accuracy.
Introducing pre-trained token embeddings to a randomly initialized model rescues its performance.
Overall, our analysis demonstrates that appropriate pre-trained representations (e.g., Fourier features) can unlock the ability of Transformers to learn precise mechanisms for algorithmic tasks. Tianyi Zhou 0011, Deqing Fu, Vatsal Sharan, Robin Jia |
NeurIPS | 4 |
| 2023 | Data Curation Alone Can Stabilize In-context LearningabstractIn-context learning (ICL) enables large language models (LLMs) to perform new tasks by prompting them with a sequence of training examples.However, it is known that ICL is very sensitive to the choice of training examples: randomly sampling examples from a training set leads to high variance in performance.In this paper, we show that carefully curating a subset of training data greatly stabilizes ICL performance without any other changes to the ICL algorithm (e.g., prompt retrieval or calibration).We introduce two methods to choose training subsets-both score training examples individually, then select the highest-scoring ones.CONDACC scores a training example by its average dev-set ICL accuracy when combined with random training examples, while DATAMODELS learns linear regressors that estimate how the presence of each training example influences LLM outputs.Across five tasks and two LLMs, sampling from stable subsets selected by CONDACC and DATAMODELS improves average accuracy over sampling from the entire training set by 7.7% and 6.3%, respectively.Surprisingly, the stable subset examples are not especially diverse in content or low in perplexity, in contrast with other work suggesting that diversity and perplexity are important when prompting LLMs. Ting-Yun Chang, Robin Jia |
ACL (1) | 2 |
| 2023 | Do Question Answering Modeling Improvements Hold Across Benchmarks?abstractDo question answering (QA) modeling improvements (e.g., choice of architecture and training procedure) hold consistently across the diverse landscape of QA benchmarks?To study this question, we introduce the notion of concurrence-two benchmarks have high concurrence on a set of modeling approaches if they rank the modeling approaches similarly.We measure the concurrence between 32 QA benchmarks on a set of 20 diverse modeling approaches and find that human-constructed benchmarks have high concurrence amongst themselves, even if their passage and question distributions are very different.Surprisingly, even downsampled human-constructed benchmarks (i.e., collecting less data) and programmatically-generated benchmarks (e.g., cloze-formatted examples) have high concurrence with human-constructed benchmarks.These results indicate that, despite years of intense community focus on a small number of benchmarks, the modeling improvements studied hold broadly. Nelson F. Liu, Robin Jia, Percy Liang |
ACL (1) | 3 |
| 2023 | Contrastive Novelty-Augmented Learning: Anticipating Outliers with Large Language ModelsabstractIn many task settings, text classification models are likely to encounter examples from novel classes on which they cannot predict correctly.Selective prediction, in which models abstain on low-confidence examples, provides a possible solution, but existing models are often overly confident on unseen classes.To remedy this overconfidence, we introduce Contrastive Novelty-Augmented Learning (CoNAL), a twostep method that generates OOD examples representative of novel classes, then trains to decrease confidence on them.First, we generate OOD examples by prompting a large language model twice: we prompt it to enumerate relevant novel classes, then generate examples from each novel class matching the task format.Second, we train a classifier with a novel contrastive objective that encourages lower confidence on generated OOD examples than training examples.When trained with CoNAL, classifiers improve in their ability to detect and abstain on novel class examples over prior methods by an average of 2.3% in terms of accuracy under the accuracy-coverage curve (AUAC) and 5.5% AUROC across 4 NLP datasets, with no cost to in-distribution accuracy.1 Albert Xu, Xiang Ren 0001, Robin Jia |
ACL (1) | 3 |
| 2023 | Chain-of-Questions Training with Latent Answers for Robust Multistep Question AnsweringabstractWe propose Chain-of-Questions, a framework that trains a model to robustly answer multistep questions by generating and answering sub-questions.We obtain supervision for subquestions from human-annotated question decomposition meaning representation (QDMR), but QDMR does not include annotated answers to sub-questions.To overcome this technical challenge, we treat sub-answers as latent variables and infer them with a novel dynamic mixture of Hard-EM and MAPO.Chain-of-Questions is effective and robust, greatly outperforming strong neuro-symbolic methods by 9.0 F1 on a DROP contrast set and GPT-3.5 by 24.3 F1 on a HOTPOTQA adversarial set. Wang Zhu 0001, Jesse Thomason, Robin Jia |
EMNLP | 3 |
| 2023 | SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative ExamplesabstractDetecting negatives (such as non-entailment relationships, unanswerable questions, and false claims) is an important and challenging aspect of many natural language understanding tasks.Though manually collecting challenging negative examples can help models detect them, it is both costly and domain-specific.In this work, we propose Self-labeled Counterfactuals for Extrapolating to Negative Examples (SCENE), an automatic method for synthesizing training data that greatly improves models' ability to detect challenging negative examples.In contrast with standard data augmentation, which synthesizes new examples for existing labels, SCENE can synthesize negative examples zero-shot from only positive ones.Given a positive example, SCENE perturbs it with a mask infilling model, then determines whether the resulting example is negative based on a self-training heuristic.With access to only answerable training examples, SCENE can close 69.6% of the performance gap on SQuAD 2.0, a dataset where half of the evaluation examples are unanswerable, compared to a model trained on SQuAD 2.0.Our method also extends to boolean question answering and recognizing textual entailment, and improves generalization from SQuAD to ACE-whQA, an out-of-domain extractive QA benchmark. Deqing Fu, Ameya Godbole, Robin Jia |
EMNLP | 3 |
| 2022 | On Continual Model Refinement in Out-of-Distribution Data StreamsabstractReal-world natural language processing (NLP) models need to be continually updated to fix the prediction errors in out-of-distribution (OOD) data streams while overcoming catastrophic forgetting.However, existing continual learning (CL) problem setups cannot cover such a realistic and complex scenario.In response to this, we propose a new CL problem formulation dubbed continual model refinement (CMR).Compared to prior CL settings, CMR is more practical and introduces unique challenges (boundary-agnostic and non-stationary distribution shift, diverse mixtures of multiple OOD data clusters, error-centric streams, etc.).We extend several existing CL approaches to the CMR setting and evaluate them extensively.For benchmarking and analysis, we propose a general sampling algorithm to obtain dynamic OOD data streams with controllable nonstationarity, as well as a suite of metrics measuring various aspects of online performance.Our experiments and detailed analysis reveal the promise and challenges of the CMR problem, supporting that studying CMR in dynamic OOD streams can benefit the longevity of deployed NLP models in production. 1 Bill Y. Lin, Sida I. Wang, Xi Victoria Lin, Robin Jia, Xiang Ren 0001, Scott Yih |
ACL (1) | 4 |
| 2022 | Knowledge Base Question Answering by Case-based Reasoning over SubgraphsabstractQuestion answering (QA) over knowledge bases (KBs) is challenging because of the diverse, essentially unbounded, types of reasoning patterns needed. However, we hypothesize in a large KB, reasoning patterns required to answer a query type reoccur for various entities in their respective subgraph neighborhoods. Leveraging this structural similarity between local neighborhoods of different subgraphs, we introduce a semiparametric model (CBR-SUBG) with (i) a nonparametric component that for each query, dynamically retrieves other similar $k$-nearest neighbor (KNN) training queries along with query-specific subgraphs and (ii) a parametric component that is trained to identify the (latent) reasoning patterns from the subgraphs of KNN queries and then apply them to the subgraph of the target query. We also propose an adaptive subgraph collection strategy to select a query-specific compact subgraph, allowing us to scale to full Freebase KB containing billions of facts. We show that CBR-SUBG can answer queries requiring subgraph reasoning patterns and performs competitively with the best models on several KBQA benchmarks. Our subgraph collection strategy also produces more compact subgraphs (e.g. 55% reduction in size for WebQSP while increasing answer recall by 4.85%)\footnote{Code, model, and subgraphs are available at \url{https://github.com/rajarshd/CBR-SUBG}}. Rajarshi Das, Ameya Godbole, Ankita Naik, Elliot Tower, Manzil Zaheer, Hannaneh Hajishirzi, Robin Jia, Andrew McCallum |
ICML | 7 |
| 2022 | Models in the Loop: Aiding Crowdworkers with Generative Annotation AssistantsabstractMax Bartolo, Tristan Thrush, Sebastian Riedel, Pontus Stenetorp, Robin Jia, Douwe Kiela. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Max Bartolo, Tristan Thrush, Sebastian Riedel 0001, Pontus Stenetorp, Robin Jia, Douwe Kiela |
NAACL-HLT | 5 |
| 2022 | On the Robustness of Reading Comprehension Models to Entity RenamingabstractJun Yan, Yang Xiao, Sagnik Mukherjee, Bill Yuchen Lin, Robin Jia, Xiang Ren. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jun Yan 0012, Sagnik Mukherjee, Bill Y. Lin, Robin Jia, Xiang Ren 0001 |
NAACL-HLT | 5 |
| 2021 | Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?abstractPedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, Jordan Boyd-Graber. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Pedro Rodríguez 0001, Joe Barrow, Alexander Miserlis Hoyle, John Lalor, Robin Jia, Jordan L. Boyd-Graber |
ACL/IJCNLP (1) | 5 |
| 2021 | The statistical advantage of automatic NLG metrics at the system levelabstractJohnny Wei, Robin Jia. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Johnny Tian-Zheng Wei, Robin Jia |
ACL/IJCNLP (1) | 2 |
| 2021 | Improving Question Answering Model Robustness with Synthetic Adversarial Data GenerationabstractDespite recent progress, state-of-the-art question answering models remain vulnerable to a variety of adversarial attacks. While dynamic adversarial data collection, in which a human annotator tries to write examples that fool a model-in-the-loop, can improve model robustness, this process is expensive which limits the scale of the collected data. In this work, we are the first to use synthetic adversarial data generation to make question answering models more robust to human adversaries. We develop a data generation pipeline that selects source passages, identifies candidate answers, generates questions, then finally filters or re-labels them to improve quality. Using this approach, we amplify a smaller human-written adversarial dataset to a much larger set of synthetic question-answer pairs. By incorporating our synthetic data, we improve the state-of-the-art on the AdversarialQA dataset by 3.7F1 and improve model generalisation on nine of the twelve MRQA datasets. We further conduct a novel human-in-the-loop evaluation to show that our models are considerably more robust to new human-written adversarial examples: crowdworkers can fool our model only 8.8% of the time on average, compared to 17.6% for a model trained without synthetic data. Max Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel 0001, Pontus Stenetorp, Douwe Kiela |
EMNLP (1) | 3 |
| 2021 | Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleabstractA possible explanation for the impressive performance of masked language model (MLM) pre-training is that such models have learned to represent the syntactic structures prevalent in classical NLP pipelines.In this paper, we propose a different explanation: MLMs succeed on downstream tasks mostly due to their ability to model higher-order word cooccurrence statistics.To demonstrate this, we pre-train MLMs on sentences with randomly shuffled word order, and we show that these models still achieve high accuracy after finetuning on many downstream tasks -including tasks specifically designed to be challenging for models that ignore word order.Our models also perform surprisingly well according to some parametric syntactic probes, indicating possible deficiencies in how we test representations for syntactic information.Overall, our results show that purely distributional information largely explains the success of pretraining, and they underscore the importance of curating challenging evaluation datasets that require deeper linguistic knowledge. Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, Douwe Kiela |
EMNLP (1) | 2 |
| 2021 | Dynabench: Rethinking Benchmarking in NLPabstractDouwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel 0001, Zeerak Talat, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams |
NAACL-HLT | 16 |
| 2021 | Swords: A Benchmark for Lexical Substitution with Improved Data Coverage and QualityabstractMina Lee, Chris Donahue, Robin Jia, Alexander Iyabor, Percy Liang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Mina Lee 0002, Chris Donahue, Robin Jia, Alexander Iyabor, Percy Liang |
NAACL-HLT | 3 |
| 2021 | Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation BenchmarkingabstractWe introduce Dynaboard, an evaluation-as-a-service framework for hosting benchmarks and conducting holistic model comparison, integrated with the Dynabench platform. Our platform evaluates NLP models directly instead of relying on self-reported metrics or predictions on a single dataset. Under this paradigm, models are submitted to be evaluated in the cloud, circumventing the issues of reproducibility, accessibility, and backwards compatibility that often hinder benchmarking in NLP. This allows users to interact with uploaded models in real time to assess their quality, and permits the collection of additional metrics such as memory use, throughput, and robustness, which -- despite their importance to practitioners -- have traditionally been absent from leaderboards. On each task, models are ranked according to the Dynascore, a novel utility-based aggregation of these statistics, which users can customize to better reflect their preferences, placing more/less weight on a particular axis of evaluation or dataset. As state-of-the-art NLP models push the limits of traditional benchmarks, Dynaboard offers a standardized solution for a more diverse and comprehensive evaluation of model quality. Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain, Ledell Wu, Robin Jia, Christopher Potts, Adina Williams, Douwe Kiela |
NeurIPS | 6 |
| 2020 | Robust Encodings: A Framework for Combating Adversarial TyposabstractDespite excellent performance on many tasks, NLP systems are easily fooled by small adversarial perturbations of inputs.Existing procedures to defend against such perturbations are either (i) heuristic in nature and susceptible to stronger attacks or (ii) provide guaranteed robustness to worst-case attacks, but are incompatible with state-of-the-art models like BERT.In this work, we introduce robust encodings (RobEn): a simple framework that confers guaranteed robustness, without making compromises on model architecture.The core component of RobEn is an encoding function, which maps sentences to a smaller, discrete space of encodings.Systems using these encodings as a bottleneck confer guaranteed robustness with standard training, and the same encodings can be used across multiple tasks.We identify two desiderata to construct robust encoding functions: perturbations of a sentence should map to a small set of encodings (stability), and models using encodings should still perform well (fidelity).We instantiate RobEn to defend against a large family of adversarial typos.Across six tasks from GLUE, our instantiation of RobEn paired with BERT achieves an average robust accuracy of 71.3% against all adversarial typos in the family considered, while previous work using a typo-corrector achieves only 35.3% accuracy against a simple greedy attack. Erik Jones, Robin Jia, Aditi Raghunathan, Percy Liang |
ACL | 2 |
| 2020 | Selective Question Answering under Domain ShiftabstractTo avoid giving wrong answers, question answering (QA) models need to know when to abstain from answering.Moreover, users often ask questions that diverge from the model's training data, making errors more likely and thus abstention more critical.In this work, we propose the setting of selective question answering under domain shift, in which a QA model is tested on a mixture of in-domain and out-of-domain data, and must answer (i.e., not abstain on) as many questions as possible while maintaining high accuracy.Abstention policies based solely on the model's softmax probabilities fare poorly, since models are overconfident on out-of-domain inputs.Instead, we train a calibrator to identify inputs on which the QA model errs, and abstain when it predicts an error is likely.Crucially, the calibrator benefits from observing the model's behavior on out-of-domain data, even if from a different domain than the test data.We combine this method with a SQuADtrained QA model and evaluate on mixtures of SQuAD and five other QA datasets.Our method answers 56% of questions while maintaining 80% accuracy; in contrast, directly using the model's probabilities only answers 48% at 80% accuracy. Amita Kamath, Robin Jia, Percy Liang |
ACL | 2 |
| 2020 | With Little Power Comes Great ResponsibilityabstractDespite its importance to experimental design, statistical power (the probability that, given a real effect, an experiment will reject the null hypothesis) has largely been ignored by the NLP community.Underpowered experiments make it more difficult to discern the difference between statistical noise and meaningful model improvements, and increase the chances of exaggerated findings.By metaanalyzing a set of existing NLP papers and datasets, we characterize typical power for a variety of settings and conclude that underpowered experiments are common in the NLP literature.In particular, for several tasks in the popular GLUE benchmark, small test sets mean that most attempted comparisons to state of the art models will not be adequately powered.Similarly, based on reasonable assumptions, we find that the most typical experimental design for human rating studies will be underpowered to detect small model differences, of the sort that are frequently studied.For machine translation, we find that typical test sets of 2000 sentences have approximately 75% power to detect differences of 1 BLEU point.To improve the situation going forward, we give an overview of best practices for power analysis in NLP and release a series of notebooks to assist with future power analyses.1 Dallas Card, Peter Henderson 0002, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, Daniel Jurafsky |
EMNLP (1) | 4 |
| 2019 | Certified Robustness to Adversarial Word SubstitutionsabstractRobin Jia, Aditi Raghunathan, Kerem Göksel, Percy Liang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Robin Jia, Aditi Raghunathan, Kerem Göksel, Percy Liang |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Delete, Retrieve, Generate: a Simple Approach to Sentiment and Style TransferabstractJuncen Li, Robin Jia, He He, Percy Liang. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Juncen Li, Robin Jia, He He 0001, Percy Liang |
NAACL-HLT | 2 |
| 2017 | Adversarial Examples for Evaluating Reading Comprehension SystemsabstractStandard accuracy metrics indicate that reading comprehension systems are making rapid progress, but the extent to which these systems truly understand language remains unclear.To reward systems with real language understanding abilities, we propose an adversarial evaluation scheme for the Stanford Question Answering Dataset (SQuAD).Our method tests whether systems can answer questions about paragraphs that contain adversarially inserted sentences, which are automatically generated to distract computer systems without changing the correct answer or misleading humans.In this adversarial setting, the accuracy of sixteen published models drops from an average of 75% F1 score to 36%; when the adversary is allowed to add ungrammatical sequences of words, average accuracy on four models decreases further to 7%.We hope our insights will motivate the development of new models that understand language more precisely. Robin Jia, Percy Liang |
EMNLP | 1 |
| 2017 | Learning concepts through conversations in spoken dialogue systemsabstractSpoken dialogue systems must be able to recover gracefully from unexpected user inputs. In many cases, these unexpected utterances may be within the scope of the system, but include previously unseen phrases that the system cannot interpret. In this work, we augment a spoken dialogue system with the ability to learn about new concepts by conversing with the user in natural language. We present a novel model that detects phrases corresponding to such concepts, using information from a neural slotfiller as well as syntactic cues. The system then prompts the user for a definition of the detected phrases, and uses these definitions to re-parse the original utterance. We demonstrate significant gains by learning from the user, compared to a baseline system. Robin Jia, Larry Heck, Dilek Hakkani-Tür, Georgi Nikolov |
ICASSP | 1 |
| 2016 | Data Recombination for Neural Semantic ParsingabstractModeling crisp logical regularities is crucial in semantic parsing, making it difficult for neural models with no task-specific prior knowledge to achieve good results.In this paper, we introduce data recombination, a novel framework for injecting such prior knowledge into a model.From the training data, we induce a highprecision synchronous context-free grammar, which captures important conditional independence properties commonly found in semantic parsing.We then train a sequence-to-sequence recurrent network (RNN) model with a novel attention-based copying mechanism on datapoints sampled from this grammar, thereby teaching the model about these structural properties.Data recombination improves the accuracy of our RNN model on three semantic parsing datasets, leading to new state-of-the-art performance on the standard GeoQuery dataset for models with comparable supervision. Robin Jia, Percy Liang |
ACL (1) | 1 |