Mandar Joshi

dblp:85/1261 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
11since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2024 On Scaling Up a Multilingual Vision and Language Model
abstract
We explore the boundaries of scaling up a multilingual vision and language model, both in terms of size of the components and the breadth of its training task mixture. Our model achieves new levels of performance on a wide-range of varied and complex tasks, including multiple image-based captioning and question-answering tasks, image-based document understanding and few-shot (in-context) learning, as well as object detection, video question answering, and video captioning. Our model advances the state-of-the-art on most vision-and-language benchmarks considered (20+ of them). Finally, we observe emerging capabilities, such as complex counting and multilingual object detection, tasks that are not explicitly in the training mix.
Xi Chen 0071, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Carlos Riquelme, Sebastian Goodman, Xiao Wang 0038, Yi Tay, Siamak Shakeri, Mostafa Dehghani 0001, Daniel Salz, Mario Lucic, Michael Tschannen, Arsha Nagrani, Hexiang Hu, Mandar Joshi, Bo Pang 0001, Ceslee Montgomery, Paulina Pietrzyk, Marvin Ritter, A. J. Piergiovanni, Matthias Minderer, Filip Pavetic, Austin Waters, Gang Li 0021, Ibrahim Alabdulmohsin, Lucas Beyer, Julien Amelot, Kenton Lee, Andreas Steiner 0001, Yang Li 0058, Daniel Keysers, Anurag Arnab, Yuanzhong Xu, Keran Rong, Alexander Kolesnikov 0003, Mojtaba Seyedhosseini, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, Radu Soricut
CVPR18
2024 BAGEL: Bootstrapping Agents by Guiding Exploration with Language
abstract
Following natural language instructions by executing actions in digital environments (e.g. web-browsers and REST APIs) is a challenging task for language model (LM) agents. Unfortunately, LM agents often fail to generalize to new environments without human demonstrations. This work presents BAGEL, a method for bootstrapping LM agents without human supervision. BAGEL converts a seed set of randomly explored trajectories to synthetic demonstrations via round-trips between two noisy LM components: an LM labeler which converts a trajectory into a synthetic instruction, and a zero-shot LM agent which maps the synthetic instruction into a refined trajectory. By performing these round-trips iteratively, BAGEL quickly converts the initial distribution of trajectories towards those that are well-described by natural language. We adapt the base LM agent at test time with in-context learning by retrieving relevant BAGEL demonstrations based on the instruction, and find improvements of over 2-13% absolute on ToolQA and MiniWob++, with up to 13x reduction in execution failures.
Shikhar Murty, Christopher D. Manning, Peter Shaw 0004, Mandar Joshi, Kenton Lee
ICML4
2024 Efficient End-to-End Visual Document Understanding with Rationale Distillation
abstract
Wang Zhu, Alekh Agarwal, Mandar Joshi, Robin Jia, Jesse Thomason, Kristina Toutanova. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Wang Zhu 0001, Alekh Agarwal, Mandar Joshi, Robin Jia, Jesse Thomason, Kristina Toutanova
NAACL-HLT3
2023 MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering
abstract
Fangyu Liu, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Yasemin Altun, Nigel Collier, Julian Eisenschlos. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Fangyu Liu 0001, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Yasemin Altun, Nigel Collier, Julian Martin Eisenschlos
ACL (1)6
2023 Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia Entities
abstract
Large-scale multi-modal pre-training models such as CLIP [30] and PaLI [8] exhibit strong generalization on various visual domains and tasks. However, existing image classification benchmarks often evaluate recognition on a specific domain (e.g., outdoor images) or a specific task (e.g., classifying plant species), which falls short of evaluating whether pre-trained foundational models are universal visual recognizers. To address this, we formally present the task of Open-domain Visual Entity recognitioN (Oven), where a model need to link an image onto a Wikipedia entity with respect to a text query. We construct Oven-Wiki‡by repurposing 14 existing datasets with all labels grounded onto one single label space: Wikipedia entities. Oven-Wiki challenges models to select among six million possible Wikipedia entities, making it a general visual recognition benchmark with the largest number of labels. Our study on state-ofthe-art pre-trained models reveals large headroom in generalizing to the massive-scale label space. We show that a PaLI-based auto-regressive visual recognition model performs surprisingly well, even on Wikipedia entities that have never been seen during fine-tuning. We also find existing pretrained models yield different strengths: while PaLI-based models obtain higher overall performance, CLIP-based models are better at recognizing tail entities.
Hexiang Hu, Yi Luan, Yang Chen 0065, Urvashi Khandelwal, Mandar Joshi, Kenton Lee, Kristina Toutanova, Ming-Wei Chang
ICCV5
2023 Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
abstract
Visually-situated language is ubiquitous---sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms. Perhaps due to this diversity, previous work has typically relied on domain-specific recipes with limited sharing of the underlying data, model architectures, and objectives. We present Pix2Struct, a pretrained image-to-text model for purely visual language understanding, which can be finetuned on tasks containing visually-situated language. Pix2Struct is pretrained by learning to parse masked screenshots of web pages into simplified HTML. The web, with its richness of visual elements cleanly reflected in the HTML structure, provides a large source of pretraining data well suited to the diversity of downstream tasks. Intuitively, this objective subsumes common pretraining signals such as OCR, language modeling, and image captioning. In addition to the novel pretraining strategy, we introduce a variable-resolution input representation and a more flexible integration of language and vision inputs, where language prompts such as questions are rendered directly on top of the input image. For the first time, we show that a single pretrained model can achieve state-of-the-art results in six out of nine tasks across four domains: documents, illustrations, user interfaces, and natural images.
Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu 0001, Julian Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw 0004, Ming-Wei Chang, Kristina Toutanova
ICML2
2023 From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces
abstract
Much of the previous work towards digital agents for graphical user interfaces (GUIs) has relied on text-based representations (derived from HTML or other structured data sources), which are not always readily available. These input representations have been often coupled with custom, task-specific action spaces. This paper focuses on creating agents that interact with the digital world using the same conceptual interface that humans commonly use — via pixel-based screenshots and a generic action space corresponding to keyboard and mouse actions. Building upon recent progress in pixel-based pretraining, we show, for the first time, that it is possible for such agents to outperform human crowdworkers on the MiniWob++ benchmark of GUI-based instruction following tasks.
Peter Shaw 0004, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, Kristina Toutanova
NeurIPS2
2022 Improving Passage Retrieval with Zero-Shot Question Generation
abstract
We propose a simple and effective re-ranking method for improving passage retrieval in open question answering.The re-ranker re-scores retrieved passages with a zero-shot question generation model, which uses a pre-trained language model to compute the probability of the input question conditioned on a retrieved passage.This approach can be applied on top of any retrieval method (e.g.neural or keywordbased), does not require any domain-or taskspecific training (and therefore is expected to generalize better to data distribution shifts), and provides rich cross-attention between query and passage (i.e. it must explain every token in the question).When evaluated on a number of open-domain retrieval datasets, our re-ranker improves strong unsupervised retrieval models by 6%-18% absolute and strong supervised models by up to 12% in terms of top-20 passage retrieval accuracy.We also obtain new stateof-the-art results on full open-domain question answering by simply adding the new re-ranker to existing models with no further changes.1
Devendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan, Scott Yih, Joelle Pineau, Luke Zettlemoyer
EMNLP3
2022 HTLM: Hyper-Text Pre-Training and Prompting of Language Models
Armen Aghajanyan, Dmytro Okhonko, Mike Lewis, Mandar Joshi, Hu Xu 0001, Gargi Ghosh, Luke Zettlemoyer
ICLR4
2021 DESCGEN: A Distantly Supervised Datasetfor Generating Entity Descriptions
abstract
Weijia Shi, Mandar Joshi, Luke Zettlemoyer. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Mandar Joshi, Luke Zettlemoyer
ACL/IJCNLP (1)2
2021 FEWS: Large-Scale, Low-Shot Word Sense Disambiguation with the Dictionary
abstract
Current models for Word Sense Disambiguation (WSD) struggle to disambiguate rare senses, despite reaching human performance on global WSD metrics.This stems from a lack of data for both modeling and evaluating rare senses in existing WSD datasets.In this paper, we introduce FEWS (Few-shot Examples of Word Senses), a new low-shot WSD dataset automatically extracted from example sentences in Wiktionary.FEWS has high sense coverage across different natural language domains and provides: (1) a large training set that covers many more senses than previous datasets and (2) a comprehensive evaluation set containing few-and zero-shot examples of a wide variety of senses.We establish baselines on FEWS with knowledgebased and neural WSD approaches and present transfer learning experiments demonstrating that models additionally trained with FEWS better capture rare senses in existing WSD datasets.Finally, we find humans outperform the best baseline models on FEWS, indicating that FEWS will support significant future work on low-shot WSD.
Terra Blevins, Mandar Joshi, Luke Zettlemoyer
EACL2
2020 An Information Bottleneck Approach for Controlling Conciseness in Rationale Extraction
abstract
Decisions of complex models for language understanding can be explained by limiting the inputs they are provided to a relevant subsequence of the original text -a rationale.Models that condition predictions on a concise rationale, while being more interpretable, tend to be less accurate than models that are able to use the entire context.In this paper, we show that it is possible to better manage the trade-off between concise explanations and high task accuracy by optimizing a bound on the Information Bottleneck (IB) objective.Our approach jointly learns an explainer that predicts sparse binary masks over input sentences without explicit supervision, and an end-task predictor that considers only the residual sentences.Using IB, we derive a learning objective that allows direct control of mask sparsity levels through a tunable sparse prior.Experiments on the ERASER benchmark demonstrate significant gains over previous work for both task performance and agreement with human rationales.Furthermore, we find that in the semi-supervised setting, a modest amount of gold rationales (25% of training examples with gold masks) can close the performance gap with a model that uses the full input.1
Bhargavi Paranjape, Mandar Joshi, John Thickstun, Hannaneh Hajishirzi, Luke Zettlemoyer
EMNLP (1)2
2020 SpanBERT: Improving Pre-training by Representing and Predicting Spans
abstract
We present SpanBERT, a pre-training method that is designed to better represent and predict spans of text. Our approach extends BERT by (1) masking contiguous random spans, rather than random tokens, and (2) training the span boundary representations to predict the entire content of the masked span, without relying on the individual token representations within it. SpanBERT consistently outperforms BERT and our better-tuned baselines, with substantial gains on span selection tasks such as question answering and coreference resolution. In particular, with the same training data and model size as BERT large , our single model obtains 94.6% and 88.7% F1 on SQuAD 1.1 and 2.0 respectively. We also achieve a new state of the art on the OntoNotes coreference resolution task (79.6% F1), strong performance on the TACRED relation extraction benchmark, and even gains on GLUE. 1
Mandar Joshi, Danqi Chen 0001, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, Omer Levy
Trans. Assoc. Comput. Linguistics1
2019 BERT for Coreference Resolution: Baselines and Analysis
abstract
Mandar Joshi, Omer Levy, Luke Zettlemoyer, Daniel Weld. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mandar Joshi, Omer Levy, Luke Zettlemoyer, Daniel S. Weld
EMNLP/IJCNLP (1)1
2017 TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
abstract
We present TriviaQA, a challenging reading comprehension dataset containing over 650K question-answer-evidence triples.TriviaQA includes 95K questionanswer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high quality distant supervision for answering the questions.We show that, in comparison to other recently introduced large-scale datasets, TriviaQA (1) has relatively complex, compositional questions, (2) has considerable syntactic and lexical variability between questions and corresponding answer-evidence sentences, and (3) requires more cross sentence reasoning to find answers.We also present two baseline algorithms: a featurebased classifier and a state-of-the-art neural network, that performs well on SQuAD reading comprehension.Neither approach comes close to human performance (23% and 40% vs. 80%), suggesting that Trivi-aQA is a challenging testbed that is worth significant future study. 1
Mandar Joshi, Eunsol Choi, Daniel S. Weld, Luke Zettlemoyer
ACL (1)1
2014 Knowledge Graph and Corpus Driven Segmentation and Answer Inference for Telegraphic Entity-seeking Queries
abstract
Much recent work focuses on formal in-terpretation of natural question utterances, with the goal of executing the resulting structured queries on knowledge graphs (KGs) such as Freebase. Here we address two limitations of this approach when ap-plied to open-domain, entity-oriented Web queries. First, Web queries are rarely well-formed questions. They are “telegraphic”, with missing verbs, prepositions, clauses, case and phrase clues. Second, the KG is always incomplete, unable to directly an-swer many queries. We propose a novel technique to segment a telegraphic query and assign a coarse-grained purpose to each segment: a base entity e1, a rela-tion type r, a target entity type t2, and contextual words s. The query seeks en-tity e2 ∈ t2 where r(e1, e2) holds, fur-ther evidenced by schema-agnostic words s. Query segmentation is integrated with the KG and an unstructured corpus where mentions of entities have been linked to the KG. We do not trust the best or any specific query segmentation. Instead, evi-dence in favor of candidate e2s are aggre-gated across several segmentations. Ex-tensive experiments on the ClueWeb cor-pus and parts of Freebase as our KG, us-ing over a thousand telegraphic queries adapted from TREC, INEX, and Web-Questions, show the efficacy of our ap-proach. For one benchmark, MAP im-proves from 0.2–0.29 (competitive base-lines) to 0.42 (our system).
Mandar Joshi, Uma Sawant, Soumen Chakrabarti
EMNLP1
2012 Object-Oriented Representation and Hierarchical Reinforcement Learning in Infinite Mario
abstract
In this work, we analyze and improve upon reinforcement learning techniques used to build agents that can learn to play Infinite Mario, an action game. We extend the object-oriented representation by introducing the concept of object classes which can be effectively used to constrain state spaces. We then use this representation combined with the hierarchical reinforcement learning model as a learning framework. We also extend the idea of hierarchical RL by designing a hierarchy in action selection using domain specific knowledge. With the help of experimental results, we show that this approach facilitates faster and efficient learning for this domain.
Mandar Joshi, Rakesh Khobragade, Saurabh Sarda, Umesh A. Deshpande 0001, Shiwali Mohan
ICTAI1