EDBT 2026 Demo / reviewers in the wild / expert
Akari Asai
dblp:213/8066
· DBLP profile ↗
20ranked-venue papers
10as first author
15since 2021 · last 2025
0009-0001-3244-6598ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 10 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Quantifying the Influence of Evaluation Aspects on Long-Form Response AssessmentabstractEvaluating the outputs of large language models (LLMs) on long-form generative tasks remains challenging. While fine-grained, aspect-wise evaluations provide valuable diagnostic information, they are difficult to design exhaustively, and each aspect’s contribution to the overall acceptability of an answer is unclear. In this study, we propose a method to compute an overall quality score as a weighted average of three key aspects: factuality, informative- ness, and formality. This approach achieves stronger correlations with human judgments compared to previous metrics. Our analysis identifies factuality as the most predictive aspect of overall quality. Additionally, we release a dataset of 1.2k long-form QA answers annotated with both absolute judgments and relative preferences in overall and aspect-wise schemes to aid future research in evaluation practices. Go Kamoda, Akari Asai, Ana Brassard, Keisuke Sakaguchi |
COLING | 2 |
| 2025 | Pangea: A Fully Open Multilingual Multimodal LLM for 39 LanguagesabstractDespite recent advances in multimodal large language models (MLLMs), their development has predominantly focused on English- and western-centric datasets and tasks, leaving most of the world's languages and diverse cultural contexts underrepresented.
This paper introduces PANGEA, a multilingual multimodal LLM trained on PANGEAINS, a diverse 6M instruction dataset spanning 39 languages. PANGEAINS features: 1) high-quality English instructions, 2) carefully machine-translated instructions, and 3) culturally relevant multimodal tasks to ensure cross-cultural coverage.
To rigorously assess models' capabilities, we introduce PANGEABENCH, a holistic evaluation suite encompassing 14 datasets covering 47 languages.
Results show that PANGEA significantly outperforms existing open-source models in multilingual settings and diverse cultural contexts. Ablation studies further reveal the importance of English data proportions, language popularity, and the number of multimodal training samples on overall performance. We fully open-source our data, code, and trained checkpoints, to facilitate the development of inclusive and robust multilingual MLLMs, promoting equity and accessibility across a broader linguistic and cultural spectrum. Xiang Yue, Yueqi Song, Akari Asai, Seungone Kim, Jean de Dieu Nyandwi, Simran Khanuja, Anjali Kantharuban, Lintang Sutawika, Sathyanarayanan Ramamoorthy, Graham Neubig |
ICLR | 3 |
| 2024 | CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model GenerationabstractTong Chen, Akari Asai, Niloofar Mireshghallah, Sewon Min, James Grimmelmann, Yejin Choi, Hannaneh Hajishirzi, Luke Zettlemoyer, Pang Wei Koh. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Tong Chen 0005, Akari Asai, Niloofar Mireshghallah, Sewon Min, James Grimmelmann, Yejin Choi 0001, Hannaneh Hajishirzi, Luke Zettlemoyer, Pang Wei Koh |
EMNLP | 2 |
| 2024 | Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionabstractDespite their remarkable capabilities, large language models (LLMs) often produce responses containing factual inaccuracies due to their sole reliance on the parametric knowledge they encapsulate. Retrieval-Augmented Generation (RAG), an ad hoc approach that augments LMs with retrieval of relevant knowledge, decreases such issues. However, indiscriminately retrieving and incorporating a fixed number of retrieved passages, regardless of whether retrieval is necessary, or passages are relevant, diminishes LM versatility or can lead to unhelpful response generation. We introduce a new framework called **Self-Reflective Retrieval-Augmented Generation (Self-RAG)** that enhances an LM's quality and factuality through retrieval and self-reflection.
Our framework trains a single arbitrary LM that adaptively retrieves passages on-demand, and generates and reflects on retrieved passages and its generations using special tokens, called {\it reflection} tokens. Generating reflection tokens makes the LM controllable during the inference phase, enabling it to tailor its behavior to diverse task requirements.
Experiments show that Self-RAG (7B and 13B parameters) significantly outperforms state-of-the-art LLMs and retrieval-augmented models on a diverse set of tasks.
Specifically, Self-RAG outperforms ChatGPT and retrieval-augmented Llama2-chat on Open-domain QA, reasoning, and fact verification tasks, and it shows significant gains in improving factuality and citation accuracy for long-form generations relative to these models. Our code and trained models are available at https://selfrag.github.io/ Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, Hannaneh Hajishirzi |
ICLR | 1 |
| 2024 | BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual TransferabstractAkari Asai, Sneha Kudugunta, Xinyan Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, Hannaneh Hajishirzi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Akari Asai, Sneha Reddy Kudugunta, Xinyan Yu 0001, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, Hannaneh Hajishirzi |
NAACL-HLT | 1 |
| 2024 | Scaling Retrieval-Based Language Models with a Trillion-Token DatastoreabstractScaling laws with respect to the amount of training data and the number of parameters allow us to predict the cost-benefit trade-offs of pretraining language models (LMs) in different configurations. In this paper, we consider another dimension of scaling: the amount of data available at inference time. Specifically, we find that increasing the size of the datastore used by a retrieval-based LM monotonically improves language modeling and several downstream tasks without obvious saturation, such that a smaller model augmented with a large datastore outperforms a larger LM-only model on knowledge-intensive tasks. By plotting compute-optimal scaling curves with varied datastore, model, and pretraining data sizes, we show that using larger datastores can significantly improve model performance for the same training compute budget. We carry out our study by constructing a 1.4 trillion-token datastore named MassiveDS, which is the largest and the most diverse open-sourced datastore for retrieval-based LMs to date, and designing an efficient pipeline for studying datastore scaling in an accessible manner. Finally, we analyze the effect of improving the retriever, datastore quality filtering, and other design choices on our observed scaling trends. Overall, our results show that datastore size should be considered as an integral part of LM efficiency and performance trade-offs. To facilitate future research, we open-source our datastore and code at https://github.com/RulinShao/retrieval-scaling. Rulin Shao, Jacqueline He, Akari Asai, Tim Dettmers, Sewon Min, Luke Zettlemoyer, Pang Wei Koh |
NeurIPS | 3 |
| 2023 | When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesabstractAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, Hannaneh Hajishirzi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, Hannaneh Hajishirzi |
ACL (1) | 2 |
| 2023 | TaskWeb: Selecting Better Source Tasks for Multi-task NLPabstractRecent work in NLP has shown promising results in training models on large amounts of tasks to achieve better generalization.However, it is not well-understood how tasks are related, and how helpful training tasks can be chosen for a new task.In this work, we investigate whether knowing task relationships via pairwise task transfer improves choosing one or more source tasks that help to learn a new target task.We provide TASKWEB, a largescale benchmark of pairwise task transfers for 22 NLP tasks using three different model types, sizes, and adaptation methods, spanning about 25,000 experiments.Then, we design a new method TASKSHOP based on our analysis of TASKWEB.TASKSHOP uses TASKWEB to estimate the benefit of using a source task for learning a new target task, and to choose a subset of helpful training tasks for multi-task training.Our method improves overall rankings and top-k precision of source tasks by 10% and 38%, respectively.We also use TASKSHOP to build much smaller multi-task training sets that improve zero-shot performances across 11 different target tasks by at least 4.3%. Joongwon Kim, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi |
EMNLP | 2 |
| 2023 | RealTime QA: What's the Answer Right Now?abstractWe introduce RealTime QA, a dynamic question answering (QA) platform that announces questions and evaluates systems on a regular basis (weekly in this version). RealTime QA inquires about the current world, and QA systems need to answer questions about novel events or information. It therefore challenges static, conventional assumptions in open-domain QA datasets and pursues instantaneous applications. We build strong baseline models upon large pretrained language models, including GPT-3 and T5. Our benchmark is an ongoing effort, and this paper presents real-time evaluation results over the past year. Our experimental results show that GPT-3 can often properly update its generation results, based on newly-retrieved documents, highlighting the importance of up-to-date information retrieval. Nonetheless, we find that GPT-3 tends to return outdated answers when retrieved documents do not provide sufficient information to find an answer. This suggests an important avenue for future research: can an open-domain QA system identify such unanswerable cases and communicate with the user or even the retrieval module to modify the retrieval results? We hope that RealTime QA will spur progress in instantaneous applications of question answering and beyond. Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Ronan Le Bras 0001, Akari Asai, Xinyan Yu 0001, Dragomir R. Radev, Noah A. Smith, Yejin Choi 0001, Kentaro Inui |
NeurIPS | 5 |
| 2022 | ATTEMPT: Parameter-Efficient Multi-task Tuning via Attentional Mixtures of Soft PromptsabstractThis work introduces a new multi-task, parameter-efficient language model (LM) tuning method that learns to transfer knowledge across different tasks via a mixture of soft prompts-small prefix embedding vectors pretrained for different tasks.Our method, called ATTEMPT (ATTEntional Mixtures of Prompt Tuning), obtains source prompts as encodings of large-scale source tasks into a small number of parameters and trains an attention module to interpolate the source prompts and a newly initialized target prompt for every instance in the target task.During training, only the target task prompt and the attention weights, which are shared between tasks in multi-task training, are updated, while the original LM and source prompts are intact.ATTEMPT is highly parameter-efficient (e.g., updates 2,300 times fewer parameters than full fine-tuning), while achieving high task performance using knowledge from high-resource tasks.Moreover, it is modular using pre-trained soft prompts and can flexibly add or remove source prompts for effective knowledge transfer.Our experimental results across 21 diverse NLP datasets show that ATTEMPT significantly outperforms prompt tuning and outperforms or matches fully finetuned or other parameter-efficient tuning approaches that use over ten times more parameters.Finally, ATTEMPT outperforms previous work in few-shot learning settings. Akari Asai, Mohammadreza Salehi, Matthew E. Peters, Hannaneh Hajishirzi |
EMNLP | 1 |
| 2022 | Evidentiality-guided Generation for Knowledge-Intensive NLP TasksabstractRetrieval-augmented generation models have shown state-of-the-art performance across many knowledge-intensive NLP tasks such as open-domain question answering and fact verification.These models are trained to generate a final output given retrieved passages that can be irrelevant to an input query, leading to learning spurious cues or memorization.This work introduces a method to incorporate evidentiality of passages-whether a passage contains correct evidence to support the outputinto training the generator.We introduce a multi-task learning framework to jointly generate the final output and predict the evidentiality of each passage.Furthermore, we introduce a new task-agnostic method for obtaining high-quality silver evidentiality labels, addressing the issues of gold evidentiality labels being unavailable in most domains.Our experiments on five datasets across three knowledgeintensive tasks show that our new evidentialityguided generator significantly outperforms its direct counterpart on all of them, and advances the state of the art on three of them.Our analysis shows that the multi-task learning and silver evidentiality mining play key roles. Akari Asai, Matt Gardner 0001, Hannaneh Hajishirzi |
NAACL-HLT | 1 |
| 2021 | Challenges in Information-Seeking QA: Unanswerable Questions and Paragraph RetrievalabstractAkari Asai, Eunsol Choi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Akari Asai, Eunsol Choi |
ACL/IJCNLP (1) | 1 |
| 2021 | MultiModalQA: complex question answering over text, tables and images
Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, Jonathan Berant |
ICLR | 6 |
| 2021 | XOR QA: Cross-lingual Open-Retrieval Question AnsweringabstractAkari Asai, Jungo Kasai, Jonathan Clark, Kenton Lee, Eunsol Choi, Hannaneh Hajishirzi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Akari Asai, Jungo Kasai, Jonathan H. Clark, Kenton Lee, Eunsol Choi, Hannaneh Hajishirzi |
NAACL-HLT | 1 |
| 2021 | One Question Answering Model for Many Languages with Cross-lingual Dense Passage RetrievalabstractWe present Cross-lingual Open-Retrieval Answer Generation (CORA), the first unified many-to-many question answering (QA) model that can answer questions across many languages, even for ones without language-specific annotated data or knowledge sources.We introduce a new dense passage retrieval algorithm that is trained to retrieve documents across languages for a question.Combined with a multilingual autoregressive generation model, CORA answers directly in the target language without any translation or in-language retrieval modules as used in prior work. We propose an iterative training method that automatically extends annotated data available only in high-resource languages to low-resource ones. Our results show that CORA substantially outperforms the previous state of the art on multilingual open QA benchmarks across 26 languages, 9 of which are unseen during training. Our analyses show the significance of cross-lingual retrieval and generation in many languages, particularly under low-resource settings. Akari Asai, Xinyan Yu 0001, Jungo Kasai, Hannaneh Hajishirzi |
NeurIPS | 1 |
| 2020 | Logic-Guided Data Augmentation and Regularization for Consistent Question AnsweringabstractMany natural language questions require qualitative, quantitative or logical comparisons between two entities or events.This paper addresses the problem of improving the accuracy and consistency of responses to comparison questions by integrating logic rules and neural models.Our method leverages logical and linguistic knowledge to augment labeled training data and then uses a consistency-based regularizer to train the model.Improving the global consistency of predictions, our approach achieves large improvements over previous methods in a variety of question answering (QA) tasks including multiple-choice qualitative reasoning, cause-effect reasoning, and extractive machine reading comprehension.In particular, our method significantly improves the performance of RoBERTa-based models by 1-5% across datasets.We advance state of the art by around 5-8% on WIQA and QuaRel and reduce consistency violations by 58% on HotpotQA.We further demonstrate that our approach can learn effectively from limited data. 1 Q: The ceramic vase was less flexible than the plastic ball so it was A: more breakable Q: The ceramic vase was more flexible than the plastic ball so it was A: less breakable Q: If it is silent, does the outer ear collect less sound waves?A: more [positive causal relationship] Q: If the outer ear collect less sound waves, is less sound being detected?A: more [positive causal relationship] Q: If it is silent, is less sound being detected?A: more [positive causal relationship] Akari Asai, Hannaneh Hajishirzi |
ACL | 1 |
| 2020 | LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attentionabstractEntity representations are useful in natural language tasks involving entities.In this paper, we propose new pretrained contextualized representations of words and entities based on the bidirectional transformer (Vaswani et al., 2017).The proposed model treats words and entities in a given text as independent tokens, and outputs contextualized representations of them.Our model is trained using a new pretraining task based on the masked language model of BERT (Devlin et al., 2019).The task involves predicting randomly masked words and entities in a large entity-annotated corpus retrieved from Wikipedia.We also propose an entity-aware self-attention mechanism that is an extension of the self-attention mechanism of the transformer, and considers the types of tokens (words or entities) when computing attention scores.The proposed model achieves impressive empirical performance on a wide range of entity-related tasks.In particular, it obtains state-of-the-art results on five well-known datasets: Open Entity (entity typing), TACRED (relation classification), CoNLL-2003 (named entity recognition), ReCoRD (cloze-style question answering), and SQuAD 1.1 (extractive question answering).Our source code and pretrained representations are available at https: //github.com/studio-ousia/luke. Ikuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda 0001, Yuji Matsumoto 0001 |
EMNLP (1) | 2 |
| 2020 | Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question Answering
Akari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher, Caiming Xiong |
ICLR | 1 |
| 2020 | The Aleatoric Uncertainty Estimation Using a Separate Formulation with Virtual ResidualsabstractWe propose a new optimization framework for aleatoric uncertainty estimation in regression problems. Existing methods can quantify the error in the target estimation, but they tend to underestimate it. To obtain the predictive uncertainty inherent in an observation, we propose a new separable formulation for the estimation of a signal and of its uncertainty, avoiding the effect of overfitting. By decoupling target estimation and uncertainty estimation, we also control the balance between signal estimation and uncertainty estimation. We conduct three types of experiments: regression with simulation data, age estimation, and depth estimation. We demonstrate that the proposed method outperforms a state-of-the-art technique for signal and uncertainty estimation. Takumi Kawashima, Qing Yu 0013, Akari Asai, Daiki Ikami, Kiyoharu Aizawa |
ICPR | 3 |
| 2018 | HappyDB: A Corpus of 100, 000 Crowdsourced Happy Moments
Akari Asai, Sara Evensen, Behzad Golshan, Alon Y. Halevy, Vivian Li, Andrei Lopatenko, Daniela Stepanov, Yoshihiko Suhara, Wang Chiew Tan, Yinzhan Xu |
LREC | 1 |