VLDB 2026 Research / reviewers in the wild / expert
Saku Sugawara
dblp:195/8158
· DBLP profile ↗
28ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0002-0061-0680ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 6 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language ModelsabstractLanguage models (LMs) behave more like humans when their cognitive resources are restricted, particularly in predicting sentence processing costs such as reading times.However, it remains unclear whether such constraints similarly affect sentence comprehension strategies.Besides, existing methods do not directly target the balance between memory storage and sentence processing, which is central to human working memory.To address this issue, we propose a dual-task paradigm that combines an arithmetic computation task with a sentence comprehension task, such as "The 2 cocktail + blended 3 =..." Our experiments show that under dual-task conditions, GPT-4o, o3-mini, and o4-mini shift toward plausibility-based comprehension, mirroring humans' rational inference.Specifically, these models show a greater accuracy gap between plausible sentences (e.g., "The cocktail was blended by the bartender") and implausible sentences (e.g., "The bartender was blended by the cocktail") in the dual-task condition compared to the single-task conditions.These findings suggest that constraints on the balance between memory and processing resources promote rational inference in LMs.More broadly, they support the view that human-like sentence comprehension fundamentally arises from the allocation of limited cognitive resources. Rei Emura, Saku Sugawara |
ACL (1) | 2 |
| 2026 | C2: Scalable Rubric-Augmented Reward Modeling from Binary PreferencesabstractRubric-augmented verification guides reward models with explicit evaluation criteria, yielding more reliable judgments than single-model verification.However, most existing methods require costly rubric annotations, limiting scalability.Moreover, we find that rubric generation is vulnerable to a failure of cooperation; lowquality rubrics actively mislead reward models rather than help.Inspired by the principle of cooperative communication, we propose Cooperative yet Critical reward modeling (C2), a framework that significantly improves reward model judgments by having the reward model critically collaborate with a rubric generator trained solely from binary preferences.In C2, we synthesize helpful and misleading rubric pairs by measuring how each rubric shifts the reward model toward or away from the correct preference.Using these contrastive pairs, we train a cooperative rubric generator to propose helpful rubrics, and a critical verifier to assess rubric validity before making its judgment, following only rubrics it deems helpful at inference time.C2 outperforms reasoning reward models trained on the same binary preferences, with gains of up to 6.5 points on RM-Bench and 6.0 points length-controlled win rate on Al-pacaEval 2.0.Without external rubric annotations, C2 enables an 8B reward model to match performance achieved with rubrics from a 4× larger model.Overall, our work demonstrates that eliciting deliberate cooperation in rubricaugmented verification makes reward models more trustworthy in a scalable way. 1 Akira Kawabata, Saku Sugawara |
ACL (1) | 2 |
| 2026 | CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language ModelsabstractUnderstanding language acquisition in language models remains an open question, yet many benchmarks focus on grammatical acceptability, with far less attention to interpreting meanings conveyed by grammatical forms.We introduce the Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models (CxMP), grounded in Construction Grammar, which treats form-meaning pairings (constructions) as fundamental linguistic units.It evaluates whether models interpret the semantic information implied by constructions, using a controlled minimal-pairs across nine types.Our results show that constructional understanding develops more gradually and remains limited for some constructions even in large language models (LLMs), whereas performance on grammatical acceptability emerges earlier, with shallow heuristics in CxMP exhibiting a U-shaped pattern.These findings highlight the need to broaden existing linguistic evaluations to capture meanings encoded in linguistic form. 1 Miyu Oba, Saku Sugawara |
ACL (1) | 2 |
| 2026 | Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-JudgeabstractLarge language models (LLMs) are increasingly used as automated evaluators (LLM-as-a-Judge). This work challenges its reliability by showing that trust judgments by LLMs are biased by disclosed source labels. Using a counterfactual design, we find that both humans and LLM judges assign higher trust to information labeled as human-authored than to the same content labeled as AI-generated. Eye-tracking data reveal that humans rely heavily on source labels as heuristic cues for judgments. We analyze LLM internal states during judgment. Across label conditions, models allocate denser attention to the label region than the content region, and this label dominance is stronger under Human labels than AI labels, consistent with the human gaze patterns. Besides, decision uncertainty measured by logits is higher under AI labels than Human labels. These results indicate that the source label is a salient heuristic cue for both humans and LLMs. It raises validity concerns for label-sensitive LLM-as-a-Judge evaluation, and we cautiously raise that aligning models with human preferences may propagate human heuristic reliance into models, motivating debiased evaluation and alignment. Xin Sun 0016, Sijing Qin, Isao Echizen, Abdallah El Ali, Saku Sugawara |
ACL (1) | 6 |
| 2026 | Eyes Can't Always Tell: Fusing Eye Tracking and User Priors for User Modeling under AI Advice ConditionsabstractModeling users' cognitive states (e.g., cognitive load and decision confidence) is essential for building adaptive AI in high-stakes decision-making. While eye tracking provides non-invasive behavioral signals correlated with cognitive effort, prior work has not systematically examined how AI assistance contexts, specifically varying advice reliability and user heterogeneity, can alter the mapping between gaze signals and cognitive states. We conducted a within-subject lab eye-tracking study (N=54) on factual verification tasks under three conditions: No-AI, Correct-AI advice, and Incorrect-AI advice. We analyze condition-dependent changes in self-reports and eye-tracking patterns and evaluate the robustness of eye-tracking-based user modeling. Results show that AI advice increases decision confidence compared to No-AI, while Correct-AI is associated with lower perceived cognitive load and more efficient gaze behavior. Crucially, predictive modeling is context-sensitive: the relationship between eye-tracking signals and cognitive states shifts across AI conditions. Finally, fusing eye-tracking features with user priors (demographics, AI literacy/experience, and propensity to trust technology) improves cross-participant generalization. These findings support condition-aware and personalized user modeling for cognitively aligned adaptive AI systems. Xin Sun 0016, Shu Wei, Jos A. Bosch, Isao Echizen, Abdallah El Ali, Saku Sugawara |
UMAP | 8 |
| 2025 | Development of Numerical Error Detection Tasks to Analyze the Numerical Capabilities of Language ModelsabstractNumbers are used to describe quantities in various scenarios in daily life; therefore, numerical errors can significantly affect the meaning of the entire sentence, and even a single-letter error can be fatal. Detecting numerical errors often requires a high level of commonsense and is difficult even with the recent large language models (LLMs). In this study, we create a benchmark dataset of numerical error detection that uses automatically generated numerical errors. In our analysis, we classify the numerical errors based on the properties of the errors and investigate the ability of the model from several perspectives, including the error class, error size, and passage domain. The experimental results indicate that GPT-3.5, GPT-4, and Llama-3-Instruct (8B) perform well in the numerical error detection task; however, they are not as accurate as humans. We find that the LLMs misidentified correct numbers as errors more frequently than the humans did. In particular, the analysis demonstrates that the current LLMs still need improvement for detecting numerical errors requiring calculations or extensive prior knowledge. Taku Sakamoto, Saku Sugawara, Akiko Aizawa |
COLING | 2 |
| 2025 | Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?abstractAutomatic evaluation of generative tasks using large language models faces challenges due to ambiguous criteria.Although automatic checklist generation is a potentially promising approach, its usefulness remains underexplored.We investigate whether checklists should be used for all questions or selectively, generate them using six methods, evaluate their effectiveness across eight model sizes, and identify checklist items that correlate with human evaluations.Through experiments on pairwise comparison and direct scoring tasks, we find that selective checklist use tends to improve evaluation performance in pairwise settings, while its benefits are less consistent in direct scoring.Our analysis also shows that even checklist items with low correlation to human scores often reflect human-written criteria, indicating potential inconsistencies in human evaluation.These findings highlight the need to more clearly define objective evaluation criteria to guide both human and automatic evaluations. 1 Momoka Furuhashi, Kouta Nakayama, Takashi Kodama, Saku Sugawara |
EMNLP | 4 |
| 2025 | TactfulToM: Do LLMs have the Theory of Mind ability to understand White Lies?abstractWhile recent studies explore Large Language Models' (LLMs) performance on Theory of Mind (ToM) reasoning tasks, research on ToM abilities that require more nuanced social context is limited, such as white lies.We introduce TactfulToM, a novel English benchmark designed to evaluate LLMs' ability to understand white lies within real-life conversations and reason about prosocial motivations behind them, particularly when they are used used to spare others' feelings and maintain social harmony.Our benchmark is generated through a multi-stage human-in-the-loop pipeline where LLMs expand manually designed seed stories into conversations to maintain the information asymmetry between participants necessary for authentic white lies.We show that Tactful-ToM is challenging for state-of-the-art models, which perform substantially below humans, revealing shortcomings in their ability to fully comprehend the ToM reasoning that enables true understanding of white lies. 1 Emma Pretty, Saku Sugawara |
EMNLP | 4 |
| 2025 | Measuring Human Involvement in AI-Generated Text: A Case Study on Academic WritingabstractContent creation has dramatically progressed with the rapid advancement of large language models like ChatGPT and Claude. While this progress has greatly enhanced various aspects of life and work, it has also negatively affected certain areas of society. A recent survey revealed that nearly 30% of college students use generative AI to help write academic papers and reports. Most countermeasures treat the detection of AI-generated text as a binary classification task and thus lack robustness. This approach overlooks human involvement in the generation of content even though human-machine collaboration is becoming mainstream. Besides generating entire texts, people may use machines to complete or revise texts. Such human involvement varies case by case, which makes binary classification a less than satisfactory approach. We refer to this situation as participation detection obfuscation. We propose using BERTScore as a metric to measure human involvement in the generation process and a multi-task RoBERTa-based regressor trained on a token classification task to address this problem. To evaluate the effectiveness of this approach, we simulated academic-based scenarios and created a continuous dataset reflecting various levels of human involvement. All of the existing detectors we examined failed to detect the level of human involvement on this dataset. Our method, however, succeeded (F1 score of 0.9423 and a regressor mean squared error of 0.004). Moreover, it demonstrated some generalizability across generative models. Our code is available at https://github.com/gyc-nii/CAS-CS-and-dual-head-detector Zhicheng Dou, Huy H. Nguyen, Ching-Chun Chang, Saku Sugawara, Isao Echizen |
IJCNN | 5 |
| 2025 | Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
Futa Waseda, Saku Sugawara, Isao Echizen |
ACM Multimedia | 2 |
| 2024 | Rationale-Aware Answer Verification by Pairwise Self-EvaluationabstractAnswer verification identifies correct solutions among candidates generated by large language models (LLMs).Current approaches typically train verifier models by labeling solutions as correct or incorrect based solely on whether the final answer matches the gold answer.However, this approach neglects any flawed rationale in the solution yielding the correct answer, undermining the verifier's ability to distinguish between sound and flawed rationales.We empirically show that in StrategyQA, only 19% of LLM-generated solutions with correct answers have valid rationales.Furthermore, we demonstrate that training a verifier on valid rationales significantly improves its ability to distinguish valid and flawed rationales.To make a better verifier without extra human supervision, we introduce REPS (Rationale Enhancement through Pairwise Selection), a method for selecting valid rationales from candidates by iteratively applying pairwise self-evaluation using the same LLM that generates the solutions.Verifiers trained on solutions selected by REPS outperform those trained using conventional training methods on three reasoning benchmarks (ARC-Challenge, DROP, and StrategyQA).Our results suggest that training reliable verifiers requires ensuring the validity of rationales in addition to the correctness of the final answers, which would be critical for models assisting humans in solving complex reasoning tasks. Akira Kawabata, Saku Sugawara |
EMNLP | 2 |
| 2024 | Can Language Models Induce Grammatical Knowledge from Indirect Evidence?abstractWhat kinds of and how much data is necessary for language models to induce grammatical knowledge to judge sentence acceptability?Recent language models still have much room for improvement in their data efficiency compared to humans.This paper investigates whether language models efficiently use indirect data (indirect evidence), from which they infer sentence acceptability.In contrast, humans use indirect evidence efficiently, which is considered one of the inductive biases contributing to efficient language acquisition.To explore this question, we introduce the Wug In-Direct Evidence Test (WIDET), a dataset consisting of training instances inserted into the pre-training data and evaluation instances.We inject synthetic instances with newly coined wug words into pretraining data and explore the model's behavior on evaluation data that assesses grammatical acceptability regarding those words.We prepare the injected instances by varying their levels of indirectness and quantity.Our experiments surprisingly show that language models do not induce grammatical knowledge even after repeated exposure to instances with the same structure but differing only in lexical items from evaluation instances in certain language phenomena.Our findings suggest a potential direction for future research: developing models that use latent indirect evidence to induce grammatical knowledge. Miyu Oba, Yohei Oseki, Akiyo Fukatsu, Akari Haga, Hiroki Ouchi, Taro Watanabe, Saku Sugawara |
EMNLP | 7 |
| 2023 | Which Shortcut Solution Do Question Answering Models Prefer to Learn?abstractQuestion answering (QA) models for reading comprehension tend to exploit spurious correlations in training sets and thus learn shortcut solutions rather than the solutions intended by QA datasets. QA models that have learned shortcut solutions can achieve human-level performance in shortcut examples where shortcuts are valid, but these same behaviors degrade generalization potential on anti-shortcut examples where shortcuts are invalid. Various methods have been proposed to mitigate this problem, but they do not fully take the characteristics of shortcuts themselves into account. We assume that the learnability of shortcuts, i.e., how easy it is to learn a shortcut, is useful to mitigate the problem. Thus, we first examine the learnability of the representative shortcuts on extractive and multiple-choice QA datasets. Behavioral tests using biased training sets reveal that shortcuts that exploit answer positions and word-label correlations are preferentially learned for extractive and multiple-choice QA, respectively. We find that the more learnable a shortcut is, the flatter and deeper the loss landscape is around the shortcut solution in the parameter space. We also find that the availability of the preferred shortcuts tends to make the task easier to perform from an information-theoretic viewpoint. Lastly, we experimentally show that the learnability of shortcuts can be utilized to construct an effective QA training set; the more learnable a shortcut is, the smaller the proportion of anti-shortcut examples required to achieve comparable performance on shortcut and anti-shortcut examples. We claim that the learnability of shortcuts should be considered when designing mitigation methods. Kazutoshi Shinoda, Saku Sugawara, Akiko Aizawa |
AAAI | 2 |
| 2023 | PROPRES: Investigating the Projectivity of Presupposition with Various Triggers and EnvironmentsabstractWhat makes a presupposition of an utteranceinformation taken for granted by its speakerdifferent from other pragmatic inferences such as an entailment is projectivity (e.g., the negative sentence the boy did not stop shedding tears presupposes the boy had shed tears before).The projectivity may vary depending on the combination of presupposition triggers and environments.However, prior natural language understanding studies fail to take it into account as they either use no human baseline or include only negation as an entailment-canceling environment to evaluate models' performance.The current study attempts to reconcile these issues.We introduce a new dataset, projectivity of presupposition (PROPRES), which includes 12k premise-hypothesis pairs crossing six triggers involving some lexical variety with five environments.Our human evaluation reveals that humans exhibit variable projectivity in some cases.However, the model evaluation shows that the best-performed model, DeBERTa, does not fully capture it.Our findings suggest that probing studies on pragmatic inferences should take extra care of the human judgment variability and the combination of linguistic items. Daiki Asami, Saku Sugawara |
CoNLL | 2 |
| 2023 | Evaluating the Rationale Understanding of Critical Reasoning in Logical Reading ComprehensionabstractTo precisely evaluate a language model's capability for logical reading comprehension, we present a dataset for testing the understanding of the rationale behind critical reasoning.For questions taken from an existing multiplechoice logical reading comprehension dataset, we crowdsource rationale texts that explain why we should select or eliminate answer options, resulting in 3,003 multiple-choice subquestions that are associated with 943 main questions.Experiments on our dataset show that recent large language models (e.g., InstructGPT) struggle to answer the subquestions even if they are able to answer the main questions correctly.We find that the models perform particularly poorly in answering subquestions written for the incorrect options of the main questions, implying that the models have a limited capability for explaining why incorrect alternatives should be eliminated.These results suggest that our dataset encourages further investigation into the critical reasoning ability of language models while focusing on the elimination process of relevant alternatives.In a given context, you'll be given a question, an answer, and four rationales.Your task is to identify the rationale that explains the correctness of the provided option the best.If the option is wrong, choose the rationale that explains why it is wrong.Conversely, if the option is correct, choose the rationale that explains why it is correct.Context: Teachers should not do anything to cause their students to lose respect for them.And students can sense when someone is trying to hide his or her ignorance.Therefore, a teacher who does not know the answer to a question a student has asked should not pretend to know the answer.Question: The conclusion is properly drawn if which one of the following is assumed?Question: The conclusion is properly drawn if which one of the following is assumed?Option: Students' respect for a teacher is independent of the amount of knowledge they attribute to that teacher.Rationale0: The ranking of students' respect for honesty is not relevant to the conclusion of a teacher shouldn't pretend to know an answer to question they don't know the answer to. Rationale1:The assumption is that students' respect for the teacher is based on how much knowledge the teacher has.Rationale2: The conclusion is that teachers shouldn't pretend to know the answer to a question that they don't know, so the assumption is that student's respect for a teacher is interlinked to the student's perceived knowledge of the teacher.Rationale3: The conclusion does not have anything to do with a teacher being effective.Answer: The answer is Rationale2Context: Miguel has four family members who plan to come to his graduation on Sunday afternoon, but it is likely that only three of them will be allowed to attend.Normally graduation is held in the football stadium, where there is no limit on the number of family members who can attend.However, the ceremony is relocated to the gymnasium if it rains, and each graduate receives just three admission tickets for use by family members.Question: The conclusion of the argument is most strongly supported if which one of the following is assumed?Option: The weather service has indicated that there is a very high likelihood of rain on Sunday afternoon.Rationale0: No mention is made of whether un-needed spaces can be transferred between students, and so this cannot be assumed to impact the number of spaces available to Miguel's family.Rationale1: Abnormally large class size may not preclude Miguel from having more than three family members attend, as the football stadium is a possible venue and has no limitation on the number who may attend.Rationale2: A family member who cannot attend the graduation has no relevance to how many may be allowed to attend.Rationale3: Rain would preclude the use of the stadium which has no limit of the number of family members attending and force the use of the gymnasium, which limits the number attending to three. Akira Kawabata, Saku Sugawara |
EMNLP | 2 |
| 2023 | Improving Translation of Case Descriptions into Logical Fact Formulas using LegalCaseNERabstractThe automated translation of natural language text into structured logical representations is a critical task in various applications, including legal reasoning and decision-making. This paper presents a Name Entity Recognition (NER) based approach for translating the legal case descriptions written in natural language into PROLEG fact formulas. The approach comprises (1) extracting legal entities from the case description using a specialized NER model, namely LegalCaseNER and (2) transforming the extracted entities into PROLEG fact formulas using PROLEG rules. The experimental results demonstrate the efficacy of our proposed approach in accurately extracting relevant entities from legal case descriptions and translating them into the appropriate PROLEG fact formulas. Our approach provides a promising solution for handling complex and diverse case descriptions, enabling their representation in a structured format. This work provides a foundation for future research in the application of logical fact formulas in legal reasoning and decision-making. May Myo Zin, Ha-Thanh Nguyen, Ken Satoh, Saku Sugawara, Fumihito Nishino |
ICAIL | 4 |
| 2023 | Information Extraction from Lengthy Legal Contracts: Leveraging Query-Based Summarization and GPT-3.5abstractIn the legal domain, extracting information from contracts poses significant challenges, primarily due to the scarcity of annotated data. In such situations, leveraging large language models (LLMs), such as the Generative Pretrained Transformer (GPT) models, offers a promising solution. However, the inherent token limitations of these models can be a bottleneck for processing lengthy legal contracts. This paper presents an unsupervised two-step approach to address these challenges. First, we propose a query-based summarization model that extracts sentences pertinent to predefined queries, concisely representing lengthy contracts. This summarization ensures that the core information remains intact while simultaneously addressing the token limitation issue. Subsequently, the generated summary is fed to GPT-3.5 for precise information extraction. Our approach effectively overcomes the challenges of token limitations and zero resources, enabling efficient and scalable information extraction from legal contracts. We compare our results with those obtained from supervised models that have been fine-tuned on domain-specific annotated data. Experimental results demonstrate the remarkable effectiveness of our approach, as it achieves state-of-the-art performance without the need for domain-specific training data. May Myo Zin, Ha-Thanh Nguyen, Ken Satoh, Saku Sugawara, Fumihito Nishino |
JURIX | 4 |
| 2022 | What Makes Reading Comprehension Questions Difficult?abstractFor a natural language understanding benchmark to be useful in research, it has to consist of examples that are diverse and difficult enough to discriminate among current and near-future state-of-the-art systems.However, we do not yet know how best to select text sources to collect a variety of challenging examples.In this study, we crowdsource multiple-choice reading comprehension questions for passages taken from seven qualitatively distinct sources, analyzing what attributes of passages contribute to the difficulty and question types of the collected examples.To our surprise, we find that passage source, length, and readability measures do not significantly affect question difficulty.Through our manual annotation of seven reasoning types, we observe several trends between passage sources and reasoning types, e.g., logical reasoning is more often required in questions written for technical passages.These results suggest that when creating a new benchmark dataset, selecting a diverse set of passages can help ensure a diverse range of question types, but that passage difficulty need not be a priority. Saku Sugawara, Nikita Nangia, Alex Warstadt, Samuel R. Bowman |
ACL (1) | 1 |
| 2022 | Possible Stories: Evaluating Situated Commonsense Reasoning under Multiple Possible ScenariosabstractThe possible consequences for the same context may vary depending on the situation we refer to. However, current studies in natural language processing do not focus on situated commonsense reasoning under multiple possible scenarios. This study frames this task by asking multiple questions with the same set of possible endings as candidate answers, given a short story text. Our resulting dataset, Possible Stories, consists of more than 4.5K questions over 1.3K story texts in English. We discover that even current strong pretrained language models struggle to answer the questions consistently, highlighting that the highest accuracy in an unsupervised setting (60.2%) is far behind human accuracy (92.5%). Through a comparison with existing datasets, we observe that the questions in our dataset contain minimal annotation artifacts in the answer options. In addition, our dataset includes examples that require counterfactual reasoning, as well as those requiring readers’ reactions and fictional information, suggesting that our dataset can serve as a challenging testbed for future studies on situated commonsense reasoning. Mana Ashida, Saku Sugawara |
COLING | 2 |
| 2022 | Debiasing Masks: A New Framework for Shortcut Mitigation in NLUabstractDebiasing language models from unwanted behaviors in Natural Language Understanding tasks is a topic with rapidly increasing interest in the NLP community.Spurious statistical correlations in the data allow models to perform shortcuts and avoid uncovering more advanced and desirable linguistic features.A multitude of effective debiasing approaches has been proposed, but flexibility remains a major issue.For the most part, models must be retrained to find a new set of weights with debiased behavior.We propose a new debiasing method in which we identify debiased pruning masks that can be applied to a finetuned model.This enables the selective and conditional application of debiasing behaviors.We assume that bias is caused by a certain subset of weights in the network; our method is, in essence, a mask search to identify and remove biased weights.Our masks show equivalent or superior performance to the standard counterparts, while offering important benefits.Pruning masks can be stored with high efficiency in memory, and it becomes possible to switch among several debiasing behaviors (or revert back to the original biased model) at inference time.Finally, it opens the doors to further research on how biases are acquired by studying the generated masks.For example, we observed that the early layers and attention heads were pruned more aggressively, possibly hinting towards the location in which biases may be encoded. Johannes Mario Meissner, Saku Sugawara, Akiko Aizawa |
EMNLP | 2 |
| 2022 | Cross-Modal Similarity-Based Curriculum Learning for Image CaptioningabstractImage captioning models require the high-level generalization ability to describe the contents of various images in words.Most existing approaches treat the image-caption pairs equally in their training without considering the differences in their learning difficulties.Several image captioning approaches introduce curriculum learning methods that present training data with increasing levels of difficulty.However, their difficulty measurements are either based on domain-specific features or prior model training.In this paper, we propose a simple yet efficient difficulty measurement for image captioning using cross-modal similarity calculated by a pretrained vision-language model.Experiments on the COCO and Flickr30k datasets show that our proposed approach achieves superior performance and competitive convergence speed to baselines without requiring heuristics or incurring additional training costs.Moreover, the higher model performance on difficult examples and unseen data also demonstrates the generalization ability. References:-A pizza with multiple toppings including an egg.-A white plate holding a pizza next to plate of fries.-A pizza sits on top of a pan covered in vegetables.-A pizza pie with vegetables of some sort on it.-A full veggie pizza is ready to be eaten. Hongkuan Zhang, Saku Sugawara, Akiko Aizawa, Ryohei Sasano, Koichi Takeda 0003 |
EMNLP | 2 |
| 2021 | What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?abstractNikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, Samuel R. Bowman |
ACL/IJCNLP (1) | 2 |
| 2021 | Benchmarking Machine Reading Comprehension: A Psychological PerspectiveabstractMachine reading comprehension (MRC) has received considerable attention as a benchmark for natural language understanding.However, the conventional task design of MRC lacks explainability beyond the model interpretation, i.e., reading comprehension by a model cannot be explained in human terms.To this end, this position paper provides a theoretical basis for the design of MRC datasets based on psychology as well as psychometrics, and summarizes it in terms of the prerequisites for benchmarking MRC.We conclude that future datasets should (i) evaluate the capability of the model for constructing a coherent and grounded representation to understand contextdependent situations and (ii) ensure substantive validity by shortcut-proof questions and explanation as a part of the task design. Saku Sugawara, Pontus Stenetorp, Akiko Aizawa |
EACL | 1 |
| 2020 | Assessing the Benchmarking Capacity of Machine Reading Comprehension DatasetsabstractExisting analysis work in machine reading comprehension (MRC) is largely concerned with evaluating the capabilities of systems. However, the capabilities of datasets are not assessed for benchmarking language understanding precisely. We propose a semi-automated, ablation-based methodology for this challenge; By checking whether questions can be solved even after removing features associated with a skill requisite for language understanding, we evaluate to what degree the questions do not require the skill. Experiments on 10 datasets (e.g., CoQA, SQuAD v2.0, and RACE) with a strong baseline model show that, for example, the relative scores of the baseline model provided with content words only and with shuffled sentence words in the context are on average 89.2% and 78.5% of the original scores, respectively. These results suggest that most of the questions already answered correctly by the model do not necessarily require grammatical and complex reasoning. For precise benchmarking, MRC datasets will need to take extra care in their design to ensure that questions can correctly evaluate the intended skills. Saku Sugawara, Pontus Stenetorp, Kentaro Inui, Akiko Aizawa |
AAAI | 1 |
| 2020 | Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning StepsabstractA multi-hop question answering (QA) dataset aims to test reasoning and inference skills by requiring a model to read multiple paragraphs to answer a given question.However, current datasets do not provide a complete explanation for the reasoning process from the question to the answer.Further, previous studies revealed that many examples in existing multi-hop datasets do not require multi-hop reasoning to answer a question.In this study, we present a new multihop QA dataset, called 2WikiMultiHopQA, which uses structured and unstructured data.In our dataset, we introduce the evidence information containing a reasoning path for multi-hop questions.The evidence information has two benefits: (i) providing a comprehensive explanation for predictions and (ii) evaluating the reasoning skills of a model.We carefully design a pipeline and a set of templates when generating a question-answer pair that guarantees the multi-hop steps and the quality of the questions.We also exploit the structured format in Wikidata and use logical rules to create questions that are natural but still require multi-hop reasoning.Through experiments, we demonstrate that our dataset is challenging for multi-hop models and it ensures that multi-hop reasoning is required. Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, Akiko Aizawa |
COLING | 3 |
| 2018 | What Makes Reading Comprehension Questions Easier?abstractA challenge in creating a dataset for machine reading comprehension (MRC) is to collect questions that require a sophisticated understanding of language to answer beyond using superficial cues.In this work, we investigate what makes questions easier across recent 12 MRC datasets with three question styles (answer extraction, description, and multiple choice).We propose to employ simple heuristics to split each dataset into easy and hard subsets and examine the performance of two baseline models for each of the subsets.We then manually annotate questions sampled from each subset with both validity and requisite reasoning skills to investigate which skills explain the difference between easy and hard questions.From this study, we observed that (i) the baseline performances for the hard subsets remarkably degrade compared to those of entire datasets, (ii) hard questions require knowledge inference and multiple-sentence reasoning in comparison with easy questions, and (iii) multiplechoice questions tend to require a broader range of reasoning skills than answer extraction and description questions.These results suggest that one might overestimate recent advances in MRC. Saku Sugawara, Kentaro Inui, Satoshi Sekine, Akiko Aizawa |
EMNLP | 1 |
| 2017 | Prerequisite Skills for Reading Comprehension: Multi-Perspective Analysis of MCTest Datasets and SystemsabstractOne of the main goals of natural language processing (NLP) is synthetic understanding of natural language documents, especially reading comprehension (RC). An obstacle to the further development of RC systems is the absence of a synthetic methodology to analyze their performance. It is difficult to examine the performance of systems based solely on their results for tasks because the process of natural language understanding is complex. In order to tackle this problem, we propose in this paper a methodology inspired by unit testing in software engineering that enables the examination of RC systems from multiple aspects. Our methodology consists of three steps. First, we define a set of prerequisite skills for RC based on existing NLP tasks. We assume that RC capability can be divided into these skills. Second, we manually annotate a dataset for an RC task with information regarding the skills needed to answer each question. Finally, we analyze the performance of RC systems for each skill based on the annotation. The last two steps highlight two aspects: the characteristics of the dataset, and the weaknesses in and differences among RC systems. We tested the effectiveness of our methodology by annotating the Machine Comprehension Test (MCTest) dataset and analyzing four existing systems (including a neural system) on it. The results of the annotations showed that answering questions requires a combination of skills, and clarified the kinds of capabilities that systems need to understand natural language. We conclude that the set of prerequisite skills we define are promising for the decomposition and analysis of RC. Saku Sugawara, Hikaru Yokono, Akiko Aizawa |
AAAI | 1 |
| 2017 | Evaluation Metrics for Machine Reading Comprehension: Prerequisite Skills and ReadabilityabstractKnowing the quality of reading comprehension (RC) datasets is important for the development of natural-language understanding systems.In this study, two classes of metrics were adopted for evaluating RC datasets: prerequisite skills and readability.We applied these classes to six existing datasets, including MCTest and SQuAD, and highlighted the characteristics of the datasets according to each metric and the correlation between the two classes.Our dataset analysis suggests that the readability of RC datasets does not directly affect the question difficulty and that it is possible to create an RC dataset that is easy to read but difficult to answer. Saku Sugawara, Yusuke Kido, Hikaru Yokono, Akiko Aizawa |
ACL (1) | 1 |