EDBT 2026 Demo / reviewers in the wild / expert
Keisuke Sakaguchi
dblp:127/0185
· DBLP profile ↗
32ranked-venue papers
6as first author
18since 2021 · last 2025
0000-0002-3809-1732ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE datasetabstractDiana Galvan-Sosa, Gabrielle Gaudeau, Pride Kavumba, Yunmeng Li, Hongyi Gu, Zheng Yuan, Keisuke Sakaguchi, Paula Buttery. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Diana Galván-Sosa, Gabrielle Gaudeau, Pride Kavumba, Yunmeng Li, Hongyi Gu, Zheng Yuan 0003, Keisuke Sakaguchi, Paula Buttery |
ACL (1) | 7 |
| 2025 | Annotating Errors in English Learners' Written Language Production: Advancing Automated Written Feedback Systems
Steven Coyne, Diana Galván-Sosa, Ryan Spring, Camélia Guerraoui, Michael Zock, Keisuke Sakaguchi, Kentaro Inui |
AIED (4) | 6 |
| 2025 | Quantifying the Influence of Evaluation Aspects on Long-Form Response AssessmentabstractEvaluating the outputs of large language models (LLMs) on long-form generative tasks remains challenging. While fine-grained, aspect-wise evaluations provide valuable diagnostic information, they are difficult to design exhaustively, and each aspect’s contribution to the overall acceptability of an answer is unclear. In this study, we propose a method to compute an overall quality score as a weighted average of three key aspects: factuality, informative- ness, and formality. This approach achieves stronger correlations with human judgments compared to previous metrics. Our analysis identifies factuality as the most predictive aspect of overall quality. Additionally, we release a dataset of 1.2k long-form QA answers annotated with both absolute judgments and relative preferences in overall and aspect-wise schemes to aid future research in evaluation practices. Go Kamoda, Akari Asai, Ana Brassard, Keisuke Sakaguchi |
COLING | 4 |
| 2025 | Sketch2Diagram: Generating Vector Diagrams from Hand-Drawn SketchesabstractWe address the challenge of automatically generating high-quality vector diagrams from hand-drawn sketches.
Vector diagrams are essential for communicating complex ideas across various fields, offering flexibility and scalability.
While recent research has progressed in generating diagrams from text descriptions, converting hand-drawn sketches into vector diagrams remains largely unexplored due to the lack of suitable datasets.
To address this gap, we introduce SketikZ, a dataset comprising 3,231 pairs of hand-drawn sketches and thier corresponding TikZ codes as well as reference diagrams.
Our evaluations reveal the limitations of state-of-the-art vision and language models (VLMs), positioning SketikZ as a key benchmark for future research in sketch-to-diagram conversion.
Along with SketikZ, we present ImgTikZ, an image-to-TikZ model that integrates a 6.7B parameter code-specialized open-source large language model (LLM) with a pre-trained vision encoder.
Despite its relatively compact size, ImgTikZ performs comparably to GPT-4o.
This success is driven by using our two data augmentation techniques and a multi-candidate inference strategy.
Our findings open promising directions for future research in sketch-to-diagram conversion and broader image-to-code generation tasks. SketikZ is publicly available. Itsumi Saito, Haruto Yoshida, Keisuke Sakaguchi |
ICLR | 3 |
| 2025 | Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale EvaluationabstractJaehyeok Lee, Keisuke Sakaguchi, JinYeong Bak. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jaehyeok Lee, Keisuke Sakaguchi, JinYeong Bak |
NAACL (Long Papers) | 2 |
| 2025 | Language Models can Categorize System Inputs for Performance AnalysisabstractDominic Sobhani, Ruiqi Zhong, Edison Marrese-Taylor, Keisuke Sakaguchi, Yutaka Matsuo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dominic Sobhani, Ruiqi Zhong, Edison Marrese-Taylor, Keisuke Sakaguchi, Yutaka Matsuo |
NAACL (Long Papers) | 4 |
| 2024 | A Call for Clarity in Beam Search: How It Works and When It StopsabstractText generation with beam search has proven successful in a wide range of applications. We point out that, though largely overlooked in the literature, the commonly-used implementation of beam decoding (e.g., Hugging Face Transformers and fairseq) uses a first come, first served heuristic: it keeps a set of already completed sequences over time steps and stops when the size of this set reaches the beam size. Based on this finding, we introduce a patience factor, a simple modification to this beam decoding implementation, that generalizes the stopping criterion and provides flexibility to the depth of search. Empirical results demonstrate that adjusting this patience factor improves decoding performance of strong pretrained models on news text summarization and machine translation over diverse language pairs, with a negligible inference slowdown. Our approach only modifies one line of code and can be thus readily incorporated in any implementation. Further, we find that different versions of beam decoding result in large performance differences in summarization, demonstrating the need for clarity in specifying the beam search implementation in research work. Our code will be available upon publication. Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras 0001, Dragomir R. Radev, Yejin Choi 0001, Noah A. Smith |
LREC/COLING | 2 |
| 2024 | First Heuristic Then Rational: Dynamic Use of Heuristics in Language Model ReasoningabstractYoichi Aoki, Keito Kudo, Tatsuki Kuribayashi, Shusaku Sone, Masaya Taniguchi, Keisuke Sakaguchi, Kentaro Inui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yoichi Aoki, Keito Kudo, Tatsuki Kuribayashi, Shusaku Sone, Masaya Taniguchi, Keisuke Sakaguchi, Kentaro Inui |
EMNLP | 6 |
| 2024 | PlaSma: Procedural Knowledge Models for Language-based Planning and Re-PlanningabstractProcedural planning, which entails decomposing a high-level goal into a sequence of temporally ordered steps, is an important yet intricate task for machines. It involves integrating common-sense knowledge to reason about complex and often contextualized situations, e.g. ``scheduling a doctor's appointment without a phone''. While current approaches show encouraging results using large language models (LLMs), they are hindered by drawbacks such as costly API calls and reproducibility issues. In this paper, we advocate planning using smaller language models. We present PlaSma, a novel two-pronged approach to endow small language models with procedural knowledge and (constrained) language-based planning capabilities. More concretely, we develop *symbolic procedural knowledge distillation* to enhance the commonsense knowledge in small language models and an *inference-time algorithm* to facilitate more structured and accurate reasoning. In addition, we introduce a new related task, *Replanning*, that requires a revision of a plan to cope with a constrained situation. In both the planning and replanning settings, we show that orders-of-magnitude smaller models (770M-11B parameters) can compete and often surpass their larger teacher models' capabilities. Finally, we showcase successful application of PlaSma in an embodied environment, VirtualHome. Faeze Brahman, Chandra Bhagavatula, Valentina Pyatkin, Jena D. Hwang, Xiang Li 0069, Hirona Jacqueline Arai, Soumya Sanyal 0001, Keisuke Sakaguchi, Xiang Ren 0001, Yejin Choi 0001 |
ICLR | 8 |
| 2024 | A Multimodal Dialogue System to Lead Consensus Building with Emotion-DisplayingabstractShinnosuke Nozue, Yuto Nakano, Shoji Moriya, Tomoki Ariyama, Kazuma Kokuta, Suchun Xie, Kai Sato, Shusaku Sone, Ryohei Kamei, Reina Akama, Yuichiroh Matsubayashi, Keisuke Sakaguchi. Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2024. Shinnosuke Nozue, Yuto Nakano, Shoji Moriya, Tomoki Ariyama, Kazuma Kokuta, Suchun Xie, Kai Sato, Shusaku Sone, Ryohei Kamei, Reina Akama, Yuichiroh Matsubayashi, Keisuke Sakaguchi |
SIGDIAL | 12 |
| 2023 | ELQA: A Corpus of Metalinguistic Questions and Answers about EnglishabstractWe present ELQA, a corpus of questions and answers in and about the English language.Collected from two online forums, the >70k questions (from English learners and others) cover wide-ranging topics including grammar, meaning, fluency, and etymology.The answers include descriptions of general properties of English vocabulary and grammar as well as explanations about specific (correct and incorrect) usage examples.Unlike most NLP datasets, this corpus is metalinguistic-it consists of language about language.As such, it can facilitate investigations of the metalinguistic capabilities of NLU models, as well as educational applications in the language learning domain.To study this, we define a free-form question answering task on our dataset and conduct evaluations on multiple LLMs (Large Language Models) to analyze their capacity to generate metalinguistic answers. Shabnam Behzad, Keisuke Sakaguchi, Nathan Schneider 0001, Amir Zeldes |
ACL (1) | 2 |
| 2023 | I2D2: Inductive Knowledge Distillation with NeuroLogic and Self-ImitationabstractChandra Bhagavatula, Jena D. Hwang, Doug Downey, Ronan Le Bras, Ximing Lu, Lianhui Qin, Keisuke Sakaguchi, Swabha Swayamdipta, Peter West, Yejin Choi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Chandra Bhagavatula, Jena D. Hwang, Doug Downey, Ronan Le Bras 0001, Ximing Lu, Lianhui Qin, Keisuke Sakaguchi, Swabha Swayamdipta, Peter West, Yejin Choi 0001 |
ACL (1) | 7 |
| 2023 | Do Deep Neural Networks Capture Compositionality in Arithmetic Reasoning?abstractKeito Kudo, Yoichi Aoki, Tatsuki Kuribayashi, Ana Brassard, Masashi Yoshikawa, Keisuke Sakaguchi, Kentaro Inui. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Keito Kudo, Yoichi Aoki, Tatsuki Kuribayashi, Ana Brassard, Masashi Yoshikawa, Keisuke Sakaguchi, Kentaro Inui |
EACL | 6 |
| 2023 | RealTime QA: What's the Answer Right Now?abstractWe introduce RealTime QA, a dynamic question answering (QA) platform that announces questions and evaluates systems on a regular basis (weekly in this version). RealTime QA inquires about the current world, and QA systems need to answer questions about novel events or information. It therefore challenges static, conventional assumptions in open-domain QA datasets and pursues instantaneous applications. We build strong baseline models upon large pretrained language models, including GPT-3 and T5. Our benchmark is an ongoing effort, and this paper presents real-time evaluation results over the past year. Our experimental results show that GPT-3 can often properly update its generation results, based on newly-retrieved documents, highlighting the importance of up-to-date information retrieval. Nonetheless, we find that GPT-3 tends to return outdated answers when retrieved documents do not provide sufficient information to find an answer. This suggests an important avenue for future research: can an open-domain QA system identify such unanswerable cases and communicate with the user or even the retrieval module to modify the retrieval results? We hope that RealTime QA will spur progress in instantaneous applications of question answering and beyond. Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Ronan Le Bras 0001, Akari Asai, Xinyan Yu 0001, Dragomir R. Radev, Noah A. Smith, Yejin Choi 0001, Kentaro Inui |
NeurIPS | 2 |
| 2022 | Twist Decoding: Diverse Generators Guide Each OtherabstractJungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Hao Peng, Ximing Lu, Dragomir Radev, Yejin Choi, Noah A. Smith. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras 0001, Hao Peng 0009, Ximing Lu, Dragomir R. Radev, Yejin Choi 0001, Noah A. Smith |
EMNLP | 2 |
| 2022 | Bidimensional Leaderboards: Generate and Evaluate Language Hand in HandabstractJungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander Fabbri, Yejin Choi, Noah Smith. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras 0001, Lavinia Dunagan, Jacob Morrison, Alexander R. Fabbri, Yejin Choi 0001, Noah A. Smith |
NAACL-HLT | 2 |
| 2022 | Transparent Human Evaluation for Image CaptioningabstractJungo Kasai, Keisuke Sakaguchi, Lavinia Dunagan, Jacob Morrison, Ronan Le Bras, Yejin Choi, Noah Smith. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jungo Kasai, Keisuke Sakaguchi, Lavinia Dunagan, Jacob Morrison, Ronan Le Bras 0001, Yejin Choi 0001, Noah A. Smith |
NAACL-HLT | 2 |
| 2021 | (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsabstractRecent years have brought about a renewed interest in commonsense representation and reasoning in the field of natural language understanding. The development of new commonsense knowledge graphs (CSKG) has been central to these advances as their diverse facts can be used and referenced by machine learning models for tackling new and challenging tasks. At the same time, there remain questions about the quality and coverage of these resources due to the massive scale required to comprehensively encompass general commonsense knowledge. In this work, we posit that manually constructed CSKGs will never achieve the coverage necessary to be applicable in all situations encountered by NLP agents. Therefore, we propose a new evaluation framework for testing the utility of KGs based on how effectively implicit knowledge representations can be learned from them. With this new goal, we propose Atomic 2020, a new CSKG of general-purpose commonsense knowledge containing knowledge that is not readily available in pretrained language models. We evaluate its properties in comparison with other leading CSKGs, performing the first large-scale pairwise study of commonsense knowledge resources. Next, we show that Atomic 2020 is better suited for training knowledge models that can generate accurate, representative knowledge for new, unseen entities and events. Finally, through human evaluation, we show that the few-shot performance of GPT-3 (175B parameters), while impressive, remains ~12 absolute points lower than a BART-based knowledge model trained on Atomic 2020 despite using over 430x fewer parameters. Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras 0001, Jeff Da, Keisuke Sakaguchi, Antoine Bosselut, Yejin Choi 0001 |
AAAI | 5 |
| 2020 | WinoGrande: An Adversarial Winograd Schema Challenge at ScaleabstractThe Winograd Schema Challenge (WSC) (Levesque, Davis, and Morgenstern 2011), a benchmark for commonsense reasoning, is a set of 273 expert-crafted pronoun resolution problems originally designed to be unsolvable for statistical models that rely on selectional preferences or word associations. However, recent advances in neural language models have already reached around 90% accuracy on variants of WSC. This raises an important question whether these models have truly acquired robust commonsense capabilities or whether they rely on spurious biases in the datasets that lead to an overestimation of the true capabilities of machine commonsense.To investigate this question, we introduce WinoGrande, a large-scale dataset of 44k problems, inspired by the original WSC design, but adjusted to improve both the scale and the hardness of the dataset. The key steps of the dataset construction consist of (1) a carefully designed crowdsourcing procedure, followed by (2) systematic bias reduction using a novel AfLite algorithm that generalizes human-detectable word associations to machine-detectable embedding associations. The best state-of-the-art methods on WinoGrande achieve 59.4 – 79.1%, which are ∼15-35% (absolute) below human performance of 94.0%, depending on the amount of the training data allowed (2% – 100% respectively).Furthermore, we establish new state-of-the-art results on five related benchmarks — WSC (→ 90.1%), DPR (→ 93.1%), COPA(→ 90.6%), KnowRef (→ 85.6%), and Winogender (→ 97.1%). These results have dual implications: on one hand, they demonstrate the effectiveness of WinoGrande when used as a resource for transfer learning. On the other hand, they raise a concern that we are likely to be overestimating the true capabilities of machine commonsense across all these benchmarks. We emphasize the importance of algorithmic bias reduction in existing and future benchmarks to mitigate such overestimation. Keisuke Sakaguchi, Ronan Le Bras 0001, Chandra Bhagavatula, Yejin Choi 0001 |
AAAI | 1 |
| 2020 | Uncertain Natural Language InferenceabstractWe introduce Uncertain Natural Language Inference (UNLI), a refinement of Natural Language Inference (NLI) that shifts away from categorical labels, targeting instead the direct prediction of subjective probability assessments.We demonstrate the feasibility of collecting annotations for UNLI by relabeling a portion of the SNLI dataset under a probabilistic scale, where items even with the same categorical label differ in how likely people judge them to be true given a premise.We describe a direct scalar regression modeling approach, and find that existing categorically labeled NLI data can be used in pre-training.Our best models approach human performance, demonstrating models may be capable of more subtle inferences than the categorical bin assignment employed in current NLI tasks. Tongfei Chen, Zhengping Jiang, Adam Poliak, Keisuke Sakaguchi, Benjamin Van Durme |
ACL | 4 |
| 2020 | A Dataset for Tracking Entities in Open Domain Procedural TextabstractNiket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal, Peter Clark, Michal Guerquin, Kyle Richardson, Eduard Hovy. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Niket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal, Peter Clark, Michal Guerquin, Kyle Richardson 0001, Eduard H. Hovy |
EMNLP (1) | 2 |
| 2020 | Abductive Commonsense Reasoning
Chandra Bhagavatula, Ronan Le Bras 0001, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Yih, Yejin Choi 0001 |
ICLR | 4 |
| 2020 | The Universal Decompositional Semantics Dataset and Decomp ToolkitabstractWe present the Universal Decompositional Semantics (UDS) dataset (v1.0), which is bundled with the Decomp toolkit (v0.1). UDS1.0 unifies five high-quality, decompositional semantics-aligned annotation sets within a single semantic graph specification—with graph structures defined by the predicative patterns produced by the PredPatt tool and real-valued node and edge attributes constructed using sophisticated normalization procedures. The Decomp toolkit provides a suite of Python 3 tools for querying UDS graphs using SPARQL. Both UDS1.0 and Decomp0.1 are publicly available at http://decomp.io. Aaron Steven White, Elias Stengel-Eskin, Siddharth Vashishtha, Venkata Subrahmanyan Govindarajan, Dee Ann Reisinger, Tim Vieira, Keisuke Sakaguchi, Sheng Zhang 0012, Francis Ferraro, Rachel Rudinger, Kyle Rawlins, Benjamin Van Durme |
LREC | 7 |
| 2019 | WIQA: A dataset for "What if..." reasoning over procedural textabstractNiket Tandon, Bhavana Dalvi, Keisuke Sakaguchi, Peter Clark, Antoine Bosselut. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Niket Tandon, Bhavana Dalvi, Keisuke Sakaguchi, Peter Clark, Antoine Bosselut |
EMNLP/IJCNLP (1) | 3 |
| 2018 | Efficient Online Scalar Annotation with Bounded SupportabstractWe describe a novel method for efficiently eliciting scalar annotations for dataset construction and system quality estimation by human judgments.We contrast direct assessment (annotators assign scores to items directly), online pairwise ranking aggregation (scores derive from annotator comparison of items), and a hybrid approach (EASL: Efficient Annotation of Scalar Labels) proposed here.Our proposal leads to increased correlation with ground truth, at far greater annotator efficiency, suggesting this strategy as an improved mechanism for dataset creation and manual system evaluation. Keisuke Sakaguchi, Benjamin Van Durme |
ACL (1) | 1 |
| 2017 | Robsut Wrod Reocginiton via Semi-Character Recurrent Neural NetworkabstractLanguage processing mechanism by humans is generally more robust than computers. The Cmabrigde Uinervtisy (Cambridge University) effect from the psycholinguistics literature has demonstrated such a robust word processing mechanism, where jumbled words (e.g. Cmabrigde / Cambridge) are recognized with little cost. On the other hand, computational models for word recognition (e.g. spelling checkers) perform poorly on data with such noise. Inspired by the findings from the Cmabrigde Uinervtisy effect, we propose a word recognition model based on a semi-character level recurrent neural network (scRNN). In our experiments, we demonstrate that scRNN has significantly more robust performance in word spelling correction (i.e. word recognition) compared to existing spelling checkers and character-based convolutional neural network. Furthermore, we demonstrate that the model is cognitively plausible by replicating a psycholinguistics experiment about human reading difficulty using our model. Keisuke Sakaguchi, Kevin Duh, Matt Post, Benjamin Van Durme |
AAAI | 1 |
| 2016 | Phrase Structure Annotation and Parsing for Learner English
Ryo Nagata, Keisuke Sakaguchi |
ACL (1) | 2 |
| 2016 | There's No Comparison: Reference-less Evaluation Metrics in Grammatical Error CorrectionabstractCurrent methods for automatically evaluating grammatical error correction (GEC) systems rely on gold-standard references.However, these methods suffer from penalizing grammatical edits that are correct but not in the gold standard.We show that reference-less grammaticality metrics correlate very strongly with human judgments and are competitive with the leading reference-based evaluation metrics.By interpolating both methods, we achieve state-of-the-art correlation with human judgments.Finally, we show that GEC metrics are much more reliable when they are calculated at the sentence level instead of the corpus level.We have set up a CodaLab site for benchmarking GEC output using a common dataset and different evaluation metrics. Courtney Napoles, Keisuke Sakaguchi, Joel R. Tetreault |
EMNLP | 2 |
| 2016 | Universal Decompositional Semantics on Universal DependenciesabstractAaron Steven White, Drew Reisinger, Keisuke Sakaguchi, Tim Vieira, Sheng Zhang, Rachel Rudinger, Kyle Rawlins, Benjamin Van Durme. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2016. Aaron Steven White, Dee Ann Reisinger, Keisuke Sakaguchi, Tim Vieira, Sheng Zhang 0012, Rachel Rudinger, Kyle Rawlins, Benjamin Van Durme |
EMNLP | 3 |
| 2016 | Reassessing the Goals of Grammatical Error Correction: Fluency Instead of GrammaticalityabstractThe field of grammatical error correction (GEC) has grown substantially in recent years, with research directed at both evaluation metrics and improved system performance against those metrics. One unvisited assumption, however, is the reliance of GEC evaluation on error-coded corpora, which contain specific labeled corrections. We examine current practices and show that GEC’s reliance on such corpora unnaturally constrains annotation and automatic evaluation, resulting in (a) sentences that do not sound acceptable to native speakers and (b) system rankings that do not correlate with human judgments. In light of this, we propose an alternate approach that jettisons costly error coding in favor of unannotated, whole-sentence rewrites. We compare the performance of existing metrics over different gold-standard annotations, and show that automatic evaluation with our new annotation scheme has very strong correlation with expert rankings (ρ = 0.82). As a result, we advocate for a fundamental and necessary shift in the goal of GEC, from correcting small, labeled error types, to producing text that has native fluency. Keisuke Sakaguchi, Courtney Napoles, Matt Post, Joel R. Tetreault |
Trans. Assoc. Comput. Linguistics | 1 |
| 2015 | Effective Feature Integration for Automated Short Answer ScoringabstractA major opportunity for NLP to have a realworld impact is in helping educators score student writing, particularly content-based writing (i.e., the task of automated short answer scoring).A major challenge in this enterprise is that scored responses to a particular question (i.e., labeled data) are valuable for modeling but limited in quantity.Additional information from the scoring guidelines for humans, such as exemplars for each score level and descriptions of key concepts, can also be used.Here, we explore methods for integrating scoring guidelines and labeled responses, and we find that stacked generalization (Wolpert, 1992) improves performance, especially for small training sets. Keisuke Sakaguchi, Michael Heilman, Nitin Madnani |
HLT-NAACL | 1 |
| 2012 | Joint English Spelling Error Correction and POS Tagging for Language Learners Writing
Keisuke Sakaguchi, Tomoya Mizumoto, Mamoru Komachi, Yuji Matsumoto 0001 |
COLING | 1 |