EDBT 2026 Demo / reviewers in the wild / expert
Masaki Uto
dblp:163/7928
· DBLP profile ↗
22ranked-venue papers
12as first author
12since 2021 · last 2026
0000-0002-9330-5158ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 17 · 11 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 11 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-Based Virtual Standardized Patients with Response Excessiveness Suppression via Direct Preference Optimization for Medical Interview Examinations
Naoki Shindo, Masaki Uto |
AIED (1) | 2 |
| 2026 | Has Automated Essay Scoring Reached Sufficient Accuracy? Deriving Achievable QWK Ceilings from Classical Test Theory
Masaki Uto |
AIED (1) | 1 |
| 2026 | An Item Response Theory Model for Addressing Halo Effects in Performance Assessment
Masaki Uto, Ryota Kitakaze |
AIED (1) | 1 |
| 2024 | Enhancing Diversity in Difficulty-Controllable Question Generation for Reading Comprehension via Extended T5abstractRecently, automatic generation of reading comprehension questions with controllable difficulty levels has attracted growing interest for educational purposes. The latest method for difficulty-controllable question generation employs a two-stage mechanism utilizing two independent large language models. Specifically, given a reading passage and a difficulty level as inputs, it first produces a reference answer using BERT and then generates a corresponding question using GPT-2. However, this two-stage approach has the limitation that the questions generated depend strongly on the reference answers produced beforehand, restricting the diversity of questions. To overcome this limitation, we propose an end-to-end method that enables the simultaneous generation of questions and reference answers by extending T5, a large language model with an encoder mechanism equivalent to BERT and a decoder mechanism equivalent to GPT-2. In our method, T5 is extended to generate answers from its encoder and questions from its decoder, with the encoder's output vector passed to the decoder. Experiments using a benchmark dataset demonstrate that our method significantly improves the diversity of both questions and answers compared with the conventional method while maintaining difficulty controllability. Teruyoshi Goto, Yuto Tomikawa, Masaki Uto |
ICCE | 3 |
| 2024 | Difficulty-Controllable Reading Comprehension Question Generation Considering the Difficulty of Reading PassagesabstractIn recent years, various question generation (QG) methods for reading comprehension have been that automatically generate questions related to given reading passages. Specifically, QG methods based on deep neural networks have succeeded in generating high-quality questions. To apply such QG methods in educational systems, such as intelligent tutoring systems and adaptive learning systems, it is crucial to generate questions with difficulty levels that are appropriate for each learner's reading ability. To meet this need, several difficulty-controllable QG have been proposed recently. However, a limitation of existing difficulty- controllable methods is that they overlook the difficulty of the reading passages, which are given as the input context for QG. Since the difficulty of reading passages can affect the difficulty of the generated questions, selecting reading passages with appropriate difficulty is crucial. Therefore, in this study, we develop a difficulty-controllable QG that includes a mechanism for selecting reading passages with appropriate difficulty for each learner. Our approach begins with the of a new item response theory (IRT) model capable of simultaneously estimating the difficulty of both questions and reading passages. Using the developed IRT model and the latest IRT- based difficulty-controllable QG method, we propose a framework to select reading passages and generate questions that are appropriate for each learner's reading ability. Yuto Tomikawa, Masaki Uto |
ICCE | 2 |
| 2024 | Estimating Scores of Critical Thinking Ability Using Essay Text AssessmentsabstractThe possibility of predicting individual scores of critical thinking ability from estimated ratings of essay texts was examined using machine learning techniques. First, the feasibility of predicting critical thinking ability scores from scores of personality and literacy of science and technology surveys was confirmed. Scores for personal characteristics were estimated using automated assessments which were based on a pre-trained language processing model of two types of essay texts. Second, variables employed in the surveys used to predict scores of critical thinking ability were replaced step by step with estimated values from the essay texts. Prediction performance decreased gradually, and the accuracy with all estimated values was above the level of performance which was predicted using the essay text analysis. As a result, the possibility of predicting scores from essay text assessments was examined. Minoru Nakayama, Masaki Uto, Satoru Kikuchi, Hiroh Yamamoto |
IV | 2 |
| 2023 | Neural Automated Essay Scoring Considering Logical Structure
Misato Yamaura, Itsuki Fukuda, Masaki Uto |
AIED | 3 |
| 2023 | Neural Automated Short-Answer Grading Considering Examinee-Specific FeaturesabstractAutomated short-answer grading (ASAG) is the task of automatically assigning scores to examinees' textual responses to short-answer questions. Recently, various ASAG models based on deep neural networks (DNNs) have been proposed and some have achieved high accuracy. Conventional ASAG models are generally trained and used independently for each short-answer question, even when a given test contains multiple short-answer questions. However, because tests are tools for evaluating particular examinee traits, multiple questions on the same test are generally designed to measure similar latent traits in examinees. Latent examinee traits across multiple short-answer questions may thus function as effective auxiliary features for ASAG. We therefore propose a new DNN-based ASAG model with an examinee-aware architecture that extracts examinee-specific features, including latent traits, from answer texts across multiple questions. Masaki Uto |
ICALT | 1 |
| 2023 | Feasibility of Prediction of Student's Characteristics Using Texts of Essays Written During a Fully Online CourseabstractThe possibility of predicting anticipated scores representing student's characteristics from their essay texts using a machine learning technique was examined. Student's characteristics include critical thinking disposition ability, disaster-prevention conscientiousness, and personality. First, the potential for predicting essay assessment scores from essay texts using machine learning with each set of essay reports or essay comments was confirmed. However, prediction performance of one other set of essays and other scores of student's characteristics was insufficient. When the training procedure was re-designed using both sets of essays, performance for some scores improved. The limited performance was examined and further improvement procedure was discussed. Minoru Nakayama, Masaki Uto, Satoru Kikuchi, Hiroh Yamamoto |
IV | 2 |
| 2022 | Analytic Automated Essay Scoring Based on Deep Neural Networks Integrating Multidimensional Item Response TheoryabstractEssay exams have been attracting attention as a way of measuring the higher-order abilities of examinees, but they have two major drawbacks in that grading them is expensive and raises questions about fairness. As an approach to overcome these problems, automated essay scoring (AES) is in increasing need. Many AES models based on deep neural networks have been proposed in recent years and have achieved high accuracy, but most of these models are designed to predict only a single overall score. However, to provide detailed feedback in practical situations, we often require not only the overall score but also analytic scores corresponding to various aspects of the essay. Several neural AES models that can predict both the analytic scores and the overall score have also been proposed for this very purpose. However, conventional models are designed to have complex neural architectures for each analytic score, which makes interpreting the score prediction difficult. To improve the interpretability of the prediction while maintaining scoring accuracy, we propose a new neural model for automated analytic scoring that integrates a multidimensional item response theory model, which is a popular psychometric model. Takumi Shibata, Masaki Uto |
COLING | 2 |
| 2021 | Integration of Automated Essay Scoring Models Using Item Response Theory
Itsuki Aomi, Emiko Tsutsumi, Masaki Uto, Maomi Ueno |
AIED (2) | 3 |
| 2021 | A Multidimensional Item Response Theory Model for Rubric-Based Writing Assessment
Masaki Uto |
AIED (1) | 1 |
| 2020 | Robust Neural Automated Essay Scoring Using Item Response Theory
Masaki Uto, Masashi Okano |
AIED (1) | 1 |
| 2020 | Automated Short-Answer Grading Using Deep Neural Networks and Item Response Theory
Masaki Uto, Yuto Uchida |
AIED (2) | 1 |
| 2020 | Neural Automated Essay Scoring Incorporating Handcrafted FeaturesabstractAutomated essay scoring (AES) is the task of automatically assigning scores to essays as an alternative to grading by human raters.Conventional AES typically relies on handcrafted features, whereas recent studies have proposed AES models based on deep neural networks (DNNs) to obviate the need for feature engineering.Furthermore, hybrid methods that integrate handcrafted features in a DNN-AES model have been recently developed and have achieved state-of-the-art accuracy.One of the most popular hybrid methods is formulated as a DNN-AES model with an additional recurrent neural network (RNN) that processes a sequence of handcrafted sentencelevel features.However, this method has the following problems: 1) It cannot incorporate effective essay-level features developed in previous AES research.2) It greatly increases the numbers of model parameters and tuning parameters, increasing the difficulty of model training.3) It has an additional RNN to process sentence-level features, enabling extension to various DNN-AES models complex.To resolve these problems, we propose a new hybrid method that integrates handcrafted essay-level features into a DNN-AES model.Specifically, our method concatenates handcrafted essay-level features to a distributed essay representation vector, which is obtained from an intermediate layer of a DNN-AES model.Our method is a simple DNN-AES extension, but significantly improves scoring accuracy. Masaki Uto, Yikuan Xie, Maomi Ueno |
COLING | 1 |
| 2020 | Impact of the number of peers on a mutual assessment as learner's performance in a simulated MOOC environment using the IRT modelabstractWe discuss the problem of setting the best number of peers to which a given evaluation job should be assigned, in a Peer Assessment setting. The Peer Assessment is supposed to happen in a large scale class, such as in the case of Massive Open Online Courses. We use a dataset that simulate a large class (1000 students), based on Gaussian distributions of the Student Model features. Such features are related to the student's proficiency, and assessment capability. The number of peers assigned to the same evaluation job was controlled from 3 to 50 in 6 steps using 10-point scale. The abilities of participants were estimated using Item Response Theory. All parameters of IRT models, which is called as Generalized Partial Credit Model, such as "ability", "consistency", and "strictness", were estimated well using MCMC technique; their standard deviation errors gradually decrease with the number of peers. As a preliminary result of optimisation, an appropriate number of peers was 15 as comparing the stadardised errors across the conditions. Minoru Nakayama, Filippo Sciarrone, Masaki Uto, Marco Temperini |
IV | 3 |
| 2019 | Rater-Effect IRT Model Integrating Supervised LDA for Accurate Measurement of Essay Writing Ability
Masaki Uto |
AIED (1) | 1 |
| 2018 | Item Response Theory Without Restriction of Equal Interval Scale for Rater's Score
Masaki Uto, Maomi Ueno |
AIED (2) | 1 |
| 2017 | Group Optimization to Maximize Peer Assessment Accuracy Using Item Response Theory
Masaki Uto, Nguyen Duc Thien, Maomi Ueno |
AIED | 1 |
| 2015 | Item Response Model with Lower Order Parameters for Peer Assessment
Masaki Uto, Maomi Ueno |
AIED | 1 |
| 2015 | Academic Writing Support System Using Bayesian NetworksabstractFor academic writing, elaborating an argument particularly addressing an argument strength is important to establish causal relations between sentences. However, when an argument becomes large or complex, elaborating an argument considering the argument strength is difficult. To solve this problem, this article presents a proposal for an argument elaboration support system using a Bayesian network representation of the Toulmin model. Using that Bayesian network representation, the proposed system can estimate argument strength, sentence validity, and sentence influence. Moreover, it can generate optimal advice for revising the argument. Masaki Uto, Maomi Ueno |
ICALT | 1 |
| 2015 | Reliable Peer Assessment for Team-project-based Learning using Item Response Theory
Nguyen Duc Thien, Masaki Uto, Yu Abe, Maomi Ueno |
ICCE | 2 |