EDBT 2026 Demo / reviewers in the wild / expert
Kentaro Inui
dblp:90/3315
· DBLP profile ↗
151ranked-venue papers
3as first author
50since 2021 · last 2026
0000-0001-6510-604XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 140 · 3 first-author · 42 since 2021Human-computer interaction and ubiquitous computing · 9 · 7 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cell-Based Representation of Relational Binding in Language ModelsabstractUnderstanding a discourse requires tracking entities and the relations that hold between them. While Large Language Models (LLMs) perform well on relational reasoning, the mechanism by which they bind entities, relations, and attributes remains unclear. We study discourse-level relational binding and show that LLMs encode it via a Cell-based Binding Representation (CBR): a low-dimensional linear subspace in which each “cell” corresponds to an entity–relation index pair, and bound attributes are retrieved from the corresponding cell during inference. Using controlled multi-sentence data annotated with entity and relation indices, we identify the CBR subspace by decoding these indices from attribute-token activations with Partial Least Squares regression. Across domains and two model families, the indices are linearly decodable and form a grid-like geometry in the projected space. We further find that context-specific CBR representations are related by translation vectors in activation space, enabling cross-context transfer. Finally, activation patching shows that manipulating this subspace systematically changes relational predictions and that perturbing it disrupts performance, providing causal evidence that LLMs rely on CBR for relational binding. Qin Dai, Benjamin Heinzerling, Kentaro Inui |
ACL (1) | 3 |
| 2026 | Why Mean Pooling Works: Quantifying Second-Order Collapse in Text EmbeddingsabstractFor constructing text embeddings, mean pooling, which averages token embeddings, is the standard approach.This paper examines whether mean pooling actually works well in real text encoders.First, we note that mean pooling can collapse information beyond the first-order statistics of the token embeddings, such as second-order statistics that capture their spatial structure, potentially mapping distinct token embedding distributions to similar text embeddings.Motivated by this concern, we propose a simple metric to quantify such a collapse induced by mean pooling.Then, using this metric, we empirically measure how often this collapse arises in actual models and texts, and find that mean pooling works well in modern text encoders.In particular, this collapse is less likely to arise in contrastive fine-tuned text encoders than in their pretrained backbone models.We also find that the robustness of these text encoders to collapse stems from the concentration of token embeddings within each text.In addition, we find that robustness to this collapse, as quantified by our proposed metric, correlates with downstream task performance.Overall, our findings help explain why modern text encoders remain effective despite relying on seemingly coarse mean pooling. Tomomasa Hara, Hiroto Kurita, Masaaki Imaizumi, Kentaro Inui, Sho Yokoi |
ACL (1) | 4 |
| 2025 | Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal CyclesabstractMunachiso S Nwadike, Zangir Iklassov, Toluwani Aremu, Tatsuya Hiraoka, Benjamin Heinzerling, Velibor Bojkovic, Hilal AlQuabeh, Martin Takáč, Kentaro Inui. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Munachiso Nwadike, Zangir Iklassov, Toluwani Aremu, Tatsuya Hiraoka, Benjamin Heinzerling, Velibor Bojkovic, Hilal AlQuabeh, Martin Takác 0001, Kentaro Inui |
ACL (1) | 9 |
| 2025 | Annotating Errors in English Learners' Written Language Production: Advancing Automated Written Feedback Systems
Steven Coyne, Diana Galván-Sosa, Ryan Spring, Camélia Guerraoui, Michael Zock, Keisuke Sakaguchi, Kentaro Inui |
AIED (4) | 7 |
| 2025 | Beyond Click to Cognition: Effective Interventions for Promoting Examination of False Beliefs in Misinformation
Yuko Tanaka, Hiromi Arai, Miwa Inuzuka, Yoichi Takahashi, Minao Kukita, Ryuta Iseki, Kentaro Inui |
CHI | 7 |
| 2025 | MQM-Chat: Multidimensional Quality Metrics for Chat TranslationabstractThe complexities of chats, such as the stylized contents specific to source segments and dialogue consistency, pose significant challenges for machine translation. Recognizing the need for a precise evaluation metric to address the issues associated with chat translation, this study introduces Multidimensional Quality Metrics for Chat Translation (MQM-Chat), which encompasses seven error types, including three specifically designed for chat translations: ambiguity and disambiguation, buzzword or loanword issues, and dialogue inconsistency. In this study, human annotations were applied to the translations of chat data generated by five translation models. Based on the error distribution of MQM-Chat and the performance of relabeling errors into chat-specific types, we concluded that MQM-Chat effectively classified the errors while highlighting chat-specific issues explicitly. The results demonstrate that MQM-Chat can qualify both the lexical accuracy and semantical accuracy of translation models in chat translation tasks. Yunmeng Li, Jun Suzuki 0001, Makoto Morishita, Kaori Abe, Kentaro Inui |
COLING | 5 |
| 2025 | SPIRIT: Patching Speech Language Models against Jailbreak AttacksabstractSpeech Language Models (SLMs) enable natural interactions via spoken instructions, which more effectively capture user intent by detecting nuances in speech.The richer speech signal introduces new security risks compared to textbased models, as adversaries can better bypass safety mechanisms by injecting imperceptible noise to speech.We analyze adversarial attacks under white-box access and find that SLMs are substantially more vulnerable to jailbreak attacks, which can achieve a perfect 100% attack success rate in some instances.To improve security, we propose post-hoc patching defenses used to intervene during inference by modifying the SLM's activations that improve robustness up to 99% with (i) negligible impact on utility and (ii) without any re-training.We conduct ablation studies to maximize the efficacy of our defenses and improve the utility/security trade-off, validated with largescale benchmarks unique to SLMs.The code is available at: https://github.com/ mbzuai-nlp/spirit-breaking.git Amirbek Djanibekov, Nurdaulet Mukhituly, Kentaro Inui, Hanan Aldarmaki, Nils Lukas |
EMNLP | 3 |
| 2025 | Identification of Multiple Logical Interpretations in Counter-ArgumentsabstractWenzhi Wang, Paul Reisert, Shoichi Naito, Naoya Inoue, Machi Shimmei, Surawat Pothong, Jungmin Choi, Kentaro Inui. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Paul Reisert, Shoichi Naito, Naoya Inoue, Machi Shimmei, Surawat Pothong, Jungmin Choi, Kentaro Inui |
EMNLP | 8 |
| 2025 | Uncovering the Spectral Bias in Diagonal State Space ModelsabstractCurrent methods for initializing state space models (SSMs) parameters mainly rely on the \textit{HiPPO framework}, which is based on an online approximation of orthogonal polynomials. Recently, diagonal alternatives have shown to reach a similar level of performance while being significantly more efficient due to the simplification in the kernel computation. However, the \textit{HiPPO framework} does not explicitly study the role of its diagonal variants. In this paper, we take a further step to investigate the role of diagonal SSM initialization schemes from the frequency perspective. Our work seeks to systematically understand how to parameterize these models and uncover the learning biases inherent in such diagonal state-space models. Based on our observations, we propose a diagonal initialization on the discrete Fourier domain \textit{S4D-DFouT}. The insights in the role of pole placing in the initialization enable us to further scale them and achieve state-of-the-art results on the Long Range Arena benchmark, allowing us to train from scratch on very large datasets as PathX-256. Ruben Solozabal, Velibor Bojkovic, Hilal AlQuabeh, Kentaro Inui, Martin Takác 0001 |
NeurIPS | 4 |
| 2025 | Large Language Models Are Human-Like InternallyabstractAbstract Recent cognitive modeling studies have reported that larger language models (LMs) exhibit a poorer fit to human reading behavior (Oh and Schuler, 2023b; Shain et al., 2024; Kuribayashi et al., 2024), leading to claims of their cognitive implausibility. In this paper, we revisit this argument through the lens of mechanistic interpretability and argue that prior conclusions were skewed by an exclusive focus on the final layers of LMs. Our analysis reveals that next-word probabilities derived from internal layers of larger LMs align with human sentence processing data as well as, or better than, those from smaller LMs. This alignment holds consistently across behavioral (self-paced reading times, gaze durations, MAZE task processing times) and neurophysiological (N400 brain potentials) measures, challenging earlier mixed results and suggesting that the cognitive plausibility of larger LMs has been underestimated. Furthermore, we first identify an intriguing relationship between LM layers and human measures: Earlier layers correspond more closely with fast gaze durations, while later layers better align with relatively slower signals such as N400 potentials and MAZE processing times. Our work opens new avenues for interdisciplinary research at the intersection of mechanistic interpretability and cognitive modeling.1 Tatsuki Kuribayashi, Yohei Oseki, Souhaib Ben Taieb, Kentaro Inui, Timothy Baldwin |
Trans. Assoc. Comput. Linguistics | 4 |
| 2024 | To Drop or Not to Drop? Predicting Argument Ellipsis Judgments: A Case Study in JapaneseabstractSpeakers sometimes omit certain arguments of a predicate in a sentence; such omission is especially frequent in pro-drop languages. This study addresses a question about ellipsis—what can explain the native speakers’ ellipsis decisions?—motivated by the interest in human discourse processing and writing assistance for this choice. To this end, we first collect large-scale human annotations of whether and why a particular argument should be omitted across over 2,000 data points in the balanced corpus of Japanese, a prototypical pro-drop language. The data indicate that native speakers overall share common criteria for such judgments and further clarify their quantitative characteristics, e.g., the distribution of related linguistic factors in the balanced corpus. Furthermore, the performance of the language model–based argument ellipsis judgment model is examined, and the gap between the systems’ prediction and human judgments in specific linguistic aspects is revealed. We hope our fundamental resource encourages further studies on natural human ellipsis judgment. Yukiko Ishizuki, Tatsuki Kuribayashi, Yuichiroh Matsubayashi, Ryohei Sasano, Kentaro Inui |
LREC/COLING | 5 |
| 2024 | First Heuristic Then Rational: Dynamic Use of Heuristics in Language Model ReasoningabstractYoichi Aoki, Keito Kudo, Tatsuki Kuribayashi, Shusaku Sone, Masaya Taniguchi, Keisuke Sakaguchi, Kentaro Inui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yoichi Aoki, Keito Kudo, Tatsuki Kuribayashi, Shusaku Sone, Masaya Taniguchi, Keisuke Sakaguchi, Kentaro Inui |
EMNLP | 7 |
| 2024 | Representational Analysis of Binding in Language ModelsabstractEntity tracking is essential for complex reasoning.To perform in-context entity tracking, language models (LMs) must bind an entity to its attribute (e.g., bind a container to its content) to recall attribute for a given entity.For example, given a context mentioning "The coffee is in Box Z, the stone is in Box M, the map is in Box H", to infer "Box Z contains the coffee" later, LMs must bind "Box Z" to "coffee".To explain the binding behaviour of LMs, Feng and Steinhardt (2023) introduce a Binding ID mechanism and state that LMs use a abstract concept called Binding ID (BI) to internally mark entity-attribute pairs.However, they have not captured the Ordering ID (OI) from entity activations that directly determines the binding behaviour.In this work, we provide a novel view of the BI mechanism by localizing OI and proving the causality between OI and binding behaviour.Specifically, by leveraging dimension reduction methods (e.g., PCA), we discover that there exists a low-rank subspace in the activations of LMs, that primarily encodes the order (i.e., OI) of entity and attribute.Moreover, we also discover the causal effect of OI on binding that when editing representations along the OI encoding direction, LMs tend to bind a given entity to other attributes accordingly.For example, by patching activations along the OI encoding direction we can make the LM to infer "Box Z contains the stone" and "Box Z contains the map".The code and datasets used in this paper are available at https://github.com/cl-tohoku/OI-Subspace. Qin Dai, Benjamin Heinzerling, Kentaro Inui |
EMNLP | 3 |
| 2024 | Flee the Flaw: Annotating the Underlying Logic of Fallacious Arguments Through Templates and Slot-fillingabstractIrfan Robbani, Paul Reisert, Surawat Pothong, Naoya Inoue, Camélia Guerraoui, Wenzhi Wang, Shoichi Naito, Jungmin Choi, Kentaro Inui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Irfan Robbani, Paul Reisert, Surawat Pothong, Naoya Inoue, Camélia Guerraoui, Shoichi Naito, Jungmin Choi, Kentaro Inui |
EMNLP | 9 |
| 2024 | Analyzing Feed-Forward Blocks in Transformers through the Lens of Attention MapsabstractTransformers are ubiquitous in wide tasks.
Interpreting their internals is a pivotal goal.
Nevertheless, their particular components, feed-forward (FF) blocks, have typically been less analyzed despite their substantial parameter amounts.
We analyze the input contextualization effects of FF blocks by rendering them in the attention maps as a human-friendly visualization scheme.
Our experiments with both masked- and causal-language models reveal that FF networks modify the input contextualization to emphasize specific types of linguistic compositions.
In addition, FF and its surrounding components tend to cancel out each other's effects, suggesting potential redundancy in the processing of the Transformer layer. Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro Inui |
ICLR | 4 |
| 2023 | Reducing the Cost: Cross-Prompt Pre-finetuning for Short Answer Scoring
Hiroaki Funayama, Yuya Asazuma, Yuichiroh Matsubayashi, Tomoya Mizumoto, Kentaro Inui |
AIED | 5 |
| 2023 | Who Does Not Benefit from Fact-checking Websites?: A Psychological Characteristic Predicts the Selective Avoidance of Clicking Uncongenial FactsabstractFact-checking messages are shared or ignored subjectively. Users tend to seek like-minded information and ignore information that conflicts with their preexisting beliefs, leaving like-minded misinformation uncontrolled on the Internet. To understand the factors that distract fact-checking engagement, we investigated the psychological characteristics associated with users’ selective avoidance of clicking uncongenial facts. In a pre-registered experiment, we measured participants’ (N = 506) preexisting beliefs about COVID-19-related news stimuli. We then examined whether they clicked on fact-checking links to false news that they believed to be accurate. We proposed an index that divided participants into fact-avoidance and fact-exposure groups using a mathematical baseline. The results indicated that 43% of participants selectively avoided clicking on uncongenial facts, keeping 93% of their false beliefs intact. Reflexiveness is the psychological characteristic that predicts selective avoidance. We discuss susceptibility to click bias that prevents users from utilizing fact-checking websites and the implications for future design. Yuko Tanaka, Miwa Inuzuka, Hiromi Arai, Yoichi Takahashi, Minao Kukita, Kentaro Inui |
CHI | 6 |
| 2023 | Fight Bias with Bias? Two Interventions for Mitigating the Selective Avoidance of Clicking Uncongenial Facts
Yuko Tanaka, Hiromi Arai, Miwa Inuzuka, Yoichi Takahashi, Minao Kukita, Kentaro Inui |
CogSci | 6 |
| 2023 | Do Deep Neural Networks Capture Compositionality in Arithmetic Reasoning?abstractKeito Kudo, Yoichi Aoki, Tatsuki Kuribayashi, Ana Brassard, Masashi Yoshikawa, Keisuke Sakaguchi, Kentaro Inui. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Keito Kudo, Yoichi Aoki, Tatsuki Kuribayashi, Ana Brassard, Masashi Yoshikawa, Keisuke Sakaguchi, Kentaro Inui |
EACL | 7 |
| 2023 | RealTime QA: What's the Answer Right Now?abstractWe introduce RealTime QA, a dynamic question answering (QA) platform that announces questions and evaluates systems on a regular basis (weekly in this version). RealTime QA inquires about the current world, and QA systems need to answer questions about novel events or information. It therefore challenges static, conventional assumptions in open-domain QA datasets and pursues instantaneous applications. We build strong baseline models upon large pretrained language models, including GPT-3 and T5. Our benchmark is an ongoing effort, and this paper presents real-time evaluation results over the past year. Our experimental results show that GPT-3 can often properly update its generation results, based on newly-retrieved documents, highlighting the importance of up-to-date information retrieval. Nonetheless, we find that GPT-3 tends to return outdated answers when retrieved documents do not provide sufficient information to find an answer. This suggests an important avenue for future research: can an open-domain QA system identify such unanswerable cases and communicate with the user or even the retrieval module to modify the retrieval results? We hope that RealTime QA will spur progress in instantaneous applications of question answering and beyond. Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Ronan Le Bras 0001, Akari Asai, Xinyan Yu 0001, Dragomir R. Radev, Noah A. Smith, Yejin Choi 0001, Kentaro Inui |
NeurIPS | 10 |
| 2023 | Use of an AI-powered Rewriting Support Software in Context with Other Tools: A Study of Non-Native English SpeakersabstractAcademic writing in English can be challenging for non-native English speakers (NNESs). AI-powered rewriting tools can potentially improve NNESs’ writing outcomes at a low cost. However, whether and how NNESs make valid assessments of the revisions provided by these algorithmic recommendations remains unclear. We report a study where NNESs leverage an AI-powered rewriting tool, Langsmith, to polish their drafted academic essays. We examined the participants’ interactions with the tool via user studies and interviews. Our data reveal that most participants used Langsmith in combination with other tools, such as machine translation (MT), and those who used MT had different ways of understanding and evaluating Langsmith’s suggestions than those who did not. Based on these findings, we assert that NNESs’ quality assessment in AI-powered rewriting tools is influenced by the simultaneous use of multiple tools, offering valuable insights into the design of future rewriting tools for NNESs. Takumi Ito, Naomi Yamashita, Tatsuki Kuribayashi, Masatoshi Hidaka, Jun Suzuki 0001, Ge Gao 0001, Jack Jamieson, Kentaro Inui |
UIST | 8 |
| 2023 | Examining the effect of whitening on static and contextualized word embeddingsabstractStatic word embeddings (SWE) and contextualized word embeddings (CWE) are the foundation of modern natural language processing. However, these embeddings suffer from spatial bias in the form of anisotropy, which has been demonstrated to reduce their performance. A method to alleviate the anisotropy is the “whitening” transformation. Whitening is a standard method in signal processing and other areas, however, its effect on SWE and CWE is not well understood. In this study, we conduct an experiment to elucidate the effect of whitening on SWE and CWE. The results indicate that whitening predominantly removes the word frequency bias in SWE, and biases other than the word frequency bias in CWE. Shota Sasaki, Benjamin Heinzerling, Jun Suzuki 0001, Kentaro Inui |
Inf. Process. Manag. | 4 |
| 2022 | Balancing Cost and Quality: An Exploration of Human-in-the-Loop Frameworks for Automated Short Answer Scoring
Hiroaki Funayama, Tasuku Sato, Yuichiroh Matsubayashi, Tomoya Mizumoto, Jun Suzuki 0001, Kentaro Inui |
AIED (1) | 6 |
| 2022 | Plausibility and Faithfulness of Feature Attribution-Based Explanations in Automated Short Answer Scoring
Tasuku Sato, Hiroaki Funayama, Kazuaki Hanawa, Kentaro Inui |
AIED (1) | 4 |
| 2022 | Topicalization in Language Models: A Case Study on JapaneseabstractHumans use different wordings depending on the context to facilitate efficient communication. For example, instead of completely new information, information related to the preceding context is typically placed at the sentence-initial position. In this study, we analyze whether neural language models (LMs) can capture such discourse-level preferences in text generation. Specifically, we focus on a particular aspect of discourse, namely the topic-comment structure. To analyze the linguistic knowledge of LMs separately, we chose the Japanese language, a topic-prominent language, for designing probing tasks, and we created human topicalization judgment data by crowdsourcing. Our experimental results suggest that LMs have different generalizations from humans; LMs exhibited less context-dependent behaviors toward topicalization judgment. These results highlight the need for the additional inductive biases to guide LMs to achieve successful discourse-level generalization. Riki Fujihara, Tatsuki Kuribayashi, Kaori Abe, Ryoko Tokuhisa, Kentaro Inui |
COLING | 5 |
| 2022 | Target-Guided Open-Domain Conversation PlanningabstractPrior studies addressing target-oriented conversational tasks lack a crucial notion that has been intensively studied in the context of goal-oriented artificial intelligence agents, namely, planning. In this study, we propose the task of Target-Guided Open-Domain Conversation Planning (TGCP) task to evaluate whether neural conversational agents have goal-oriented conversation planning abilities. Using the TGCP task, we investigate the conversation planning abilities of existing retrieval models and recent strong generative models. The experimental results reveal the challenges facing current technology. Yosuke Kishinami, Reina Akama, Shiki Sato, Ryoko Tokuhisa, Jun Suzuki 0001, Kentaro Inui |
COLING | 6 |
| 2022 | Iterative Span Selection: Self-Emergence of Resolving Orders in Semantic Role LabelingabstractSemantic Role Labeling (SRL) is the task of labeling semantic arguments for marked semantic predicates. Semantic arguments and their predicates are related in various distinct manners, of which certain semantic arguments are a necessity while others serve as an auxiliary to their predicates. To consider such roles and relations of the arguments in the labeling order, we introduce iterative argument identification (IAI), which combines global decoding and iterative identification for the semantic arguments. In experiments, we first realize that the model with random argument labeling orders outperforms other heuristic orders such as the conventional left-to-right labeling order. Combined with simple reinforcement learning, the proposed model spontaneously learns the optimized labeling orders that are different from existing heuristic orders. The proposed model with the IAI algorithm achieves competitive or outperforming results from the existing models in the standard benchmark datasets of span-based SRL: CoNLL-2005 and CoNLL-2012. Shuhei Kurita, Hiroki Ouchi, Kentaro Inui, Satoshi Sekine |
COLING | 3 |
| 2022 | Cross-stitching Text and Knowledge Graph Encoders for Distantly Supervised Relation ExtractionabstractBi-encoder architectures for distantlysupervised relation extraction are designed to make use of the complementary information found in text and knowledge graphs (KG).However, current architectures suffer from two drawbacks.They either do not allow any sharing between the text encoder and the KG encoder at all, or, in case of models with KG-to-text attention, only share information in one direction.Here, we introduce cross-stitch bi-encoders, which allow full interaction between the text encoder and the KG encoder via a cross-stitch mechanism.The cross-stitch mechanism allows sharing and updating representations between the two encoders at any layer, with the amount of sharing being dynamically controlled via cross-attention-based gates.Experimental results on two relation extraction benchmarks from two different domains show that enabling full interaction between the two encoders yields strong improvements. Qin Dai, Benjamin Heinzerling, Kentaro Inui |
EMNLP | 3 |
| 2022 | Context Limitations Make Neural Language Models More Human-LikeabstractLanguage models (LMs) have been used in cognitive modeling as well as engineering studiesthey compute information-theoretic complexity metrics that simulate humans' cognitive load during reading.This study highlights a limitation of modern neural LMs as the model of choice for this purpose: there is a discrepancy between their context access capacities and that of humans.Our results showed that constraining the LMs' context access improved their simulation of human reading behavior.We also showed that LM-human gaps in context access were associated with specific syntactic constructions; incorporating syntactic biases into LMs' context access might enhance their cognitive plausibility.1 Tatsuki Kuribayashi, Yohei Oseki, Ana Brassard, Kentaro Inui |
EMNLP | 4 |
| 2022 | COPA-SSE: Semi-structured Explanations for Commonsense ReasoningabstractWe present Semi-Structured Explanations for COPA (COPA-SSE), a new crowdsourced dataset of 9,747 semi-structured, English common sense explanations for Choice of Plausible Alternatives (COPA) questions. The explanations are formatted as a set of triple-like common sense statements with ConceptNet relations but freely written concepts. This semi-structured format strikes a balance between the high quality but low coverage of structured data and the lower quality but high coverage of free-form crowdsourcing. Each explanation also includes a set of human-given quality ratings. With their familiar format, the explanations are geared towards commonsense reasoners operating on knowledge graphs and serve as a starting point for ongoing work on improving such systems. The dataset is available at https://github.com/a-brassard/copa-sse. Ana Brassard, Benjamin Heinzerling, Pride Kavumba, Kentaro Inui |
LREC | 4 |
| 2022 | LPAttack: A Feasible Annotation Scheme for Capturing Logic Pattern of Attacks in ArgumentsabstractIn argumentative discourse, persuasion is often achieved by refuting or attacking others’ arguments. Attacking an argument is not always straightforward and often consists of complex rhetorical moves in which arguers may agree with a logic of an argument while attacking another logic. Furthermore, an arguer may neither deny nor agree with any logics of an argument, instead ignore them and attack the main stance of the argument by providing new logics and presupposing that the new logics have more value or importance than the logics presented in the attacked argument. However, there are no studies in computational argumentation that capture such complex rhetorical moves in attacks or the presuppositions or value judgments in them. To address this gap, we introduce LPAttack, a novel annotation scheme that captures the common modes and complex rhetorical moves in attacks along with the implicit presuppositions and value judgments. Our annotation study shows moderate inter-annotator agreement, indicating that human annotation for the proposed scheme is feasible. We publicly release our annotated corpus and the annotation guidelines. Farjana Sultana Mim, Naoya Inoue, Shoichi Naito, Keshav Singh 0003, Kentaro Inui |
LREC | 5 |
| 2022 | TYPIC: A Corpus of Template-Based Diagnostic Comments on ArgumentationabstractProviding feedback on the argumentation of the learner is essential for developing critical thinking skills, however, it requires a lot of time and effort. To mitigate the overload on teachers, we aim to automate a process of providing feedback, especially giving diagnostic comments which point out the weaknesses inherent in the argumentation. It is recommended to give specific diagnostic comments so that learners can recognize the diagnosis without misinterpretation. However, it is not obvious how the task of providing specific diagnostic comments should be formulated. We present a formulation of the task as template selection and slot filling to make an automatic evaluation easier and the behavior of the model more tractable. The key to the formulation is the possibility of creating a template set that is sufficient for practical use. In this paper, we define three criteria that a template set should satisfy: expressiveness, informativeness, and uniqueness, and verify the feasibility of creating a template set that satisfies these criteria as a first trial. We will show that it is feasible through an annotation study that converts diagnostic comments given in a text to a template format. The corpus used in the annotation study is publicly available. Shoichi Naito, Shintaro Sawada, Chihiro Nakagawa, Naoya Inoue, Kenshi Yamaguchi, Iori Shimizu, Farjana Sultana Mim, Keshav Singh 0003, Kentaro Inui |
LREC | 9 |
| 2022 | IRAC: A Domain-Specific Annotated Corpus of Implicit Reasoning in ArgumentsabstractThe task of implicit reasoning generation aims to help machines understand arguments by inferring plausible reasonings (usually implicit) between argumentative texts. While this task is easy for humans, machines still struggle to make such inferences and deduce the underlying reasoning. To solve this problem, we hypothesize that as human reasoning is guided by innate collection of domain-specific knowledge, it might be beneficial to create such a domain-specific corpus for machines. As a starting point, we create the first domain-specific resource of implicit reasonings annotated for a wide range of arguments, which can be leveraged to empower machines with better implicit reasoning generation ability. We carefully design an annotation framework to collect them on a large scale through crowdsourcing and show the feasibility of creating a such a corpus at a reasonable cost and high-quality. Our experiments indicate that models trained with domain-specific implicit reasonings significantly outperform domain-general models in both automatic and human evaluations. To facilitate further research towards implicit reasoning generation in arguments, we present an in-depth analysis of our corpus and crowdsourcing methodology, and release our materials (i.e., crowdsourcing guidelines and domain-specific resource of implicit reasonings). Keshav Singh 0003, Naoya Inoue, Farjana Sultana Mim, Shoichi Naito, Kentaro Inui |
LREC | 5 |
| 2022 | Law Retrieval with Supervised Contrastive Learning Using the Hierarchical Structure of Law
Jungmin Choi, Ukyo Honda, Taro Watanabe, Hiroki Ouchi, Kentaro Inui |
PACLIC | 5 |
| 2022 | N-best Response-based Analysis of Contradiction-awareness in Neural Response Generation ModelsabstractAvoiding the generation of responses that contradict the preceding context is a significant challenge in dialogue response generation.One feasible method is post-processing, such as filtering out contradicting responses from a resulting n-best response list.In this scenario, the quality of the n-best list considerably affects the occurrence of contradictions because the final response is chosen from this n-best list.This study quantitatively analyzes the contextual contradiction-awareness of neural response generation models using the consistency of the n-best lists.Particularly, we used polar questions as stimulus inputs for concise and quantitative analyses.Our tests illustrate the contradiction-awareness of recent neural response generation models and methodologies, followed by a discussion of their properties and limitations. Shiki Sato, Reina Akama, Hiroki Ouchi, Ryoko Tokuhisa, Jun Suzuki 0001, Kentaro Inui |
SIGDIAL | 6 |
| 2021 | Lower Perplexity is Not Always Human-LikeabstractTatsuki Kuribayashi, Yohei Oseki, Takumi Ito, Ryo Yoshida, Masayuki Asahara, Kentaro Inui. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Tatsuki Kuribayashi, Yohei Oseki, Takumi Ito, Ryo Yoshida, Masayuki Asahara, Kentaro Inui |
ACL/IJCNLP (1) | 6 |
| 2021 | Two Training Strategies for Improving Relation Extraction over Universal GraphabstractThis paper explores how the Distantly Supervised Relation Extraction (DS-RE) can benefit from the use of a Universal Graph (UG), the combination of a Knowledge Graph (KG) and a large-scale text collection.A straightforward extension of a current state-of-the-art neural model for DS-RE with a UG may lead to degradation in performance.We first report that this degradation is associated with the difficulty in learning a UG and then propose two training strategies: (1) Path Type Adaptive Pretraining, which sequentially trains the model with different types of UG paths so as to prevent the reliance on a single type of UG path; and (2) Complexity Ranking Guided Attention mechanism, which restricts the attention span according to the complexity of a UG path so as to force the model to extract features not only from simple UG paths but also from complex ones.Experimental results on both biomedical and NYT10 datasets prove the robustness of our methods and achieve a new state-ofthe-art result on the NYT10 dataset.The code and datasets used in this paper are available at https://github.com/baodaiqin/ UGDSRE. Qin Dai, Naoya Inoue, Kentaro Inui |
EACL | 4 |
| 2021 | Language Models as Knowledge Bases: On Entity Representations, Storage Capacity, and Paraphrased QueriesabstractPretrained language models have been suggested as a possible alternative or complement to structured knowledge bases.However, this emerging LM-as-KB paradigm has so far only been considered in a very limited setting, which only allows handling 21k entities whose name is found in common LM vocabularies.Furthermore, a major benefit of this paradigm, i.e., querying the KB using natural language paraphrases, is underexplored.Here we formulate two basic requirements for treating LMs as KBs: (i) the ability to store a large number facts involving a large number of entities and (ii) the ability to query stored facts.We explore three entity representations that allow LMs to handle millions of entities and present a detailed case study on paraphrased querying of facts stored in LMs, thereby providing a proof-of-concept that language models can indeed serve as knowledge bases. Benjamin Heinzerling, Kentaro Inui |
EACL | 2 |
| 2021 | Exploring Transitivity in Neural NLI Models through VeridicalityabstractDespite the recent success of deep neural networks in natural language processing, the extent to which they can demonstrate human-like generalization capacities for natural language understanding remains unclear.We explore this issue in the domain of natural language inference (NLI), focusing on the transitivity of inference relations, a fundamental property for systematically drawing inferences.A model capturing transitivity can compose basic inference patterns and draw new inferences.We introduce an analysis method using synthetic and naturalistic NLI datasets involving clauseembedding verbs to evaluate whether models can perform transitivity inferences composed of veridical inferences and arbitrary inference types.We find that current NLI models do not perform consistently well on transitivity inference tasks, suggesting that they lack the generalization capacity for drawing composite inferences from provided training examples.The data and code for our analysis are publicly available at https://github.com/ verypluming/transitivity. Hitomi Yanaka, Koji Mineshima, Kentaro Inui |
EACL | 3 |
| 2021 | Exploring Methods for Generating Feedback Comments for Writing LearningabstractThe task of generating explanatory notes for language learners is known as feedback comment generation.Although various generation techniques are available, little is known about which methods are appropriate for this task.Nagata (2019) demonstrates the effectiveness of neural-retrieval-based methods in generating feedback comments for preposition use.Retrieval-based methods have limitations in that they can only output feedback comments existing in a given training data.Furthermore, feedback comments can be made on other grammatical and writing items than preposition use, which is still unaddressed.To shed light on these points, we investigate a wider range of methods for generating many feedback comments in this study.Our close analysis of the type of task leads us to investigate three different architectures for comment generation: (i) a neural-retrieval-based method as a baseline, (ii) a pointer-generator-based generation method as a neural seq2seq method, (iii) a retrieve-and-edit method, a hybrid of (i) and (ii).Intuitively, the pointer-generator should outperform neural-retrieval, and retrieve-andedit should perform best.However, in our experiments, this expectation is completely overturned.We closely analyze the results to reveal the major causes of these counter-intuitive results and report on our findings from the experiments.1 Kazuaki Hanawa, Ryo Nagata, Kentaro Inui |
EMNLP (1) | 3 |
| 2021 | Summarize-then-Answer: Generating Concise Explanations for Multi-hop Reading ComprehensionabstractHow can we generate concise explanations for multi-hop Reading Comprehension (RC)?The current strategies of identifying supporting sentences can be seen as an extractive questionfocused summarization of the input text.However, these extractive explanations are not necessarily concise i.e. not minimally sufficient for answering a question.Instead, we advocate for an abstractive approach, where we propose to generate a question-focused, abstractive summary of input paragraphs and then feed it to an RC system.Given a limited amount of human-annotated abstractive explanations, we train the abstractive explainer in a semi-supervised manner, where we start from the supervised model and then train it further through trial and error maximizing a conciseness-promoted reward function.Our experiments demonstrate that the proposed abstractive explainer can generate more compact explanations than an extractive explainer with limited supervision (only 2k instances) while maintaining sufficiency.1 Our implementation is publicly available at https:// github.com/StonyBrookNLP/suqa.Charlie Rowe plays Billy Costa in a film based on what novel?[P1] [1] The Golden Compass is a 2007 British-American fantasy adventure film based on "Northern Lights", the first novel in Philip Pullman's trilogy "His Dark Materials". Naoya Inoue, Harsh Trivedi, Steven Sinha, Niranjan Balasubramanian, Kentaro Inui |
EMNLP (1) | 5 |
| 2021 | SHAPE : Shifted Absolute Position Embedding for TransformersabstractPosition representation is crucial for building position-aware representations in Transformers.Existing position representations suffer from a lack of generalization to test data with unseen lengths or high computational cost.We investigate shifted absolute position embedding (SHAPE) to address both issues.The basic idea of SHAPE is to achieve shift invariance, which is a key property of recent successful position representations, by randomly shifting absolute positions during training.We demonstrate that SHAPE is empirically comparable to its counterpart while being simpler and faster 1 . Shun Kiyono, Sosuke Kobayashi, Jun Suzuki 0001, Kentaro Inui |
EMNLP (1) | 4 |
| 2021 | Incorporating Residual and Normalization Layers into Analysis of Masked Language ModelsabstractTransformer architecture has become ubiquitous in the natural language processing field.To interpret the Transformer-based models, their attention patterns have been extensively analyzed.However, the Transformer architecture is not only composed of the multihead attention; other components can also contribute to Transformers' progressive performance.In this study, we extended the scope of the analysis of Transformers from solely the attention patterns to the whole attention block, i.e., multi-head attention, residual connection, and layer normalization.Our analysis of Transformer-based masked language models shows that the token-to-token interaction performed via attention has less impact on the intermediate representations than previously assumed.These results provide new intuitive explanations of existing reports; for example, discarding the learned attention patterns tends not to adversely affect the performance.The codes of our experiments are publicly available. Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro Inui |
EMNLP (1) | 4 |
| 2021 | Pseudo Zero Pronoun Resolution Improves Zero Anaphora ResolutionabstractMasked language models (MLMs) have contributed to drastic performance improvements with regard to zero anaphora resolution (ZAR).To further improve this approach, in this study, we made two proposals.The first is a new pretraining task that trains MLMs on anaphoric relations with explicit supervision, and the second proposal is a new finetuning method that remedies a notorious issue, the pretrainfinetune discrepancy.Our experiments on Japanese ZAR demonstrated that our two proposals boost the state-of-the-art performance, and our detailed analysis provides new insights on the remaining challenges. Ryuto Konno, Shun Kiyono, Yuichiroh Matsubayashi, Hiroki Ouchi, Kentaro Inui |
EMNLP (1) | 5 |
| 2021 | Transformer-based Lexically Constrained Headline GenerationabstractThis paper explores a variant of automatic headline generation methods, where a generated headline is required to include a given phrase such as a company or a product name.Previous methods using Transformer-based models generate a headline including a given phrase by providing the encoder with additional information corresponding to the given phrase.However, these methods cannot always include the phrase in the generated headline.Inspired by previous RNN-based methods generating token sequences in backward and forward directions from the given phrase, we propose a simple Transformerbased method that guarantees to include the given phrase in the high-quality generated headline.We also consider a new headline generation strategy that takes advantage of the controllable generation order of Transformer.Our experiments with the Japanese News Corpus demonstrate that our methods, which are guaranteed to include the phrase in the generated headline, achieve ROUGE scores comparable to previous Transformer-based methods.We also show that our generation strategy performs better than previous strategies. Kosuke Yamada, Yuta Hitomi, Hideaki Tamori, Ryohei Sasano, Naoaki Okazaki, Kentaro Inui, Koichi Takeda 0003 |
EMNLP (1) | 6 |
| 2021 | Evaluation of Similarity-based Explanations
Kazuaki Hanawa, Sho Yokoi, Satoshi Hara 0001, Kentaro Inui |
ICLR | 4 |
| 2021 | Learning to Learn to be Right for the Right ReasonsabstractPride Kavumba, Benjamin Heinzerling, Ana Brassard, Kentaro Inui. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Pride Kavumba, Benjamin Heinzerling, Ana Brassard, Kentaro Inui |
NAACL-HLT | 4 |
| 2021 | Instance-Based Neural Dependency ParsingabstractAbstract Interpretable rationales for model predictions are crucial in practical applications. We develop neural models that possess an interpretable inference process for dependency parsing. Our models adopt instance-based inference, where dependency edges are extracted and labeled by comparing them to edges in a training set. The training edges are explicitly used for the predictions; thus, it is easy to grasp the contribution of each edge to the predictions. Our experiments show that our instance-based models achieve competitive accuracy with standard neural models and have the reasonable plausibility of instance-based explanations. Hiroki Ouchi, Jun Suzuki 0001, Sosuke Kobayashi, Sho Yokoi, Tatsuki Kuribayashi, Masashi Yoshikawa, Kentaro Inui |
Trans. Assoc. Comput. Linguistics | 7 |
| 2021 | Corruption Is Not All Bad: Incorporating Discourse Structure Into Pre-Training via Corruption for Essay ScoringabstractExisting approaches for automated essay scoring and document representation learning typically rely on discourse parsers to incorporate discourse structure into text representation. However, the performance of parsers is not always adequate, especially when they are used on noisy texts, such as student essays. In this paper, we propose an unsupervised pre-training approach to capture discourse structure of essays in terms of coherence and cohesion that does not require any discourse parser or annotation. We introduce several types of token, sentence and paragraph-level corruption techniques for our proposed pre-training approach and augment masked language modeling pre-training with our pre-training method to leverage both contextualized and discourse information. Our proposed unsupervised approach achieves a new state-of-the-art result on the task of essay Organization scoring. Farjana Sultana Mim, Naoya Inoue, Paul Reisert, Hiroki Ouchi, Kentaro Inui |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Subword-Based Compact Reconstruction for Open-Vocabulary Neural Word EmbeddingsabstractThe methodology of neural word embeddings has become an important fundamental resource for tackling many applications in the artificial intelligence (AI) research field. They have successfully been proven to capture high-quality syntactic and semantic relationships in a vector space. Despite their significant impact, neural word embeddings have several disadvantages. In this paper, we focus on two issues regarding well-trained word embeddings: (i) the massive memory requirement and (ii) the inapplicability of out-of-vocabulary (OOV) words. To overcome these two issues, we propose a method of reconstructing pre-trained word embeddings by using subword information that can effectively represent a large number of subword embeddings in a considerably small fixed space while preventing quality degradation from the original word embeddings. The key techniques of our method are twofold: memory-shared embeddings and a variant of the key-value-query self-attention mechanism. Our experiments show that our reconstructed subword-based word embeddings can successfully imitate well-trained word embeddings in a small fixed space while preventing quality degradation across several linguistic benchmark datasets and can simultaneously predict effective embeddings of OOV words. We also demonstrate the effectiveness of our reconstruction method when it is applied to downstream tasks, such as named entity recognition and natural language inference tasks. Shota Sasaki, Jun Suzuki 0001, Kentaro Inui |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Assessing the Benchmarking Capacity of Machine Reading Comprehension DatasetsabstractExisting analysis work in machine reading comprehension (MRC) is largely concerned with evaluating the capabilities of systems. However, the capabilities of datasets are not assessed for benchmarking language understanding precisely. We propose a semi-automated, ablation-based methodology for this challenge; By checking whether questions can be solved even after removing features associated with a skill requisite for language understanding, we evaluate to what degree the questions do not require the skill. Experiments on 10 datasets (e.g., CoQA, SQuAD v2.0, and RACE) with a strong baseline model show that, for example, the relative scores of the baseline model provided with content words only and with shuffled sentence words in the context are on average 89.2% and 78.5% of the original scores, respectively. These results suggest that most of the questions already answered correctly by the model do not necessarily require grammatical and complex reasoning. For precise benchmarking, MRC datasets will need to take extra care in their design to ensure that questions can correctly evaluate the intended skills. Saku Sugawara, Pontus Stenetorp, Kentaro Inui, Akiko Aizawa |
AAAI | 3 |
| 2020 | R4C: A Benchmark for Evaluating RC Systems to Get the Right Answer for the Right ReasonabstractRecent studies have revealed that reading comprehension (RC) systems learn to exploit annotation artifacts and other biases in current datasets.This prevents the community from reliably measuring the progress of RC systems.To address this issue, we introduce R 4 C, a new task for evaluating RC systems' internal reasoning.R 4 C requires giving not only answers but also derivations: explanations that justify predicted answers.We present a reliable, crowdsourced framework for scalably annotating RC datasets with derivations.We create and publicly release the R 4 C dataset, the first, quality-assured dataset consisting of 4.6k questions, each of which is annotated with 3 reference derivations (i.e.13.8k derivations).Experiments show that our automatic evaluation metrics using multiple reference derivations are reliable, and that R 4 C assesses different skills from an existing benchmark. Naoya Inoue, Pontus Stenetorp, Kentaro Inui |
ACL | 3 |
| 2020 | Encoder-Decoder Models Can Benefit from Pre-trained Masked Language Models in Grammatical Error CorrectionabstractThis paper investigates how to effectively incorporate a pre-trained masked language model (MLM), such as BERT, into an encoderdecoder (EncDec) model for grammatical error correction (GEC).The answer to this question is not as straightforward as one might expect because the previous common methods for incorporating a MLM into an EncDec model have potential drawbacks when applied to GEC.For example, the distribution of the inputs to a GEC model can be considerably different (erroneous, clumsy, etc.) from that of the corpora used for pre-training MLMs; however, this issue is not addressed in the previous methods.Our experiments show that our proposed method, where we first fine-tune a MLM with a given GEC corpus and then use the output of the finetuned MLM as additional features in the GEC model, maximizes the benefit of the MLM.The best-performing model achieves state-ofthe-art performances on the BEA-2019 and CoNLL-2014 benchmarks.Our code is publicly available at: https://github.com/ kanekomasahiro/bert-gec. Masahiro Kaneko, Masato Mita, Shun Kiyono, Jun Suzuki 0001, Kentaro Inui |
ACL | 5 |
| 2020 | Language Models as an Alternative Evaluator of Word Order Hypotheses: A Case Study in JapaneseabstractWe examine a methodology using neural language models (LMs) for analyzing the word order of language.This LM-based method has the potential to overcome the difficulties existing methods face, such as the propagation of preprocessor errors in count-based methods.In this study, we explore whether the LMbased method is valid for analyzing the word order.As a case study, this study focuses on Japanese due to its complex and flexible word order.To validate the LM-based method, we test (i) parallels between LMs and human word order preference, and (ii) consistency of the results obtained using the LM-based method with previous linguistic studies.Through our experiments, we tentatively conclude that LMs display sufficient word order knowledge for usage as an analysis tool.Finally, using the LMbased method, we demonstrate the relationship between the canonical word order and topicalization, which had yet to be analyzed by largescale experiments.C Data used in Section 5.2, Section 6, and Appendix F Tatsuki Kuribayashi, Takumi Ito, Jun Suzuki 0001, Kentaro Inui |
ACL | 4 |
| 2020 | Instance-Based Learning of Span Representations: A Case Study through Named Entity RecognitionabstractInterpretable rationales for model predictions play a critical role in practical applications.In this study, we develop models possessing interpretable inference process for structured prediction.Specifically, we present a method of instance-based learning that learns similarities between spans.At inference time, each span is assigned a class label based on its similar spans in the training set, where it is easy to understand how much each training instance contributes to the predictions.Through empirical analysis on named entity recognition, we demonstrate that our method enables to build models that have high interpretability without sacrificing performance. Hiroki Ouchi, Jun Suzuki 0001, Sosuke Kobayashi, Sho Yokoi, Tatsuki Kuribayashi, Ryuto Konno, Kentaro Inui |
ACL | 7 |
| 2020 | Evaluating Dialogue Generation Systems via Response SelectionabstractExisting automatic evaluation metrics for open-domain dialogue response generation systems correlate poorly with human evaluation.We focus on evaluating response generation systems via response selection.To evaluate systems properly via response selection, we propose a method to construct response selection test sets with well-chosen false candidates.Specifically, we propose to construct test sets filtering out some types of false candidates: (i) those unrelated to the ground-truth response and (ii) those acceptable as appropriate responses.Through experiments, we demonstrate that evaluating systems via response selection with the test set developed by our method correlates more strongly with human evaluation, compared with widely used automatic evaluation metrics such as BLEU. Shiki Sato, Reina Akama, Hiroki Ouchi, Jun Suzuki 0001, Kentaro Inui |
ACL | 5 |
| 2020 | Do Neural Models Learn Systematicity of Monotonicity Inference in Natural Language?abstractDespite the success of language models using neural networks, it remains unclear to what extent neural models have the generalization ability to perform inferences.In this paper, we introduce a method for evaluating whether neural models can learn systematicity of monotonicity inference in natural language, namely, the regularity for performing arbitrary inferences with generalization on composition.We consider four aspects of monotonicity inferences and test whether the models can systematically interpret lexical and logical phenomena on different training/test splits.A series of experiments show that three neural models systematically draw inferences on unseen combinations of lexical and logical phenomena when the syntactic structures of the sentences are similar between the training and test sets.However, the performance of the models significantly decreases when the structures are slightly changed in the test set while retaining all vocabularies and constituents already appearing in the training set.This indicates that the generalization ability of neural models is limited to cases where the syntactic structures are nearly the same as those in the training set.(1) P : Some [puppies ↑] ran.H: Some dogs ran.(2) P : No [cats ↓] ran.H: No small cats ran.(3) P : Some [puppies which chased no [cats ↓]] ran.H: Some dogs which chased no small cats ran.(5) P : Some small dogs ran ⇒ H: Some dogs ran (6) P : Several dogs ran ⇒ H: Several animals ran (7) P : No animals ran ⇒ H: No dogs ran (8) P : Several small dogs ran ⇒ H: Several dogs ran (9) P : No dogs ran ⇒ H: No small dogs ranHere, we consider a set of inferences D Q,R Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui |
ACL | 4 |
| 2020 | PheMT: A Phenomenon-wise Dataset for Machine Translation Robustness on User-Generated ContentsabstractNeural Machine Translation (NMT) has shown drastic improvement in its quality when translating clean input, such as text from the news domain.However, existing studies suggest that NMT still struggles with certain kinds of input with considerable noise, such as User-Generated Contents (UGC) on the Internet.To make better use of NMT for cross-cultural communication, one of the most promising directions is to develop a model that correctly handles these expressions.Though its importance has been recognized, it is still not clear as to what creates the great gap in performance between the translation of clean input and that of UGC.To answer the question, we present a new dataset, PheMT, for evaluating the robustness of MT systems against specific linguistic phenomena in Japanese-English translation.Our experiments with the created dataset revealed that not only our in-house models but even widely used off-the-shelf systems are greatly disturbed by the presence of certain phenomena. Ryo Fujii, Masato Mita, Kaori Abe, Kazuaki Hanawa, Makoto Morishita, Jun Suzuki 0001, Kentaro Inui |
COLING | 7 |
| 2020 | An Empirical Study of Contextual Data Augmentation for Japanese Zero Anaphora ResolutionabstractOne critical issue of zero anaphora resolution (ZAR) is the scarcity of labeled data.This study explores how effectively this problem can be alleviated by data augmentation.We adopt a state-ofthe-art data augmentation method, called the contextual data augmentation (CDA), that generates labeled training instances using a pretrained language model.The CDA has been reported to work well for several other natural language processing tasks, including text classification and machine translation (Kobayashi, 2018;Wu et al., 2019;Gao et al., 2019).This study addresses two underexplored issues on CDA, that is, how to reduce the computational cost of data augmentation and how to ensure the quality of the generated data.We also propose two methods to adapt CDA to ZAR: [MASK]-based augmentation and linguistically-controlled masking.Consequently, the experimental results on Japanese ZAR show that our methods contribute to both the accuracy gain and the computation cost reduction.Our closer analysis reveals that the proposed method can improve the quality of the augmented training data when compared to the conventional CDA. Ryuto Konno, Yuichiroh Matsubayashi, Shun Kiyono, Hiroki Ouchi, Kentaro Inui |
COLING | 6 |
| 2020 | Modeling Event Salience in Narratives via Barthes' Cardinal FunctionsabstractEvents in a narrative differ in salience: some are more important to the story than others.Estimating event salience is useful for tasks such as story generation, and as a tool for text analysis in narratology and folkloristics.To compute event salience without any annotations, we adopt Barthes' definition of event salience and propose several unsupervised methods that require only a pre-trained language model.Evaluating the proposed methods on folktales with event salience annotation, we show that the proposed methods outperform baseline methods and find fine-tuning a language model on narrative texts is a key factor in improving the proposed methods. Takaki Otake, Sho Yokoi, Naoya Inoue, Tatsuki Kuribayashi, Kentaro Inui |
COLING | 6 |
| 2020 | Filtering Noisy Dialogue Corpora by Connectivity and Content RelatednessabstractLarge-scale dialogue datasets have recently become available for training neural dialogue agents. However, these datasets have been reported to contain a non-negligible number of unacceptable utterance pairs. In this paper, we propose a method for scoring the quality of utterance pairs in terms of their connectivity and relatedness. The proposed scoring method is designed based on findings widely shared in the dialogue and linguistics research communities. We demonstrate that it has a relatively good correlation with the human judgment of dialogue quality. Furthermore, the method is applied to filter out potentially unacceptable utterance pairs from a large-scale noisy dialogue corpus to ensure its quality. We experimentally confirm that training data filtered by the proposed method improves the quality of neural dialogue agents in response generation. Reina Akama, Sho Yokoi, Jun Suzuki 0001, Kentaro Inui |
EMNLP (1) | 4 |
| 2020 | Attention is Not Only a Weight: Analyzing Transformers with Vector NormsabstractAttention is a key component of Transformers, which have recently achieved considerable success in natural language processing. Hence, attention is being extensively studied to investigate various linguistic capabilities of Transformers, focusing on analyzing the parallels between attention weights and specific linguistic phenomena. This paper shows that attention weights alone are only one of the two factors that determine the output of attention and proposes a norm-based analysis that incorporates the second factor, the norm of the transformed input vectors. The findings of our norm-based analyses of BERT and a Transformer-based neural machine translation system include the following: (i) contrary to previous studies, BERT pays poor attention to special tokens, and (ii) reasonable word alignment can be extracted from attention mechanisms of Transformer. These findings provide insights into the inner workings of Transformers. Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro Inui |
EMNLP (1) | 4 |
| 2020 | Word Rotator's DistanceabstractA key principle in assessing textual similarity is measuring the degree of semantic overlap between two texts by considering the word alignment.Such alignment-based approaches are intuitive and interpretable; however, they are empirically inferior to the simple cosine similarity between general-purpose sentence vectors.To address this issue, we focus on and demonstrate the fact that the norm of word vectors is a good proxy for word importance, and their angle is a good proxy for word similarity.Alignment-based approaches do not distinguish them, whereas sentence-vector approaches automatically use the norm as the word importance.Accordingly, we propose a method that first decouples word vectors into their norm and direction, and then computes alignment-based similarity using earth mover's distance (i.e., optimal transport cost), which we refer to as word rotator's distance.Besides, we find how to "grow" the norm and direction of word vectors (vector converter), which is a new systematic approach derived from sentence-vector estimation methods.On several textual similarity datasets, the combination of these simple proposed methods outperformed not only alignment-based approaches but also strong baselines.1 Sho Yokoi, Reina Akama, Jun Suzuki 0001, Kentaro Inui |
EMNLP (1) | 5 |
| 2020 | Creating Corpora for Research in Feedback Comment GenerationabstractIn this paper, we report on datasets that we created for research in feedback comment generation — a task of automatically generating feedback comments such as a hint or an explanatory note for writing learning. There has been almost no such corpus open to the public and accordingly there has been a very limited amount of work on this task. In this paper, we first discuss the principle and guidelines for feedback comment annotation. Then, we describe two corpora that we have manually annotated with feedback comments (approximately 50,000 general comments and 6,700 on preposition use). A part of the annotation results is now available on the web, which will facilitate research in feedback comment generation Ryo Nagata, Kentaro Inui, Shin'ichiro Ishikawa |
LREC | 2 |
| 2020 | Massive Exploration of Pseudo Data for Grammatical Error CorrectionabstractCollecting a large amount of training data for grammatical error correction (GEC) models has been an ongoing challenge in the field of GEC. Recently, it has become common to use data demanding deep neural models such as an encoder-decoder for GEC; thus, tackling the problem of data collection has become increasingly important. The incorporation of pseudo data in the training of GEC models is one of the main approaches for mitigating the problem of data scarcity. However, a consensus is lacking on experimental configurations, namely, (i) the methods for generating pseudo data, (ii) the seed corpora used as the source of the pseudo data, and (iii) the means of optimizing the model. In this study, these configurations are thoroughly explored through massive amount of experiments, with the aim of providing an improved understanding of pseudo data. Our main experimental finding is that pretraining a model with pseudo data generated by back-translation-based method is the most effective approach. Our findings are supported by the achievement of state-of-the-art performance on multiple benchmark test sets (the CoNLL-2014 test set and the official test set of the BEA-2019 shared task) without requiring any modifications to the model architecture. We also perform an in-depth analysis of our model with respect to the grammatical error type and proficiency level of the text. Finally, we suggest future directions for further improving model performance. Shun Kiyono, Jun Suzuki 0001, Tomoya Mizumoto, Kentaro Inui |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Mixture of Expert/Imitator Networks: Scalable Semi-Supervised Learning FrameworkabstractThe current success of deep neural networks (DNNs) in an increasingly broad range of tasks involving artificial intelligence strongly depends on the quality and quantity of labeled training data. In general, the scarcity of labeled data, which is often observed in many natural language processing tasks, is one of the most important issues to be addressed. Semisupervised learning (SSL) is a promising approach to overcoming this issue by incorporating a large amount of unlabeled data. In this paper, we propose a novel scalable method of SSL for text classification tasks. The unique property of our method, Mixture of Expert/Imitator Networks, is that imitator networks learn to “imitate” the estimated label distribution of the expert network over the unlabeled data, which potentially contributes a set of features for the classification. Our experiments demonstrate that the proposed method consistently improves the performance of several types of baseline DNNs. We also demonstrate that our method has the more data, better performance property with promising scalability to the amount of unlabeled data. Shun Kiyono, Jun Suzuki 0001, Kentaro Inui |
AAAI | 3 |
| 2019 | An Empirical Study of Span Representations in Argumentation Structure ParsingabstractTatsuki Kuribayashi, Hiroki Ouchi, Naoya Inoue, Paul Reisert, Toshinori Miyoshi, Jun Suzuki, Kentaro Inui. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Tatsuki Kuribayashi, Hiroki Ouchi, Naoya Inoue, Paul Reisert, Toshinori Miyoshi, Jun Suzuki 0001, Kentaro Inui |
ACL (1) | 7 |
| 2019 | An Annotation Protocol for Collecting User-Generated Counter-Arguments Using Crowdsourcing
Paul Reisert, Gisela Vallejo, Naoya Inoue, Iryna Gurevych, Kentaro Inui |
AIED (2) | 5 |
| 2019 | An Empirical Study of Incorporating Pseudo Data into Grammatical Error CorrectionabstractShun Kiyono, Jun Suzuki, Masato Mita, Tomoya Mizumoto, Kentaro Inui. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Shun Kiyono, Jun Suzuki 0001, Masato Mita, Tomoya Mizumoto, Kentaro Inui |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Transductive Learning of Neural Language Models for Syntactic and Semantic AnalysisabstractHiroki Ouchi, Jun Suzuki, Kentaro Inui. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Hiroki Ouchi, Jun Suzuki 0001, Kentaro Inui |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Select and Attend: Towards Controllable Content Selection in Text GenerationabstractXiaoyu Shen, Jun Suzuki, Kentaro Inui, Hui Su, Dietrich Klakow, Satoshi Sekine. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiaoyu Shen 0001, Jun Suzuki 0001, Kentaro Inui, Hui Su, Dietrich Klakow, Satoshi Sekine |
EMNLP/IJCNLP (1) | 3 |
| 2019 | A Large-Scale Multi-Length Headline Corpus for Analyzing Length-Constrained Headline Generation Model EvaluationabstractYuta Hitomi, Yuya Taguchi, Hideaki Tamori, Ko Kikuta, Jiro Nishitoba, Naoaki Okazaki, Kentaro Inui, Manabu Okumura. Proceedings of the 12th International Conference on Natural Language Generation. 2019. Yuta Hitomi, Yuya Taguchi, Hideaki Tamori, Ko Kikuta, Jiro Nishitoba, Naoaki Okazaki, Kentaro Inui, Manabu Okumura |
INLG | 7 |
| 2019 | Diamonds in the Rough: Generating Fluent Sentences from Early-Stage Drafts for Academic Writing AssistanceabstractThe writing process consists of several stages such as drafting, revising, editing, and proofreading. Studies on writing assistance, such as grammatical error correction (GEC), have mainly focused on sentence editing and proofreading, where surface-level issues such as typographical errors, spelling errors, or grammatical errors should be corrected. We broaden this focus to include the earlier revising stage, where sentences require adjustment to the information included or major rewriting and propose Sentence-level Revision (SentRev) as a new writing assistance task. Well-performing systems in this task can help inexperienced authors by producing fluent, complete sentences given their rough, incomplete drafts. We build a new freely available crowdsourced evaluation dataset consisting of incomplete sentences authored by non-native writers paired with their final versions extracted from published academic papers for developing and evaluating SentRev models. We also establish baseline performance on SentRev using our newly built evaluation dataset. Takumi Ito, Tatsuki Kuribayashi, Hayato Kobayashi, Ana Brassard, Masato Hagiwara, Jun Suzuki 0001, Kentaro Inui |
INLG | 7 |
| 2018 | Interpretable and Compositional Relation Learning by Joint Training with an AutoencoderabstractEmbedding models for entities and relations are extremely useful for recovering missing facts in a knowledge base.Intuitively, a relation can be modeled by a matrix mapping entity vectors.However, relations reside on low dimension sub-manifolds in the parameter space of arbitrary matrices -for one reason, composition of two relations M 1 , M 2 may match a third M 3 (e.g.composition of relations currency of country and country of film usually matches currency of film budget), which imposes compositional constraints to be satisfied by the parameters (i.e.M 1 •M 2 ≈ M 3 ).In this paper we investigate a dimension reduction technique by training relations jointly with an autoencoder, which is expected to better capture compositional constraints.We achieve state-of-the-art on Knowledge Base Completion tasks with strongly improved Mean Rank, and show that joint training with an autoencoder leads to interpretable sparse codings of relations, helps discovering compositional constraints and benefits from compositional training.Our source code is released at github.com/tianran/glimvec. Kentaro Inui |
ACL (1) | 3 |
| 2018 | Distance-Free Modeling of Multi-Predicate Interactions in End-to-End Japanese Predicate-Argument Structure AnalysisabstractCapturing interactions among multiple predicate-argument structures (PASs) is a crucial issue in the task of analyzing PAS in Japanese. In this paper, we propose new Japanese PAS analysis models that integrate the label prediction information of arguments in multiple PASs by extending the input and last layers of a standard deep bidirectional recurrent neural network (bi-RNN) model. In these models, using the mechanisms of pooling and attention, we aim to directly capture the potential interactions among multiple PASs, without being disturbed by the word order and distance. Our experiments show that the proposed models improve the prediction accuracy specifically for cases where the predicate and argument are in an indirect dependency relation and achieve a new state of the art in the overall F_1 on a standard benchmark corpus. Yuichiroh Matsubayashi, Kentaro Inui |
COLING | 2 |
| 2018 | Predicting Stances from Social Media Posts using Factorization MachinesabstractSocial media provide platforms to express, discuss, and shape opinions about events and issues in the real world. An important step to analyze the discussions on social media and to assist in healthy decision-making is stance detection. This paper presents an approach to detect the stance of a user toward a topic based on their stances toward other topics and the social media posts of the user. We apply factorization machines, a widely used method in item recommendation, to model user preferences toward topics from the social media data. The experimental results demonstrate that users’ posts are useful to model topic preferences and therefore predict stances of silent users. Akira Sasaki, Kazuaki Hanawa, Naoaki Okazaki, Kentaro Inui |
COLING | 4 |
| 2018 | What Makes Reading Comprehension Questions Easier?abstractA challenge in creating a dataset for machine reading comprehension (MRC) is to collect questions that require a sophisticated understanding of language to answer beyond using superficial cues.In this work, we investigate what makes questions easier across recent 12 MRC datasets with three question styles (answer extraction, description, and multiple choice).We propose to employ simple heuristics to split each dataset into easy and hard subsets and examine the performance of two baseline models for each of the subsets.We then manually annotate questions sampled from each subset with both validity and requisite reasoning skills to investigate which skills explain the difference between easy and hard questions.From this study, we observed that (i) the baseline performances for the hard subsets remarkably degrade compared to those of entire datasets, (ii) hard questions require knowledge inference and multiple-sentence reasoning in comparison with easy questions, and (iii) multiplechoice questions tend to require a broader range of reasoning skills than answer extraction and description questions.These results suggest that one might overestimate recent advances in MRC. Saku Sugawara, Kentaro Inui, Satoshi Sekine, Akiko Aizawa |
EMNLP | 2 |
| 2018 | Pointwise HSIC: A Linear-Time Kernelized Co-occurrence Norm for Sparse Linguistic ExpressionsabstractIn this paper, we propose a new kernel-based co-occurrence measure that can be applied to sparse linguistic expressions (e.g., sentences) with a very short learning time, as an alternative to pointwise mutual information (PMI).As well as deriving PMI from mutual information, we derive this new measure from the Hilbert-Schmidt independence criterion (HSIC); thus, we call the new measure the pointwise HSIC (PHSIC).PHSIC can be interpreted as a smoothed variant of PMI that allows various similarity metrics (e.g., sentence embeddings) to be plugged in as kernels.Moreover, PHSIC can be estimated by simple and fast (linear in the size of the data) matrix calculations regardless of whether we use linear or nonlinear kernels.Empirically, in a dialogue response selection task, PHSIC is learned thousands of times faster than an RNNbased PMI while outperforming PMI in accuracy.In addition, we also demonstrate that PH-SIC is beneficial as a criterion of a data selection task for machine translation owing to its ability to give high (low) scores to a consistent (inconsistent) pair with other pairs. Sho Yokoi, Sosuke Kobayashi, Kenji Fukumizu, Jun Suzuki 0001, Kentaro Inui |
EMNLP | 5 |
| 2018 | A Melody-Conditioned Lyrics Language ModelabstractKento Watanabe, Yuichiroh Matsubayashi, Satoru Fukayama, Masataka Goto, Kentaro Inui, Tomoyasu Nakano. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Kento Watanabe, Yuichiroh Matsubayashi, Satoru Fukayama, Masataka Goto, Kentaro Inui, Tomoyasu Nakano |
NAACL-HLT | 5 |
| 2018 | Multi-dialect Neural Machine Translation and Dialectometry
Kaori Abe, Yuichiroh Matsubayashi, Naoaki Okazaki, Kentaro Inui |
PACLIC | 4 |
| 2018 | Improving Scientific Relation Classification with Task Specific Supersense
Qin Dai, Naoya Inoue, Paul Reisert, Kentaro Inui |
PACLIC | 4 |
| 2018 | Reducing Odd Generation from Neural Headline Generation
Shun Kiyono, Sho Takase, Jun Suzuki 0001, Naoaki Okazaki, Kentaro Inui, Masaaki Nagata |
PACLIC | 5 |
| 2018 | Suspicious News Detection Using Micro Blog Text
Tsubasa Tagami, Hiroki Ouchi, Hiroki Asano, Kazuaki Hanawa, Kaori Uchiyama, Kaito Suzuki, Kentaro Inui, Atsushi Komiya, Atsuo Fujimura, Ryo Yamashita, Hitofumi Yanai, Akinori Machino |
PACLIC | 7 |
| 2017 | Other Topics You May Also Agree or Disagree: Modeling Inter-Topic Preferences using Tweets and Matrix FactorizationabstractWe present in this paper our approach for modeling inter-topic preferences of Twitter users: for example, those who agree with the Trans-Pacific Partnership (TPP) also agree with free trade.This kind of knowledge is useful not only for stance detection across multiple topics but also for various real-world applications including public opinion surveys, electoral predictions, electoral campaigns, and online debates.In order to extract users' preferences on Twitter, we design linguistic patterns in which people agree and disagree about specific topics (e.g., "A is completely wrong").By applying these linguistic patterns to a collection of tweets, we extract statements agreeing and disagreeing with various topics.Inspired by previous work on item recommendation, we formalize the task of modeling intertopic preferences as matrix factorization: representing users' preferences as a usertopic matrix and mapping both users and topics onto a latent feature space that abstracts the preferences.Our experimental results demonstrate both that our proposed approach is useful in predicting missing preferences of users and that the latent vector representations of topics successfully encode inter-topic preferences. Akira Sasaki, Kazuaki Hanawa, Naoaki Okazaki, Kentaro Inui |
ACL (1) | 4 |
| 2017 | Monitoring Geographical Entities with Temporal Awareness in Tweets
Koji Matsuda, Mizuki Sango, Naoaki Okazaki, Kentaro Inui |
CICLing (2) | 4 |
| 2017 | Neural Architectures for Fine-grained Entity Type ClassificationabstractSonse Shimaoka, Pontus Stenetorp, Kentaro Inui, Sebastian Riedel. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Sonse Shimaoka, Pontus Stenetorp, Kentaro Inui, Sebastian Riedel 0001 |
EACL (1) | 3 |
| 2017 | Learning Co-Substructures by Kernel Dependence MaximizationabstractModeling associations between items in a dataset is a problem that is frequently encountered in data and knowledge mining research. Most previous studies have simply applied a predefined fixed pattern for extracting the substructure of each item pair and then analyzed the associations between these substructures. Using such fixed patterns may not, however, capture the significant association. We, therefore, propose the novel machine learning task of extracting a strongly associated substructure pair (co-substructure) from each input item pair. We call this task dependent co-substructure extraction (DCSE), and formalize it as a dependence maximization problem. Then, we discuss critical issues with this task: the data sparsity problem and a huge search space. To address the data sparsity problem, we adopt the Hilbert--Schmidt independence criterion as an objective function. To improve search efficiency, we adopt the Metropolis--Hastings algorithm. We report the results of empirical evaluations, in which the proposed method is applied for acquiring and predicting narrative event pairs, an active task in the field of natural language processing. Sho Yokoi, Daichi Mochihashi, Naoaki Okazaki, Kentaro Inui |
IJCAI | 5 |
| 2017 | A Neural Language Model for Dynamically Representing the Meanings of Unknown Words and Entities in a DiscourseabstractThis study addresses the problem of identifying the meaning of unknown words or entities in a discourse with respect to the word embedding approaches used in neural language models. We proposed a method for on-the-fly construction and exploitation of word embeddings in both the input and output layers of a neural model by tracking contexts. This extends the dynamic entity representation used in Kobayashi et al. (2016) and incorporates a copy mechanism proposed independently by Gu et al. (2016) and Gulcehre et al. (2016). In addition, we construct a new task and dataset called Anonymized Language Modeling for evaluating the ability to capture word meanings while reading. Experiments conducted using our novel dataset show that the proposed variant of RNN language model outperformed the baseline model. Furthermore, the experiments also demonstrate that dynamic updates of an output layer help a model predict reappearing entities, whereas those of an input layer are effective to predict words following reappearing entities. Sosuke Kobayashi, Naoaki Okazaki, Kentaro Inui |
IJCNLP(1) | 3 |
| 2017 | LyriSys: An Interactive Support System for Writing Lyrics Based on Topic TransitionabstractThis paper presents LyriSys, a novel lyric-writing support system. Previous systems for lyric writing can fully automatically only generate a single line of lyrics that satisfies given constraints on accent and syllable patterns or an entire lyric. In contrast to such systems, LyriSys allows users to create and revise their work incrementally in a trial-and-error manner. Through fine-grained interactions with the system, the user can create the specifications of the musical structure and the story of the lyrics in terms of the verse-bridge-chorus structure, the number of lines, words and syllables, and most importantly, the transition over semantic topics such as "scene", "dark" and "sweet love". This paper provides an overview of the design of the system and its user interface and describes how the writing process is guided by a state-of-the-art probabilistic generative topic model that is trained without supervision. The system works for both Japanese and English. Kento Watanabe, Yuichiroh Matsubayashi, Kentaro Inui, Tomoyasu Nakano, Satoru Fukayama, Masataka Goto |
IUI | 3 |
| 2017 | A Crowdsourcing Approach for Annotating Causal Relation Instances in Wikipedia
Kazuaki Hanawa, Akira Sasaki, Naoaki Okazaki, Kentaro Inui |
PACLIC | 4 |
| 2017 | The mechanism of additive compositionabstractAdditive composition (Foltz et al. in Discourse Process 15:285–307, 1998 ; Landauer and Dumais in Psychol Rev 104(2):211, 1997 ; Mitchell and Lapata in Cognit Sci 34(8):1388–1429, 2010 ) is a widely used method for computing meanings of phrases, which takes the average of vector representations of the constituent words. In this article, we prove an upper bound for the bias of additive composition, which is the first theoretical analysis on compositional frameworks from a machine learning point of view. The bound is written in terms of collocation strength; we prove that the more exclusively two successive words tend to occur together, the more accurate one can guarantee their additive composition as an approximation to the natural phrase vector. Our proof relies on properties of natural language data that are empirically verified, and can be theoretically derived from an assumption that the data is generated from a Hierarchical Pitman–Yor Process. The theory endorses additive composition as a reasonable operation for calculating meanings of phrases, and suggests ways to improve additive compositionality, including: transforming entries of distributional word vectors by a function that meets a specific condition, constructing a novel type of vector representations to make additive composition sensitive to word order, and utilizing singular value decomposition to train word vectors. Naoaki Okazaki, Kentaro Inui |
Mach. Learn. | 3 |
| 2016 | Learning to Describe E-Commerce Images from Noisy Online Data
Takuya Yashima, Naoaki Okazaki, Kentaro Inui, Kota Yamaguchi, Takayuki Okatani |
ACCV (5) | 3 |
| 2016 | Composing Distributed Representations of Relational PatternsabstractLearning distributed representations for relation instances is a central technique in downstream NLP applications. In order to address semantic modeling of relational patterns, this paper constructs a new dataset that provides multiple similarity ratings for every pair of relational patterns on the existing dataset. In addition, we conduct a comparative study of different encoders including additive composition, RNN, LSTM, and GRU for composing distributed representations of relational patterns. We also present Gated Additive Composition, which is an enhancement of additive composition with the gating mechanism. Experiments show that the new dataset does not only enable detailed analyses of the different encoders, but also provides a gauge to predict successes of distributed representations of relational patterns in the relation classification task. Sho Takase, Naoaki Okazaki, Kentaro Inui |
ACL (1) | 3 |
| 2016 | Learning Semantically and Additively Compositional Distributional RepresentationsabstractThis paper connects a vector-based composition model to a formal semantics, the Dependency-based Compositional Semantics (DCS).We show theoretical evidence that the vector compositions in our model conform to the logic of DCS.Experimentally, we show that vector-based composition brings a strong ability to calculate similar phrases as similar vectors, achieving near state-of-the-art on a wide range of phrase similarity tasks and relation classification; meanwhile, DCS can guide building vectors for structured queries that can be directly executed.We evaluate this utility on sentence completion task and report a new state-of-the-art. Naoaki Okazaki, Kentaro Inui |
ACL (1) | 3 |
| 2016 | Modeling Context-sensitive Selectional Preference with Distributed RepresentationsabstractThis paper proposes a novel problem setting of selectional preference (SP) between a predicate and its arguments, called as context-sensitive SP (CSP). CSP models the narrative consistency between the predicate and preceding contexts of its arguments, in addition to the conventional SP based on semantic types. Furthermore, we present a novel CSP model that extends the neural SP model (Van de Cruys, 2014) to incorporate contextual information into the distributed representations of arguments. Experimental results demonstrate that the proposed CSP model successfully learns CSP and outperforms the conventional SP model in coreference cluster ranking. Naoya Inoue, Yuichiroh Matsubayashi, Masayuki Ono, Naoaki Okazaki, Kentaro Inui |
COLING | 5 |
| 2016 | Modeling Discourse Segments in Lyrics Using Repeated PatternsabstractThis study proposes a computational model of the discourse segments in lyrics to understand and to model the structure of lyrics. To test our hypothesis that discourse segmentations in lyrics strongly correlate with repeated patterns, we conduct the first large-scale corpus study on discourse segments in lyrics. Next, we propose the task to automatically identify segment boundaries in lyrics and train a logistic regression model for the task with the repeated pattern and textual features. The results of our empirical experiments illustrate the significance of capturing repeated patterns in predicting the boundaries of discourse segments in lyrics. Kento Watanabe, Yuichiroh Matsubayashi, Naho Orita, Naoaki Okazaki, Kentaro Inui, Satoru Fukayama, Tomoyasu Nakano, Jordan B. L. Smith, Masataka Goto |
COLING | 5 |
| 2016 | Question-Answering with Logic Specific to Video Games
Corentin Dumont, Kentaro Inui |
LREC | 3 |
| 2016 | Dynamic Entity Representation with Max-pooling Improves Machine ReadingabstractSosuke Kobayashi, Ran Tian, Naoaki Okazaki, Kentaro Inui. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Sosuke Kobayashi, Naoaki Okazaki, Kentaro Inui |
HLT-NAACL | 4 |
| 2016 | Recognizing Open-Vocabulary Relations between Objects in Images
Masayasu Muraoka, Sumit Maharjan, Masaki Saito, Kota Yamaguchi, Naoaki Okazaki, Takayuki Okatani, Kentaro Inui |
PACLIC | 7 |
| 2016 | Neural Joint Learning for Classifying Wikipedia Articles into Fine-grained Named Entity Types
Masatoshi Suzuki, Koji Matsuda, Satoshi Sekine, Naoaki Okazaki, Kentaro Inui |
PACLIC | 5 |
| 2016 | Toward the automatic extraction of knowledge of usable goods
Mei Uemura, Naho Orita, Naoaki Okazaki, Kentaro Inui |
PACLIC | 4 |
| 2016 | Stance Classification by Recognizing Related Events about TargetsabstractRecently, many people express their opinions using social networking services such as Twitter and Facebook. Each opinion has a stance related to something such as product, service, and politics. The task of detecting a stance is known as sentiment analysis, reputation mining, and stance detection. A popular approach for stance detection uses sentiment polarity towards a target in a text. This approach is known as targeted sentiment analysis. If a target appears in text, the detecting stance based on targeted sentiment polarity would work well. However, how can we detect stance towards an event? (e.g. "I cannot understand why man can marry only with a woman", "The problem of low birth rate becomes more severe" to the event "Allowing same-sex marriage"). To detect these stances, it is necessary to recognize a situation in which the event occurs or does not occur. To classify texts including these phenomena, we propose a classification method based on machine learning considering PRIOR-SITUATION and EFFECT. Akira Sasaki, Junta Mizuno, Naoaki Okazaki, Kentaro Inui |
WI | 4 |
| 2016 | Fine-Grained Named Entity Classification with Wikipedia Article VectorsabstractThis paper addresses the task of assigning multiple labels of fine-grained named entity (NE) types to Wikipedia articles. To address the sparseness of the input feature space, which is salient particularly in fine-grained type classification, we propose to learn article vectors (i.e. entity embeddings) from hypertext structure of Wikipedia using a Skip-gram model and incorporate them into the input feature set. To conduct large-scale practical experiments, we created a new dataset containing over 22,000 manually labeled instances. The results of our experiments show that our idea gained statistically significant improvements in classification results. Masatoshi Suzuki, Koji Matsuda, Satoshi Sekine, Naoaki Okazaki, Kentaro Inui |
WI | 5 |
| 2016 | Modeling semantic compositionality of relational patternsabstractVector representation is a common approach for expressing the meaning of a relational pattern. Most previous work obtained a vector of a relational pattern based on the distribution of its context words (e.g., arguments of the relational pattern), regarding the pattern as a single ‘word’. However, this approach suffers from the data sparseness problem, because relational patterns are productive, i.e., produced by combinations of words. To address this problem, we propose a novel method for computing the meaning of a relational pattern based on the semantic compositionality of constituent words. We extend the Skip-gram model (Mikolov et al., 2013) to handle semantic compositions of relational patterns using recursive neural networks. The experimental results show the superiority of the proposed method for modeling the meanings of relational patterns, and demonstrate the contribution of this work to the task of relation extraction. Sho Takase, Naoaki Okazaki, Kentaro Inui |
Eng. Appl. Artif. Intell. | 3 |
| 2015 | Reducing Lexical Features in Parsing by Word Embeddings
Hiroya Komatsu, Naoaki Okazaki, Kentaro Inui |
PACLIC | 4 |
| 2015 | Recognizing Complex Negation on Twitter
Junta Mizuno, Canasai Kruengkrai, Kiyonori Ohtake, Chikara Hashimoto, Kentaro Torisawa, Julien Kloetzer, Kentaro Inui |
PACLIC | 7 |
| 2015 | Fast and Large-scale Unsupervised Relation Extraction
Sho Takase, Naoaki Okazaki, Kentaro Inui |
PACLIC | 3 |
| 2014 | An Example-Based Approach to Difficult Pronoun Resolution
Canasai Kruengkrai, Naoya Inoue, Jun Sugiura, Kentaro Inui |
PACLIC | 4 |
| 2014 | Finding The Best Model Among Representative Compositional Models
Masayasu Muraoka, Sonse Shimaoka, Kazeto Yamamoto, Yotaro Watanabe, Naoaki Okazaki, Kentaro Inui |
PACLIC | 6 |
| 2014 | Modeling Structural Topic Transitions for Automatic Lyrics Generation
Kento Watanabe, Yuichiroh Matsubayashi, Kentaro Inui, Masataka Goto |
PACLIC | 3 |
| 2013 | Is a 204 cm Man Tall or Small ? Acquisition of Numerical Common Sense from the Web
Katsuma Narisawa, Yotaro Watanabe, Junta Mizuno, Naoaki Okazaki, Kentaro Inui |
ACL (1) | 5 |
| 2013 | Evidence in Automatic Error Correction Improves Learners' English Skill
Jiro Umezawa, Junta Mizuno, Naoaki Okazaki, Kentaro Inui |
CICLing (2) | 4 |
| 2013 | Discriminative Learning of First-Order Weighted Abduction from Partial Discourse Explanations
Kazeto Yamamoto, Naoya Inoue, Yotaro Watanabe, Naoaki Okazaki, Kentaro Inui |
CICLing (1) | 5 |
| 2013 | A Lexicon-based Investigation of Research Issues in Japanese Factuality Analysis
Kazuya Narita, Junta Mizuno, Kentaro Inui |
IJCNLP | 3 |
| 2013 | NICT Disaster Information Analysis System
Kiyonori Ohtake, Jun Goto, Stijn De Saeger, Kentaro Torisawa, Junta Mizuno, Kentaro Inui |
IJCNLP | 6 |
| 2013 | Inducing Context Gazetteers from Encyclopedic Databases for Named Entity Recognition
Hancheol Cho, Naoaki Okazaki, Kentaro Inui |
PAKDD (1) | 3 |
| 2012 | Coreference Resolution with ILP-based Weighted Abduction
Naoya Inoue, Ekaterina Ovchinnikova, Kentaro Inui, Jerry R. Hobbs |
COLING | 3 |
| 2012 | A Latent Discriminative Model for Compositional Entailment Relation Recognition using Natural Logic
Yotaro Watanabe, Junta Mizuno, Eric Nichols, Naoaki Okazaki, Kentaro Inui |
COLING | 5 |
| 2012 | Large-Scale Cost-Based Abduction in Full-Fledged First-Order Predicate Logic with Cutting Plane Inference
Naoya Inoue, Kentaro Inui |
JELIA | 2 |
| 2012 | Set Expansion using Sibling Relations between Semantic Categories
Sho Takase, Naoaki Okazaki, Kentaro Inui |
PACLIC | 3 |
| 2012 | Leveraging Diverse Lexical Resources for Textual Entailment RecognitionabstractSince the problem of textual entailment recognition requires capturing semantic relations between diverse expressions of language, linguistic and world knowledge play an important role. In this article, we explore the effectiveness of different types of currently available resources including synonyms, antonyms, hypernym-hyponym relations, and lexical entailment relations for the task of textual entailment recognition. In order to do so, we develop an entailment relation recognition system which utilizes diverse linguistic analyses and resources to align the linguistic units in a pair of texts and identifies entailment relations based on these alignments. We use the Japanese subset of the NTCIR-9 RITE-1 dataset for evaluation and error analysis, conducting ablation testing and evaluation on hand-crafted alignment gold standard data to evaluate the contribution of individual resources. Error analysis shows that existing knowledge sources are effective for RTE, but that their coverage is limited, especially for domain-specific and other low-frequency expressions. To increase alignment coverage on such expressions, we propose a method of alignment inference that uses syntactic and semantic dependency information to identify likely alignments without relying on external resources. Evaluation adding alignment inference to a system using all available knowledge sources shows improvements in both precision and recall of entailment relation recognition. Yotaro Watanabe, Junta Mizuno, Eric Nichols, Katsuma Narisawa, Keita Nabeshima, Naoaki Okazaki, Kentaro Inui |
ACM Trans. Asian Lang. Inf. Process. | 7 |
| 2011 | Dependency Syntax Analysis Using Grammar Induction and a Lexical Categories Precedence System
Hiram Calvo, Omar Juárez-Gambino, Alexander F. Gelbukh, Kentaro Inui |
CICLing (1) | 4 |
| 2011 | Co-related Verb Argument Selectional Preferences
Hiram Calvo, Kentaro Inui, Yuji Matsumoto 0001 |
CICLing (1) | 2 |
| 2011 | Web Spam Detection by Exploring Densely Connected SubgraphsabstractIn this paper, we present a Web spam detection algorithm that relies on link analysis. The method consists of three steps: (1) decomposition of web graphs in densely connected sub graphs and calculation of the features for each sub graph, (2) use of SVM classifiers to identify sub graphs composed of Web spam, and (3) propagation of predictions over web graphs by a biased Page Rank algorithm to expand the scope of identification. We performed experiments on a public benchmark. An empirical study of the core structure of web graphs suggests that highly ranked non-spam hosts can be identified by viewing the coreness of the web graph elements. Yutaka I. Leon-Suematsu, Kentaro Inui, Sadao Kurohashi, Yutaka Kidawara |
Web Intelligence | 2 |
| 2011 | Mining personal experiences and opinions from Web documentsabstractThis paper proposes a new UGC-oriented language technology application, which we call experience mining. Experience mining aims at automatically collecting instances of personal experiences as well as opinions from vast amounts of user generated cont Shuya Abe, Kentaro Inui, Kazuo Hara, Hiraku Morita, Chitose Sao, Megumi Eguchi, Asuka Sumida, Koji Murakami, Suguru Matsuyoshi |
Web Intell. Agent Syst. | 2 |
| 2010 | Annotating Event Mentions in Text with Modality, Focus, and Source Information
Suguru Matsuyoshi, Megumi Eguchi, Chitose Sao, Koji Murakami, Kentaro Inui, Yuji Matsumoto 0001 |
LREC | 5 |
| 2010 | Dependency Tree-based Sentiment Classification using CRFs with Hidden Variables
Tetsuji Nakagawa, Kentaro Inui, Sadao Kurohashi |
HLT-NAACL | 2 |
| 2009 | Capturing Salience with a Trainable Cache Model for Zero-anaphora Resolution
Ryu Iida, Kentaro Inui, Yuji Matsumoto 0001 |
ACL/IJCNLP | 2 |
| 2009 | Learning Co-relations of Plausible Verb Arguments with a WSM and a Distributional Thesaurus
Hiram Calvo, Kentaro Inui, Yuji Matsumoto 0001 |
CIARP | 2 |
| 2009 | Interpolated PLSI for Learning Plausible Verb Arguments
Hiram Calvo, Kentaro Inui, Yuji Matsumoto 0001 |
PACLIC | 2 |
| 2009 | Identifying Information Sender Configuration of Web PagesabstractThe source of a piece of information is a crucial element to consider when judging the credibility of that information. In this paper, we address the task of identifying the information source which is cast as a problem of identifying the {\em information sender configuration (ISC)} of a Web page. An information sender of a Web page is an entity which is involved in the publication of the information on the page. An ISC of a Web page describes the information senders of the page and the relationship among them. Information sender extraction is thus a subtask of identifying ISC, and we present a method for extracting information senders from Web pages and offer preliminary evaluation. The ISC provides a basis for deeper analysis of information on the Web. Yoshikiyo Kato, Daisuke Kawahara, Kentaro Inui, Sadao Kurohashi, Tomohide Shibata |
Web Intelligence | 3 |
| 2008 | Two-Phased Event Relation Acquisition: Coupling the Relation-Oriented and Argument-Oriented Approaches
Shuya Abe, Kentaro Inui, Yuji Matsumoto 0001 |
COLING | 2 |
| 2008 | Emotion Classification Using Massive Examples Extracted from the Web
Ryoko Tokuhisa, Kentaro Inui, Yuji Matsumoto 0001 |
COLING | 2 |
| 2008 | Acquiring Event Relation Knowledge by Learning Cooccurrence Patterns and Fertilizing Cooccurrence Samples with Verbal Nouns
Shuya Abe, Kentaro Inui, Yuji Matsumoto 0001 |
IJCNLP | 2 |
| 2008 | Experience Mining: Building a Large-Scale Database of Personal Experiences and Opinions from Web DocumentsabstractThis paper proposes a new UGC-oriented language technology application, which we call experience mining. Experience mining aims at automatically collecting instances of personal experiences as well as opinions from an explosive number of user generated contents (UGCs) such as Weblog and forum posts and storing them in an experience database with semantically rich indices. After arguing the technical issues of this new task, we focus on the central problem, factuality analysis, among others and propose a machine learning-based solution as well as the task definition itself. Our empirical evaluation indicates that our factuality analysis task is sufficiently well-defined to achieve a high inter-annotator agreement and our factorial CRF-based model considerably outperforms the baseline. We also present an application system, which currently stores over 50M experience instances extracted from 150M Japanese blog posts with semantic indices and is scheduled to start serving as an experience search engine for unrestricted users in October. Kentaro Inui, Shuya Abe, Kazuo Hara, Hiraku Morita, Chitose Sao, Megumi Eguchi, Asuka Sumida, Koji Murakami, Suguru Matsuyoshi |
Web Intelligence | 1 |
| 2008 | Grasping Major Statements and Their Contradictions Toward Information Credibility Analysis of Web ContentsabstractThe World Wide Web contains wide variety of news reports, arguments, opinions, etc. that vary widely in quality. People judge the credibility of information on the Web for decision making in daily life. At present, while the quantity of information on the Web is explosively increasing, it is necessary to develop a system that supports such judgments. We have been developing an information credibility analysis system, WISDOM that considers the viewpoints of information contents, information senders, and information appearances. In this paper, as a viewpoint of information contents, we propose a method for providing a bird's eye view of major statements on a given topic and their contradictions. We evaluate the obtained statements in our experiments, and confirm the effectiveness of our approach. Furthermore, we discuss our future objectives. Daisuke Kawahara, Sadao Kurohashi, Kentaro Inui |
Web Intelligence | 3 |
| 2007 | Extracting Aspect-Evaluation and Aspect-Of Relations in Opinion Mining
Nozomi Kobayashi, Kentaro Inui, Yuji Matsumoto 0001 |
EMNLP-CoNLL | 2 |
| 2007 | Zero-anaphora resolution by learning rich syntactic pattern featuresabstractWe approach the zero-anaphora resolution problem by decomposing it into intrasentential and intersentential zero-anaphora resolution tasks. For the former task, syntactic patterns of zeropronouns and their antecedents are useful clues. Taking Japanese as a target language, we empirically demonstrate that incorporating rich syntactic pattern features in a state-of-the-art learning-based anaphora resolution model dramatically improves the accuracy of intrasentential zero-anaphora, which consequently improves the overall performance of zero-anaphora resolution. Ryu Iida, Kentaro Inui, Yuji Matsumoto 0001 |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2006 | Exploiting Syntactic Patterns as Clues in Zero-Anaphora ResolutionabstractWe approach the zero-anaphora resolution problem by decomposing it into intra-sentential and inter-sentential zero-anaphora resolution. For the former problem, syntactic patterns of the appearance of zero-pronouns and their antecedents are useful clues. Taking Japanese as a target language, we empirically demonstrate that incorporating rich syntactic pattern features in a state-of-the-art learning-based anaphora resolution model dramatically improves the accuracy of intra-sentential zero-anaphora, which consequently improves the overall performance of zero-anaphora resolution. Ryu Iida, Kentaro Inui, Yuji Matsumoto 0001 |
ACL | 2 |
| 2006 | Augmenting a Semantic Verb Lexicon with a Large Scale Collection of Example Sentences
Kentaro Inui, Toru Hirano, Ryu Iida, Atsushi Fujita, Yuji Matsumoto 0001 |
LREC | 1 |
| 2005 | Exploiting Lexical Conceptual Structure for Paraphrase Generation
Atsushi Fujita, Kentaro Inui, Yuji Matsumoto 0001 |
IJCNLP | 2 |
| 2005 | Anaphora resolution by antecedent identification followed by anaphoricity determinationabstractWe propose a machine learning-based approach to noun-phrase anaphora resolution that combines the advantages of previous learning-based models while overcoming their drawbacks. Our anaphora resolution process reverses the order of the steps in the classification-then-search model proposed by Ng and Cardie [2002b], inheriting all the advantages of that model. We conducted experiments on resolving noun-phrase anaphora in Japanese. The results show that with the selection-then-classification-based modifications, our proposed model outperforms earlier learning-based approaches. Ryu Iida, Kentaro Inui, Yuji Matsumoto 0001 |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2005 | Acquiring causal knowledge from text using the connective marker tameabstractIn this paper, we deal with automatic knowledge acquisition from text, specifically the acquisition of causal relations . A causal relation is the relation existing between two events such that one event causes (or enables) the other event, such as “hard rain causes flooding” or “taking a train requires buying a ticket.” In previous work these relations have been classified into several types based on a variety of points of view. In this work, we consider four types of causal relations--- cause , effect , precond(ition) and means ---mainly based on agents' volitionality, as proposed in the research field of discourse understanding. The idea behind knowledge acquisition is to use resultative connective markers, such as “because,” “but,” and “if” as linguistic cues. However, there is no guarantee that a given connective marker always signals the same type of causal relation. Therefore, we need to create a computational model that is able to classify samples according to the causal relation. To examine how accurately we can automatically acquire causal knowledge, we attempted an experiment using Japanese newspaper articles, focusing on the resultative connective “tame.” By using machine-learning techniques, we achieved 80% recall with over 95% precision for the cause , precond , and means relations, and 30% recall with 90% precision for the effect relation. Furthermore, the classification results suggest that one can expect to acquire over 27,000 instances of causal relations from 1 year of Japanese newspaper articles. Takashi Inui, Kentaro Inui, Yuji Matsumoto 0001 |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2004 | Detection of Incorrect Case Assignments in Paraphrase Generation
Atsushi Fujita, Kentaro Inui, Yuji Matsumoto 0001 |
IJCNLP | 2 |
| 2004 | Collecting Evaluative Expressions for Opinion Extraction
Nozomi Kobayashi, Kentaro Inui, Yuji Matsumoto 0001, Kenji Tateishi, Toshikazu Fukushima |
IJCNLP | 2 |
| 2003 | What Kinds and Amounts of Causal Knowledge Can Be Acquired from Text by Using Connective Markers as Clues?
Takashi Inui, Kentaro Inui, Yuji Matsumoto 0001 |
Discovery Science | 2 |
| 2000 | Committee-based Decision Making in Probabiiistic Partial Parsing
Takashi Inui, Kentaro Inui |
COLING | 2 |
| 1998 | An Empirical Evaluation on Statistical Parsing of Japanese Sentences Using Lexical Association Statistics
Kiyoaki Shirai, Kentaro Inui, Takenobu Tokunaga, Hozumi Tanaka |
EMNLP | 2 |
| 1998 | Selective Sampling for Example-based Word Sense Disambiguation
Atsushi Fujii, Kentaro Inui, Takenobu Tokunaga, Hozumi Tanaka |
Comput. Linguistics | 2 |
| 1996 | To what extent does case contribute to verb sense disambiguation?
Atsushi Fujii, Kentaro Inui, Takenobu Tokunaga, Hozumi Tanaka |
COLING | 2 |
| 1996 | The Internet a "natural" channel for language learning
Kentaro Inui |
COLING | 1 |