Naoya Inoue

dblp:48/4618 · DBLP profile ↗
← Back
34ranked-venue papers
5as first author
21since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 5 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 On Effects of Steering Latent Representation for Large Language Model Unlearning
abstract
Representation Misdirection for Unlearning (RMU), which steers model representation in the intermediate layer to a target random representation, is an effective method for large language model (LLM) unlearning. Despite its high performance, the underlying cause and explanation remain underexplored. In this paper, we theoretically demonstrate that steering forget representations in the intermediate layer reduces token confidence, causing LLMs to generate wrong or nonsense responses. We investigate how the coefficient influences the alignment of forget-sample representations with the random direction and hint at the optimal coefficient values for effective unlearning across different network layers. We show that RMU unlearned models are robust against adversarial jailbreak attacks. Furthermore, our empirical analysis shows that RMU is less effective when applied to the middle and later layers in LLMs. To resolve this drawback, we propose Adaptive RMU---a simple yet effective alternative method that makes unlearning effective with most layers. Extensive experiments demonstrate that Adaptive RMU significantly improves the unlearning performance compared to prior art while incurring no additional computational cost.
Huu-Tien Dang, Tin Pham, Hoang Thanh-Tung, Naoya Inoue
AAAI4
2025 Understanding Token Probability Encoding in Output Embeddings
abstract
In this paper, we investigate the output token probability information in the output embedding of language models. We find an approximate common log-linear encoding of output token probabilities within the output embedding vectors and empirically demonstrate that it is accurate and sparse. As a causality examination, we steer the encoding in output embedding to modify the output probability distribution accurately. Moreover, the sparsity we find in output probability encoding suggests that a large number of dimensions in the output embedding do not contribute to causal language modeling. Therefore, we attempt to delete the output-unrelated dimensions and find more than 30% of the dimensions can be deleted without significant movement in output distribution and sequence generation. Additionally, in the pre-training dynamics of language models, we find that the output embeddings capture the corpus token frequency information in early steps, even before an obvious convergence of parameters starts.
Hakaze Cho, Yoshihiro Sakai, Kenshiro Tanaka, Mariko Kato, Naoya Inoue
COLING5
2025 The Transfer Neurons Hypothesis: An Underlying Mechanism for Language Latent Space Transitions in Multilingual LLMs
abstract
Recent studies have suggested a processing framework for multilingual inputs in decoder-based LLMs: early layers convert inputs into English-centric and language-agnostic representations; middle layers perform reasoning within an English-centric latent space; and final layers generate outputs by transforming these representations back into language-specific latent spaces.However, the internal dynamics of such transformation and the underlying mechanism remain underexplored.Towards a deeper understanding of this framework, we propose and empirically validate The Transfer Neurons Hypothesis: certain neurons in the MLP module are responsible for transferring representations between language-specific latent spaces and a shared semantic latent space.Furthermore, we show that one function of language-specific neurons, as identified in recent studies, is to facilitate movement between latent spaces.Finally, we show that transfer neurons are critical for reasoning in multilingual LLMs
Hinata Tezuka, Naoya Inoue
EMNLP2
2025 Identification of Multiple Logical Interpretations in Counter-Arguments
abstract
Wenzhi Wang, Paul Reisert, Shoichi Naito, Naoya Inoue, Machi Shimmei, Surawat Pothong, Jungmin Choi, Kentaro Inui. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Paul Reisert, Shoichi Naito, Naoya Inoue, Machi Shimmei, Surawat Pothong, Jungmin Choi, Kentaro Inui
EMNLP4
2025 CyLLM-DAP: Cybersecurity Domain-Adaptive Pre-Training Framework of Large Language Models
Khang Mai, Razvan Beuran, Naoya Inoue
ICISSP (2)3
2025 Revisiting In-context Learning Inference Circuit in Large Language Models
abstract
In-context Learning (ICL) is an emerging few-shot learning paradigm on Language Models (LMs) with inner mechanisms un-explored. There are already existing works describing the inner processing of ICL, while they struggle to capture all the inference phenomena in large language models. Therefore, this paper proposes a comprehensive circuit to model the inference dynamics and try to explain the observed phenomena of ICL. In detail, we divide ICL inference into 3 major operations: (1) Input Text Encode: LMs encode every input text (in the demonstrations and queries) into linear representation in the hidden states with sufficient information to solve ICL tasks. (2) Semantics Merge: LMs merge the encoded representations of demonstrations with their corresponding label tokens to produce joint representations of labels and demonstrations. (3) Feature Retrieval and Copy: LMs search the joint representations of demonstrations similar to the query representation on a task subspace, and copy the searched representations into the query. Then, language model heads capture these copied label representations to a certain extent and decode them into predicted labels. Through careful measurements, the proposed inference circuit successfully captures and unifies many fragmented phenomena observed during the ICL process, making it a comprehensive and practical explanation of the ICL inference process. Moreover, ablation analysis by disabling the proposed steps seriously damages the ICL performance, suggesting the proposed inference circuit is a dominating mechanism. Additionally, we confirm and list some bypass mechanisms that solve ICL tasks in parallel with the proposed circuit.
Hakaze Cho, Mariko Kato, Yoshihiro Sakai, Naoya Inoue
ICLR4
2025 Token-based Decision Criteria Are Suboptimal in In-context Learning
abstract
Hakaze Cho, Yoshihiro Sakai, Mariko Kato, Kenshiro Tanaka, Akira Ishii, Naoya Inoue. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Hakaze Cho, Yoshihiro Sakai, Mariko Kato, Kenshiro Tanaka, Akira Ishii, Naoya Inoue
NAACL (Long Papers)6
2025 Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
abstract
The unusual properties of in-context learning (ICL) have prompted investigations into the internal mechanisms of large language models. Prior work typically focuses on either special attention heads or task vectors at specific layers, but lacks a unified framework linking these components to the evolution of hidden states across layers that ultimately produce the model’s output. In this paper, we propose such a framework for ICL in classification tasks by analyzing two geometric factors that govern performance: the separability and alignment of query hidden states. A fine-grained analysis of layer-wise dynamics reveals a striking two-stage mechanism—separability emerges in early layers, while alignment develops in later layers. Ablation studies further show that Previous Token Heads drive separability, while Induction Heads and task vectors enhance alignment. Our findings thus bridge the gap between attention heads and task vectors, offering a unified account of ICL’s underlying mechanisms.
Haolin Yang 0003, Hakaze Cho, Yiqiao Zhong, Naoya Inoue
NeurIPS4
2025 Improve Smart Contract Vulnerability Explanation with Synthetic Data and Chain-of-Thought Prompting
Minh Le Nguyen 0001, Naoya Inoue
NLDB (1)2
2025 Non-Interactive Symbolic-Aided Chain-of-Thought for Logical Reasoning
Phuong Minh Nguyen 0001, Tien Dang, Naoya Inoue
PACLIC3
2024 JEMHopQA: Dataset for Japanese Explainable Multi-Hop Question Answering
abstract
We present JEMHopQA, a multi-hop QA dataset for the development of explainable QA systems. The dataset consists not only of question-answer pairs, but also of supporting evidence in the form of derivation triples, which contributes to making the QA task more realistic and difficult. It is created based on Japanese Wikipedia using both crowd-sourced human annotation as well as prompting a large language model (LLM), and contains a diverse set of question, answer and topic categories as compared with similar datasets released previously. We describe the details of how we built the dataset as well as the evaluation of the QA task presented by this dataset using GPT-4, and show that the dataset is sufficiently challenging for the state-of-the-art LLM while showing promise for combining such a model with existing knowledge resources to achieve better performance.
Ai Ishii, Naoya Inoue, Hisami Suzuki, Satoshi Sekine
LREC/COLING2
2024 Find-the-Common: A Benchmark for Explaining Visual Patterns from Images
abstract
Recent advances in Instruction-fine-tuned Vision and Language Models (IVLMs), such as GPT-4V and InstructBLIP, have prompted some studies have started an in-depth analysis of the reasoning capabilities of IVLMs. However, Inductive Visual Reasoning, a vital skill for text-image understanding, remains underexplored due to the absence of benchmarks. In this paper, we introduce Find-the-Common (FTC): a new vision and language task for Inductive Visual Reasoning. In this task, models are required to identify an answer that explains the common attributes across visual scenes. We create a new dataset for the FTC and assess the performance of several contemporary approaches including Image-Based Reasoning, Text-Based Reasoning, and Image-Text-Based Reasoning with various models. Extensive experiments show that even state-of-the-art models like GPT-4V can only archive with 48% accuracy on the FTC, for which, the FTC is a new challenge for the visual reasoning research community. Our dataset has been released and is available online: https://github.com/SSSSSeki/Find-the-common.
Naoya Inoue, Houjing Wei
LREC/COLING2
2024 Flee the Flaw: Annotating the Underlying Logic of Fallacious Arguments Through Templates and Slot-filling
abstract
Irfan Robbani, Paul Reisert, Surawat Pothong, Naoya Inoue, Camélia Guerraoui, Wenzhi Wang, Shoichi Naito, Jungmin Choi, Kentaro Inui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Irfan Robbani, Paul Reisert, Surawat Pothong, Naoya Inoue, Camélia Guerraoui, Shoichi Naito, Jungmin Choi, Kentaro Inui
EMNLP4
2024 A Viewpoints Embedded Diff-table System For Cross-sectional Insight Survey In a Research Task
Jinghong Li, Naoya Inoue, Shinobu Hasegawa
PACLIC2
2023 Counterfactual Adversarial Training for Improving Robustness of Pre-trained Language Models
Hoai Linh Luu, Naoya Inoue
PACLIC2
2022 LPAttack: A Feasible Annotation Scheme for Capturing Logic Pattern of Attacks in Arguments
abstract
In argumentative discourse, persuasion is often achieved by refuting or attacking others’ arguments. Attacking an argument is not always straightforward and often consists of complex rhetorical moves in which arguers may agree with a logic of an argument while attacking another logic. Furthermore, an arguer may neither deny nor agree with any logics of an argument, instead ignore them and attack the main stance of the argument by providing new logics and presupposing that the new logics have more value or importance than the logics presented in the attacked argument. However, there are no studies in computational argumentation that capture such complex rhetorical moves in attacks or the presuppositions or value judgments in them. To address this gap, we introduce LPAttack, a novel annotation scheme that captures the common modes and complex rhetorical moves in attacks along with the implicit presuppositions and value judgments. Our annotation study shows moderate inter-annotator agreement, indicating that human annotation for the proposed scheme is feasible. We publicly release our annotated corpus and the annotation guidelines.
Farjana Sultana Mim, Naoya Inoue, Shoichi Naito, Keshav Singh 0003, Kentaro Inui
LREC2
2022 TYPIC: A Corpus of Template-Based Diagnostic Comments on Argumentation
abstract
Providing feedback on the argumentation of the learner is essential for developing critical thinking skills, however, it requires a lot of time and effort. To mitigate the overload on teachers, we aim to automate a process of providing feedback, especially giving diagnostic comments which point out the weaknesses inherent in the argumentation. It is recommended to give specific diagnostic comments so that learners can recognize the diagnosis without misinterpretation. However, it is not obvious how the task of providing specific diagnostic comments should be formulated. We present a formulation of the task as template selection and slot filling to make an automatic evaluation easier and the behavior of the model more tractable. The key to the formulation is the possibility of creating a template set that is sufficient for practical use. In this paper, we define three criteria that a template set should satisfy: expressiveness, informativeness, and uniqueness, and verify the feasibility of creating a template set that satisfies these criteria as a first trial. We will show that it is feasible through an annotation study that converts diagnostic comments given in a text to a template format. The corpus used in the annotation study is publicly available.
Shoichi Naito, Shintaro Sawada, Chihiro Nakagawa, Naoya Inoue, Kenshi Yamaguchi, Iori Shimizu, Farjana Sultana Mim, Keshav Singh 0003, Kentaro Inui
LREC4
2022 IRAC: A Domain-Specific Annotated Corpus of Implicit Reasoning in Arguments
abstract
The task of implicit reasoning generation aims to help machines understand arguments by inferring plausible reasonings (usually implicit) between argumentative texts. While this task is easy for humans, machines still struggle to make such inferences and deduce the underlying reasoning. To solve this problem, we hypothesize that as human reasoning is guided by innate collection of domain-specific knowledge, it might be beneficial to create such a domain-specific corpus for machines. As a starting point, we create the first domain-specific resource of implicit reasonings annotated for a wide range of arguments, which can be leveraged to empower machines with better implicit reasoning generation ability. We carefully design an annotation framework to collect them on a large scale through crowdsourcing and show the feasibility of creating a such a corpus at a reasonable cost and high-quality. Our experiments indicate that models trained with domain-specific implicit reasonings significantly outperform domain-general models in both automatic and human evaluations. To facilitate further research towards implicit reasoning generation in arguments, we present an in-depth analysis of our corpus and crowdsourcing methodology, and release our materials (i.e., crowdsourcing guidelines and domain-specific resource of implicit reasonings).
Keshav Singh 0003, Naoya Inoue, Farjana Sultana Mim, Shoichi Naito, Kentaro Inui
LREC2
2021 Two Training Strategies for Improving Relation Extraction over Universal Graph
abstract
This paper explores how the Distantly Supervised Relation Extraction (DS-RE) can benefit from the use of a Universal Graph (UG), the combination of a Knowledge Graph (KG) and a large-scale text collection.A straightforward extension of a current state-of-the-art neural model for DS-RE with a UG may lead to degradation in performance.We first report that this degradation is associated with the difficulty in learning a UG and then propose two training strategies: (1) Path Type Adaptive Pretraining, which sequentially trains the model with different types of UG paths so as to prevent the reliance on a single type of UG path; and (2) Complexity Ranking Guided Attention mechanism, which restricts the attention span according to the complexity of a UG path so as to force the model to extract features not only from simple UG paths but also from complex ones.Experimental results on both biomedical and NYT10 datasets prove the robustness of our methods and achieve a new state-ofthe-art result on the NYT10 dataset.The code and datasets used in this paper are available at https://github.com/baodaiqin/ UGDSRE.
Qin Dai, Naoya Inoue, Kentaro Inui
EACL2
2021 Summarize-then-Answer: Generating Concise Explanations for Multi-hop Reading Comprehension
abstract
How can we generate concise explanations for multi-hop Reading Comprehension (RC)?The current strategies of identifying supporting sentences can be seen as an extractive questionfocused summarization of the input text.However, these extractive explanations are not necessarily concise i.e. not minimally sufficient for answering a question.Instead, we advocate for an abstractive approach, where we propose to generate a question-focused, abstractive summary of input paragraphs and then feed it to an RC system.Given a limited amount of human-annotated abstractive explanations, we train the abstractive explainer in a semi-supervised manner, where we start from the supervised model and then train it further through trial and error maximizing a conciseness-promoted reward function.Our experiments demonstrate that the proposed abstractive explainer can generate more compact explanations than an extractive explainer with limited supervision (only 2k instances) while maintaining sufficiency.1 Our implementation is publicly available at https:// github.com/StonyBrookNLP/suqa.Charlie Rowe plays Billy Costa in a film based on what novel?[P1] [1] The Golden Compass is a 2007 British-American fantasy adventure film based on "Northern Lights", the first novel in Philip Pullman's trilogy "His Dark Materials".
Naoya Inoue, Harsh Trivedi, Steven Sinha, Niranjan Balasubramanian, Kentaro Inui
EMNLP (1)1
2021 Corruption Is Not All Bad: Incorporating Discourse Structure Into Pre-Training via Corruption for Essay Scoring
abstract
Existing approaches for automated essay scoring and document representation learning typically rely on discourse parsers to incorporate discourse structure into text representation. However, the performance of parsers is not always adequate, especially when they are used on noisy texts, such as student essays. In this paper, we propose an unsupervised pre-training approach to capture discourse structure of essays in terms of coherence and cohesion that does not require any discourse parser or annotation. We introduce several types of token, sentence and paragraph-level corruption techniques for our proposed pre-training approach and augment masked language modeling pre-training with our pre-training method to leverage both contextualized and discourse information. Our proposed unsupervised approach achieves a new state-of-the-art result on the task of essay Organization scoring.
Farjana Sultana Mim, Naoya Inoue, Paul Reisert, Hiroki Ouchi, Kentaro Inui
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 R4C: A Benchmark for Evaluating RC Systems to Get the Right Answer for the Right Reason
abstract
Recent studies have revealed that reading comprehension (RC) systems learn to exploit annotation artifacts and other biases in current datasets.This prevents the community from reliably measuring the progress of RC systems.To address this issue, we introduce R 4 C, a new task for evaluating RC systems' internal reasoning.R 4 C requires giving not only answers but also derivations: explanations that justify predicted answers.We present a reliable, crowdsourced framework for scalably annotating RC datasets with derivations.We create and publicly release the R 4 C dataset, the first, quality-assured dataset consisting of 4.6k questions, each of which is annotated with 3 reference derivations (i.e.13.8k derivations).Experiments show that our automatic evaluation metrics using multiple reference derivations are reliable, and that R 4 C assesses different skills from an existing benchmark.
Naoya Inoue, Pontus Stenetorp, Kentaro Inui
ACL1
2020 Modeling Event Salience in Narratives via Barthes' Cardinal Functions
abstract
Events in a narrative differ in salience: some are more important to the story than others.Estimating event salience is useful for tasks such as story generation, and as a tool for text analysis in narratology and folkloristics.To compute event salience without any annotations, we adopt Barthes' definition of event salience and propose several unsupervised methods that require only a pre-trained language model.Evaluating the proposed methods on folktales with event salience annotation, we show that the proposed methods outperform baseline methods and find fine-tuning a language model on narrative texts is a key factor in improving the proposed methods.
Takaki Otake, Sho Yokoi, Naoya Inoue, Tatsuki Kuribayashi, Kentaro Inui
COLING3
2019 An Empirical Study of Span Representations in Argumentation Structure Parsing
abstract
Tatsuki Kuribayashi, Hiroki Ouchi, Naoya Inoue, Paul Reisert, Toshinori Miyoshi, Jun Suzuki, Kentaro Inui. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Tatsuki Kuribayashi, Hiroki Ouchi, Naoya Inoue, Paul Reisert, Toshinori Miyoshi, Jun Suzuki 0001, Kentaro Inui
ACL (1)3
2019 An Annotation Protocol for Collecting User-Generated Counter-Arguments Using Crowdsourcing
Paul Reisert, Gisela Vallejo, Naoya Inoue, Iryna Gurevych, Kentaro Inui
AIED (2)3
2018 Improving Scientific Relation Classification with Task Specific Supersense
Qin Dai, Naoya Inoue, Paul Reisert, Kentaro Inui
PACLIC2
2016 Modeling Context-sensitive Selectional Preference with Distributed Representations
abstract
This paper proposes a novel problem setting of selectional preference (SP) between a predicate and its arguments, called as context-sensitive SP (CSP). CSP models the narrative consistency between the predicate and preceding contexts of its arguments, in addition to the conventional SP based on semantic types. Furthermore, we present a novel CSP model that extends the neural SP model (Van de Cruys, 2014) to incorporate contextual information into the distributed representations of arguments. Experimental results demonstrate that the proposed CSP model successfully learns CSP and outperforms the conventional SP model in coreference cluster ranking.
Naoya Inoue, Yuichiroh Matsubayashi, Masayuki Ono, Naoaki Okazaki, Kentaro Inui
COLING1
2014 An Example-Based Approach to Difficult Pronoun Resolution
Canasai Kruengkrai, Naoya Inoue, Jun Sugiura, Kentaro Inui
PACLIC2
2013 Discriminative Learning of First-Order Weighted Abduction from Partial Discourse Explanations
Kazeto Yamamoto, Naoya Inoue, Yotaro Watanabe, Naoaki Okazaki, Kentaro Inui
CICLing (1)2
2012 Coreference Resolution with ILP-based Weighted Abduction
Naoya Inoue, Ekaterina Ovchinnikova, Kentaro Inui, Jerry R. Hobbs
COLING1
2012 Large-Scale Cost-Based Abduction in Full-Fledged First-Order Predicate Logic with Cutting Plane Inference
Naoya Inoue, Kentaro Inui
JELIA1
2011 Selective injection and laser manipulation of nanotool inside a specific cell using Optical pH regulation and optical tweezers
abstract
We developed Optical pH regulation using functional nanotool impregnated with photo-responsive chemical for selective cell injection of nanotool. The nanotool was modified by fluorescent dye for intracellular measurement. The nanotool was included in the fusogenic liposome. Membrane fusion of the liposome to the cell membrane was used for invasive cell injection of the nanotool. The liposome fuses to the cell in weak acidic condition. Local pH regulation inside the liposome was developed using photochromic chemical for selective cell injection of the nanotool. The nanotool was modified by Leuco crystal violet (LCV). LCV emits the proton by ultraviolet (UV) illumination. The emitted proton decreases the pH value in the liposome. This pH regulation is reversible by UV/VIS illumination. The liposome was manipulated by optical tweezers. After contact of the liposome to the cell, the liposome was adhered to the cell by UV induced membrane fusion. Injected nanotool was manipulated by optical tweezers. Intracellular temperature was detected by measuring the fluorescence intensity from the nanotool. We demonstrated optical pH regulation, selective cell injection of the nanotool, and manipulation of the nanotool in the cell.
Hisataka Maruyama, Naoya Inoue, Taisuke Masuda, Fumihito Arai
ICRA2
2010 Geometrically local isotropic independence and numerical analysis of the Mahalanobis metric in vector space
Joken Son, Naoya Inoue, Yukihiko Yamashita
Pattern Recognit. Lett.2
2008 Numerical analysis of Mahalanobis metric in vector space
abstract
The Mahalanobis metric was proposed by extending the Mahalanobis distance to provide a probabilistic distance for a non-normal distribution. The Mahalanobis metric equation is a nonlinear second order differential equation derived from the equation of geometrically local isotropic independence, which is proposed to define normal distributions in a manifold. In this paper we provide experimental results of calculating the Mahalanobis metric by the Newton-Raphson method. We add error to the original probability density function and calculate the Mahalanobis metric to investigate the effect of the error in a probability density function to the solution.
Joken Son, Naoya Inoue, Yukihiko Yamashita
ICPR2