VLDB 2026 Research / reviewers in the wild / expert
Gaku Morio
dblp:218/0681
· DBLP profile ↗
17ranked-venue papers
9as first author
11since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 9 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing DetectionabstractCompanies spend large amounts of money on public relations campaigns to project a positive brand image.However, sometimes there is a mismatch between what they say and what they do. Oil & gas companies, for example, are accused of "greenwashing" with imagery of climate-friendly initiatives.Understanding the framing, and changes in framing, at scale can help better understand the goals and nature of public relation campaigns.To address this, we introduce a benchmark dataset of expert-annotated video ads obtained from Facebook and YouTube.The dataset provides annotations for 13 framing types for more than 50 companies or advocacy groups across 20 countries.Our dataset is especially designed for the evaluation of vision-language models (VLMs), distinguishing it from past text-only framing datasets.Baseline experiments show some promising results, while leaving room for improvement for future work: GPT-4.1 can detect environmental messages with 79% F1 score, while our best model only achieves 46% F1 score on identifying framing around green innovation.We also identify challenges that VLMs must address, such as implicit framing, handling videos of various lengths, or implicit cultural backgrounds.Our dataset contributes to research in multimodal analysis of strategic communication in the energy sector. Gaku Morio, Harri Rowlands, Dominik Stammbach, Christopher D. Manning, Peter Henderson 0002 |
NeurIPS | 1 |
| 2024 | CHICOT: A Developer-Assistance Toolkit for Code Search with High-Level Contextual InformationabstractWe propose a source code search system named CHICOT (Code search with HIgh level COnText) to assist developers in reusing existing code. While previous studies have examined code search on the basis of code-level, fine-grained specifications such as functionality, logic, or implementation, CHICOT addresses a unique mission: code search with high-level contextual information, such as the purpose or domain of a developer's project. It achieves this feature by first extracting the context information from codebases and then considering this context during the search. It provides a VSCode plugin for daily coding assistance, and the built-in crawler ensures up-to-date code suggestions. The case study attests to the utility of CHICOT in real-world scenarios. Terufumi Morishita, Yuta Koreeda, Atsuki Yamaguchi, Gaku Morio, Osamu Imaichi, Yasuhiro Sogawa |
AAAI | 4 |
| 2024 | JFLD: A Japanese Benchmark for Deductive Reasoning Based on Formal LogicabstractLarge language models (LLMs) have proficiently solved a broad range of tasks with their rich knowledge but often struggle with logical reasoning. To foster the research on logical reasoning, many benchmarks have been proposed so far. However, most of these benchmarks are limited to English, hindering the evaluation of LLMs specialized for each language. To address this, we propose JFLD (Japanese Formal Logic Deduction), a deductive reasoning benchmark for Japanese. JFLD assess whether LLMs can generate logical steps to (dis-)prove a given hypothesis based on a given set of facts. Its key features are assessing pure logical reasoning abilities isolated from knowledge and assessing various reasoning rules. We evaluate various Japanese LLMs and see that they are still poor at logical reasoning, thus highlighting a substantial need for future research. Terufumi Morishita, Atsuki Yamaguchi, Gaku Morio, Hikaru Tomonari, Osamu Imaichi, Yasuhiro Sogawa |
LREC/COLING | 3 |
| 2024 | ReportParse: A Unified NLP Tool for Extracting Document Structure and Semantics of Corporate Sustainability Reporting
Gaku Morio, Soh Young In, Jungah Yoon, Harri Rowlands, Christopher D. Manning |
IJCAI | 1 |
| 2024 | Enhancing Reasoning Capabilities of LLMs via Principled Synthetic Logic CorpusabstractLarge language models (LLMs) are capable of solving a wide range of tasks, yet they have struggled with reasoning.
To address this, we propose $\textbf{Additional Logic Training (ALT)}$, which aims to enhance LLMs' reasoning capabilities by program-generated logical reasoning samples.
We first establish principles for designing high-quality samples by integrating symbolic logic theory and previous empirical insights.
Then, based on these principles, we construct a synthetic corpus named $\textbf{Formal} \ \textbf{Logic} \ \textbf{\textit{D}eduction} \ \textbf{\textit{D}iverse}$ (FLD$ _{\times2}$), comprising numerous samples of multi-step deduction with unknown facts, diverse reasoning rules, diverse linguistic expressions, and challenging distractors.
Finally, we empirically show that ALT on FLD$ _{\times2}$ substantially enhances the reasoning capabilities of state-of-the-art LLMs, including LLaMA-3.1-70B.
Improvements include gains of up to 30 points on logical reasoning benchmarks, up to 10 points on math and coding benchmarks, and 5 points on the benchmark suite BBH. Terufumi Morishita, Gaku Morio, Atsuki Yamaguchi, Yasuhiro Sogawa |
NeurIPS | 2 |
| 2023 | Learning Deductive Reasoning from Synthetic Corpus based on Formal LogicabstractWe study a synthetic corpus based approach for language models (LMs) to acquire logical deductive reasoning ability. The previous studies generated deduction examples using specific sets of deduction rules. However, these rules were limited or otherwise arbitrary. This can limit the generalizability of acquired deductive reasoning ability. We rethink this and adopt a well-grounded set of deduction rules based on formal logic theory, which can derive any other deduction rules when combined in a multistep way. We empirically verify that LMs trained on the proposed corpora, which we name $\textbf{FLD}$ ($\textbf{F}$ormal $\textbf{L}$ogic $\textbf{D}$eduction), acquire more generalizable deductive reasoning ability. Furthermore, we identify the aspects of deductive reasoning ability on which deduction corpora can enhance LMs and those on which they cannot. Finally, on the basis of these results, we discuss the future directions for applying deduction corpora or other approaches for each aspect. We release the code, data, and models. Terufumi Morishita, Gaku Morio, Atsuki Yamaguchi, Yasuhiro Sogawa |
ICML | 2 |
| 2023 | An NLP Benchmark Dataset for Assessing Corporate Climate Policy EngagementabstractAs societal awareness of climate change grows, corporate climate policy engagements are attracting attention.We propose a dataset to estimate corporate climate policy engagement from various PDF-formatted documents.Our dataset comes from LobbyMap (a platform operated by global think tank InfluenceMap) that provides engagement categories and stances on the documents.To convert the LobbyMap data into the structured dataset, we developed a pipeline using text extraction and OCR.Our contributions are: (i) Building an NLP dataset including 10K documents on corporate climate policy engagement. (ii) Analyzing the properties and challenges of the dataset. (iii) Providing experiments for the dataset using pre-trained language models.The results show that while Longformer outperforms baselines and other pre-trained models, there is still room for significant improvement.We hope our work begins to bridge research on NLP and climate change. Gaku Morio, Christopher D. Manning |
NeurIPS | 1 |
| 2022 | Rethinking Fano's Inequality in Ensemble LearningabstractWe propose a fundamental theory on ensemble learning that evaluates a given ensemble system by a well-grounded set of metrics. Previous studies used a variant of Fano’s inequality of information theory and derived a lower bound of the classification error rate on the basis of the accuracy and diversity of models. We revisit the original Fano’s inequality and argue that the studies did not take into account the information lost when multiple model predictions are combined into a final prediction. To address this issue, we generalize the previous theory to incorporate the information loss. Further, we empirically validate and demonstrate the proposed theory through extensive experiments on actual systems. The theory reveals the strengths and weaknesses of systems on each metric, which will push the theoretical understanding of ensemble learning and give us insights into designing systems. Terufumi Morishita, Gaku Morio, Shota Horiguchi, Hiroaki Ozaki, Nobuo Nukaga |
ICML | 2 |
| 2022 | End-to-end Argument Mining with Cross-corpora Multi-task LearningabstractAbstract Mining an argument structure from text is an important step for tasks such as argument search and summarization. While studies on argument(ation) mining have proposed promising neural network models, they usually suffer from a shortage of training data. To address this issue, we expand the training data with various auxiliary argument mining corpora and propose an end-to-end cross-corpus training method called Multi-Task Argument Mining (MT-AM). To evaluate our approach, we conducted experiments for the main argument mining tasks on several well-established argument mining corpora. The results demonstrate that MT-AM generally outperformed the models trained on a single corpus. Also, the smaller the target corpus was, the better the MT-AM performed. Our extensive analyses suggest that the improvement of MT-AM depends on several factors of transferability among auxiliary and target corpora. Gaku Morio, Hiroaki Ozaki, Terufumi Morishita, Kohsuke Yanai |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | i-Parser: Interactive Parser Development Kit for Natural Language ProcessingabstractThis demonstration paper presents i-Parser, a novel development kit that produces high-performance semantic parsers. i-Parser converts training graphs into sequences written in a context-free language, then our proposed model learns to generate the sequences. With interactive configuration and visualization, users can easily build their own parsers. Benchmark results of i-Parser showed high performances of various parsing tasks in natural language processing. Gaku Morio, Hiroaki Ozaki, Yuta Koreeda, Terufumi Morishita, Toshinori Miyoshi |
AAAI | 1 |
| 2021 | Project-then-Transfer: Effective Two-stage Cross-lingual Transfer for Semantic Dependency ParsingabstractThis paper describes the first report on crosslingual transfer for semantic dependency parsing.We present the insight that there are two different kinds of cross-linguality, namely surface level and semantic level, and try to capture both kinds of cross-linguality by combining annotation projection and model transfer of pre-trained language models.Our experiments showed that the performance of our graph-based semantic dependency parser almost achieved the approximated upper bound. Hiroaki Ozaki, Gaku Morio, Terufumi Morishita, Toshinori Miyoshi |
EACL | 2 |
| 2020 | Towards Better Non-Tree Argument Mining: Proposition-Level Biaffine Parsing with Task-Specific ParameterizationabstractState-of-the-art argument mining studies have advanced the techniques for predicting argument structures. However, the technology for capturing non-tree-structured arguments is still in its infancy. In this paper, we focus on non-tree argument mining with a neural model. We jointly predict proposition types and edges between propositions. Our proposed model incorporates (i) task-specific parameterization (TSP) that effectively encodes a sequence of propositions and (ii) a proposition-level biaffine attention (PLBA) that can predict a non-tree argument consisting of edges. Experimental results show that both TSP and PLBA boost edge prediction performance compared to baselines. Gaku Morio, Hiroaki Ozaki, Terufumi Morishita, Yuta Koreeda, Kohsuke Yanai |
ACL | 1 |
| 2020 | Corpus for Modeling User Interactions in Online Persuasive DiscussionsabstractPersuasions are common in online arguments such as discussion forums. To analyze persuasive strategies, it is important to understand how individuals construct posts and comments based on the semantics of the argumentative components. In addition to understanding how we construct arguments, understanding how a user post interacts with other posts (i.e., argumentative inter-post relation) still remains a challenge. Therefore, in this study, we developed a novel annotation scheme and corpus that capture both user-generated inner-post arguments and inter-post relations between users in ChangeMyView, a persuasive forum. Our corpus consists of arguments with 4612 elementary units (EUs) (i.e., propositions), 2713 EU-to-EU argumentative relations, and 605 inter-post argumentative relations in 115 threads. We analyzed the annotated corpus to identify the characteristics of online persuasive arguments, and the results revealed persuasive documents have more claims than non-persuasive ones and different interaction patterns among persuasive and non-persuasive documents. Our corpus can be used as a resource for analyzing persuasiveness and training an argument mining system to identify and extract argument structures. The annotated corpus and annotation guidelines have been made publicly available. Ryo Egawa, Gaku Morio, Katsuhide Fujita |
LREC | 2 |
| 2019 | On the Role of Syntactic Graph Convolutions for Identifying and Classifying Argument ComponentsabstractThis paper focuses on fundamental research that combines syntactic knowledge with neural studies, which utilize syntactic information in argument component identification and classification (AC-I/C) tasks in argument mining (AM). The following are our paper’s contributions: 1) We propose a way of incorporating a syntactic GCN into multi-task learning models for AC-I/C tasks. 2) We demonstrate the valid effectiveness of our proposed syntactic GCN in fair experiments in some datasets. We also found that syntactic GCNs are promising for lexically independent scenarios. Our code in the experiments is available for reproducibility.1 Gaku Morio, Katsuhide Fujita |
AAAI | 1 |
| 2019 | Revealing and Predicting Online Persuasion Strategy with Elementary UnitsabstractGaku Morio, Ryo Egawa, Katsuhide Fujita. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Gaku Morio, Ryo Egawa, Katsuhide Fujita |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Can You Give Me a Reason?: Argument-inducing Online Forum by Argument MiningabstractThis demonstration paper presents an argument-inducing online forum that stimulates participants with lack of premises for their claim in online discussions. The proposed forum provides its participants the following two subsystems: (1) Argument estimator for online discussions automatically generates a visualization of the argument structures in posts based on argument mining. The forum indicates structures such as claim-premise relations in real time by exploiting a state-of-the-art deep learning model. (2) Argument-inducing agent for online discussion (AIAD) automatically generates a reply post based on the argument estimator requesting further reasons to improve the argumentation of participants. Makiko Ida, Gaku Morio, Kosui Iwasa, Tomoyuki Tatsumi, Takaki Yasui, Katsuhide Fujita |
WWW | 2 |
| 2018 | Annotating Online Civic Discussion Threads for Argument MiningabstractArgument mining techniques have become popular in online civic discussion thread analysis to understand an enormous amount of posts and flow of discussions for consensus building. However, the existing corpora and discussion thread analysis haven't discussed argument mining schemes sufficiently. This paper proposes a novel scheme for discussion thread analysis, annotates online civic discussions, and analyzes the annotated corpus. Our scheme consists of novel inner-and inter-post schemes. The inner-post scheme considers a post as a stand-alone discourse in a thread. We perform a micro-level annotation of argument components and relations in a post. The inter-post scheme provides a micro-level inter-post interaction to capture the argumentative reply-to relation. As a result, we have an annotated corpus including 399 threads and 5559 sentences of 204 citizens that is valid and argumentative. In addition, we analyze the annotated corpus to demonstrate statistical and linguistic properties of the corpus. Gaku Morio |
WI | 1 |