EDBT 2026 Demo / reviewers in the wild / expert
Hitomi Yanaka
dblp:187/5994
· DBLP profile ↗
18ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0003-0354-6116ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 6 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-Grained Analysis of Shared Syntactic Mechanisms in Language ModelsabstractWhile language models demonstrate sophisticated syntactic capabilities, the extent to which their internal mechanisms align with crossconstructional principles studied in linguistics remains poorly understood.This study investigates whether models employ shared neural mechanisms across different syntactic constructions by applying causal interpretability methods at a granular level.Focusing on filler-gap dependencies and negative polarity item (NPI) licensing, we utilize activation patching to identify the functional roles of specific attention heads and MLP blocks.Our results reveal a highly localized and shared mechanism for filler-gap dependencies located in the early to middle layers, whereas NPI processing exhibits no such unified mechanism.Furthermore, we find that these mechanisms identified by activation patching generalize to out-of-distribution, while distributed alignment search, a supervised interpretability method, is susceptible to overfitting on narrow linguistic distributions.Finally, we validate our findings by demonstrating that the manipulation of the identified components improves model performance on acceptability judgment benchmarks. 1 Ryoma Kumon, Hitomi Yanaka |
ACL (1) | 2 |
| 2026 | J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language ModelingabstractSpoken dialogue is essential for human-AI interactions, providing expressive capabilities beyond text. Developing effective spoken dialogue systems (SDSs) requires large-scale, high-quality, and diverse spoken dialogue corpora. However, existing datasets are often limited in size, spontaneity, or linguistic coherence. To address these limitations, we introduce J-CHAT, a 76,000-hour open-source Japanese spoken dialogue corpus. Constructed using an automated, language-independent methodology, J-CHAT ensures acoustic cleanliness, diversity, and natural spontaneity. The corpus is built from YouTube and podcast data, with extensive filtering and denoising to enhance quality. Experimental results with generative spoken dialogue language models trained on J-CHAT demonstrate its effectiveness for SDS development. By providing a robust foundation for training advanced dialogue models, we anticipate that J-CHAT will drive progress in human-AI dialogue research and applications. Wataru Nakata, Kentaro Seki, Hitomi Yanaka, Yuki Saito 0001, Shinnosuke Takamichi, Hiroshi Saruwatari |
LREC | 3 |
| 2026 | Seeing the Other Side: Diagnostic Tasks for Viewpoint Reasoning in Vision-Language Models
Makoto Takenaka, Hitomi Yanaka |
LREC | 2 |
| 2026 | Developing a Guideline for the Labovian-Structural Analysis of Oral Narratives in Japanese
Amane Watahiki, Tomoki Doi, Akari Kikuchi, Hiroshi Ohata, Yuki I. Nakata, Takuya Niikawa, Taiga Shinozaki, Hitomi Yanaka |
LREC | 8 |
| 2025 | Do Large Vision-Language Models Distinguish between the Actual and Apparent Features of Illusions?
Taiga Shinozaki, Tomoki Doi, Amane Watahiki, Satoshi Nishida, Hitomi Yanaka |
CogSci | 5 |
| 2025 | Bridging Perception and Language: A Systematic Benchmark for LVLMs' Understanding of Amodal Completion Reports
Amane Watahiki, Tomoki Doi, Taiga Shinozaki, Satoshi Nishida, Takuya Niikawa, Katsunori Miyahara, Hitomi Yanaka |
CogSci | 7 |
| 2025 | Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese DatasetabstractLarge language models (LLMs) exhibit social biases, prompting the development of various debiasing methods.However, debiasing methods may degrade the capabilities of LLMs.Previous research has evaluated the impact of bias mitigation primarily through tasks measuring general language understanding, which are often unrelated to social biases.In contrast, cultural commonsense is closely related to social biases, as both are rooted in social norms and values.The impact of bias mitigation on cultural commonsense in LLMs has not been well investigated.Considering this gap, we propose SOBACO (SOcial BiAs and Cultural cOmmonsense benchmark), a Japanese benchmark designed to evaluate social biases and cultural commonsense in LLMs in a unified format.We evaluate several LLMs on SOBACO to examine how debiasing methods affect cultural commonsense in LLMs.Our results reveal that the debiasing methods degrade the performance of the LLMs on the cultural commonsense task (up to 75% accuracy deterioration).These results highlight the importance of developing debiasing methods that consider the trade-off with cultural commonsense to improve fairness and utility of LLMs.Warning: This paper contains examples of social biases that can be offensive. Taisei Yamamoto, Ryoma Kumon, Danushka Bollegala, Hitomi Yanaka |
EMNLP | 4 |
| 2025 | Analyzing the Inner Workings of Transformers in Compositional GeneralizationabstractRyoma Kumon, Hitomi Yanaka. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ryoma Kumon, Hitomi Yanaka |
NAACL (Long Papers) | 2 |
| 2024 | Visual-Textual Entailment with Quantities Using Model Checking and Knowledge InjectionabstractIn recent years, there has been great interest in multimodal inference. We concentrate on visual-textual entailment (VTE), a critical task in multimodal inference. VTE is the task of determining entailment relations between an image and a sentence. Several deep learning-based approaches have been proposed for VTE, but current approaches struggle with accurately handling quantities. On the other hand, one promising approach, one based on logical inference that can successfully deal with large quantities, has also been proposed. However, that approach uses automated theorem provers, increasing the computational cost for problems involving many entities. In addition, that approach cannot deal well with lexical differences between the semantic representations of images and sentences. In this paper, we present a logic-based VTE system that overcomes these drawbacks, using model checking for inference to increase efficiency and knowledge injection to perform more robust inference. We create a VTE dataset containing quantities and negation to assess how well VTE systems understand such phenomena. Using this dataset, we demonstrate that our system solves VTE tasks with quantities and negation more robustly than previous approaches. Nobuyuki Iokawa, Hitomi Yanaka |
LREC/COLING | 2 |
| 2024 | Exploring Intra and Inter-language Consistency in Embeddings with ICAabstractWord embeddings represent words as multidimensional real vectors, facilitating data analysis and processing, but are often challenging to interpret.Independent Component Analysis (ICA) creates clearer semantic axes by identifying independent key features.Previous research has shown ICA's potential to reveal universal semantic axes across languages.However, it lacked verification of the consistency of independent components within and across languages.We investigated the consistency of semantic axes in two ways: both within a single language and across multiple languages.We first probed into intra-language consistency, focusing on the reproducibility of axes by running the ICA algorithm multiple times and clustering the outcomes.Then, we statistically examined inter-language consistency by verifying those axes' correspondences using statistical tests.We newly applied statistical methods to establish a robust framework that ensures the reliability and universality of semantic axes. Rongzhi Li, Takeru Matsuda, Hitomi Yanaka |
EMNLP | 3 |
| 2024 | On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific NeuronsabstractTakeshi Kojima, Itsuki Okimura, Yusuke Iwasawa, Hitomi Yanaka, Yutaka Matsuo. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Takeshi Kojima, Itsuki Okimura, Yusuke Iwasawa, Hitomi Yanaka, Yutaka Matsuo |
NAACL-HLT | 4 |
| 2022 | Compositional Evaluation on Japanese Textual Entailment and SimilarityabstractAbstract Natural Language Inference (NLI) and Semantic Textual Similarity (STS) are widely used benchmark tasks for compositional evaluation of pre-trained language models. Despite growing interest in linguistic universals, most NLI/STS studies have focused almost exclusively on English. In particular, there are no available multilingual NLI/STS datasets in Japanese, which is typologically different from English and can shed light on the currently controversial behavior of language models in matters such as sensitivity to word order and case particles. Against this background, we introduce JSICK, a Japanese NLI/STS dataset that was manually translated from the English dataset SICK. We also present a stress-test dataset for compositional inference, created by transforming syntactic structures of sentences in JSICK to investigate whether language models are sensitive to word order and case particles. We conduct baseline experiments on different pre-trained language models and compare the performance of multilingual models when applied to Japanese and other languages. The results of the stress-test experiments suggest that the current pre-trained language models are insensitive to word order and case marking. Hitomi Yanaka, Koji Mineshima |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | Exploring Transitivity in Neural NLI Models through VeridicalityabstractDespite the recent success of deep neural networks in natural language processing, the extent to which they can demonstrate human-like generalization capacities for natural language understanding remains unclear.We explore this issue in the domain of natural language inference (NLI), focusing on the transitivity of inference relations, a fundamental property for systematically drawing inferences.A model capturing transitivity can compose basic inference patterns and draw new inferences.We introduce an analysis method using synthetic and naturalistic NLI datasets involving clauseembedding verbs to evaluate whether models can perform transitivity inferences composed of veridical inferences and arbitrary inference types.We find that current NLI models do not perform consistently well on transitivity inference tasks, suggesting that they lack the generalization capacity for drawing composite inferences from provided training examples.The data and code for our analysis are publicly available at https://github.com/ verypluming/transitivity. Hitomi Yanaka, Koji Mineshima, Kentaro Inui |
EACL | 1 |
| 2020 | Do Neural Models Learn Systematicity of Monotonicity Inference in Natural Language?abstractDespite the success of language models using neural networks, it remains unclear to what extent neural models have the generalization ability to perform inferences.In this paper, we introduce a method for evaluating whether neural models can learn systematicity of monotonicity inference in natural language, namely, the regularity for performing arbitrary inferences with generalization on composition.We consider four aspects of monotonicity inferences and test whether the models can systematically interpret lexical and logical phenomena on different training/test splits.A series of experiments show that three neural models systematically draw inferences on unseen combinations of lexical and logical phenomena when the syntactic structures of the sentences are similar between the training and test sets.However, the performance of the models significantly decreases when the structures are slightly changed in the test set while retaining all vocabularies and constituents already appearing in the training set.This indicates that the generalization ability of neural models is limited to cases where the syntactic structures are nearly the same as those in the training set.(1) P : Some [puppies ↑] ran.H: Some dogs ran.(2) P : No [cats ↓] ran.H: No small cats ran.(3) P : Some [puppies which chased no [cats ↓]] ran.H: Some dogs which chased no small cats ran.(5) P : Some small dogs ran ⇒ H: Some dogs ran (6) P : Several dogs ran ⇒ H: Several animals ran (7) P : No animals ran ⇒ H: No dogs ran (8) P : Several small dogs ran ⇒ H: Several dogs ran (9) P : No dogs ran ⇒ H: No small dogs ranHere, we consider a set of inferences D Q,R Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro Inui |
ACL | 1 |
| 2018 | Neural sentence generation from formal semanticsabstractSequence-to-sequence models have shown strong performance in a wide range of NLP tasks, yet their applications to sentence generation from logical representations are underdeveloped.In this paper, we present a sequence-to-sequence model for generating sentences from logical meaning representations based on event semantics.We use a semantic parsing system based on Combinatory Categorial Grammar (CCG) to obtain data annotated with logical formulas.We augment our sequence-to-sequence model with masking for predicates to constrain output sentences.We also propose a novel evaluation method for generation using Recognizing Textual Entailment (RTE).Combining parsing and generation, we test whether or not the output sentence entails the original text and vice versa.Experiments showed that our model outperformed a baseline with respect to both BLEU scores and accuracies in RTE. Kana Manome, Masashi Yoshikawa, Hitomi Yanaka, Pascual Martínez-Gómez, Koji Mineshima, Daisuke Bekki |
INLG | 3 |
| 2018 | Acquisition of Phrase Correspondences Using Natural Deduction ProofsabstractHitomi Yanaka, Koji Mineshima, Pascual Martínez-Gómez, Daisuke Bekki. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Hitomi Yanaka, Koji Mineshima, Pascual Martínez-Gómez, Daisuke Bekki |
NAACL-HLT | 1 |
| 2017 | Determining Semantic Textual Similarity using Natural Deduction ProofsabstractDetermining semantic textual similarity is a core research subject in natural language processing.Since vector-based models for sentence representation often use shallow information, capturing accurate semantics is difficult.By contrast, logical semantic representations capture deeper levels of sentence semantics, but their symbolic nature does not offer graded notions of textual similarity.We propose a method for determining semantic textual similarity by combining shallow features with features extracted from natural deduction proofs of bidirectional entailment relations between sentence pairs.For the natural deduction proofs, we use ccg2lambda, a higherorder automatic inference system, which converts Combinatory Categorial Grammar (CCG) derivation trees into semantic representations and conducts natural deduction proofs.Experiments show that our system was able to outperform other logicbased systems and that features derived from the proofs are effective for learning textual similarity. Hitomi Yanaka, Koji Mineshima, Pascual Martínez-Gómez, Daisuke Bekki |
EMNLP | 1 |
| 2016 | Clustering Documents on Case Vectors Represented by Predicate-argument Structures - Applied for Eliciting Technological Problems from PatentsabstractPatent analysis is useful to understand the trends of technological problems and develop strategies for technologies.Here patent classification is a method to support the analysis.The purpose of this study is to propose a method for patent classification, with the use of hierarchical clustering based on the structural similarity of problems to be solved.The structural similarity can be calculated with case vectors based on predicate-argument structures of the contents of the patents.The interview survey indicated that this classification plays an essential role in analogical problem solving, by allowing visualization of similar technological problems. Hitomi Yanaka, Yukio Ohsawa |
FedCSIS | 1 |