Sho Yokoi

dblp:184/8316 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
YearPublicationVenuePosition
2026 Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings
abstract
For constructing text embeddings, mean pooling, which averages token embeddings, is the standard approach.This paper examines whether mean pooling actually works well in real text encoders.First, we note that mean pooling can collapse information beyond the first-order statistics of the token embeddings, such as second-order statistics that capture their spatial structure, potentially mapping distinct token embedding distributions to similar text embeddings.Motivated by this concern, we propose a simple metric to quantify such a collapse induced by mean pooling.Then, using this metric, we empirically measure how often this collapse arises in actual models and texts, and find that mean pooling works well in modern text encoders.In particular, this collapse is less likely to arise in contrastive fine-tuned text encoders than in their pretrained backbone models.We also find that the robustness of these text encoders to collapse stems from the concentration of token embeddings within each text.In addition, we find that robustness to this collapse, as quantified by our proposed metric, correlates with downstream task performance.Overall, our findings help explain why modern text encoders remain effective despite relying on seemingly coarse mean pooling.
Tomomasa Hara, Hiroto Kurita, Masaaki Imaizumi, Kentaro Inui, Sho Yokoi
ACL (1)5
2025 Quantifying Lexical Semantic Shift via Unbalanced Optimal Transport
abstract
Lexical semantic change detection aims to identify shifts in word meanings over time. While existing methods using embeddings from a diachronic corpus pair estimate the degree of change for target words, they offer limited insight into changes at the level of individual usage instances. To address this, we apply Unbalanced Optimal Transport (UOT) to sets of contextualized word embeddings, capturing semantic change through the excess and deficit in the alignment between usage instances. In particular, we propose Sense Usage Shift (SUS), a measure that quantifies changes in the usage frequency of a word sense at each usage instance. By leveraging SUS, we demonstrate that several challenges in semantic change detection can be addressed in a unified manner, including quantifying instance-level semantic change and word-level tasks such as measuring the magnitude of semantic change and the broadening or narrowing of meaning.
Ryo Kishino, Hiroaki Yamagiwa, Ryo Nagata, Sho Yokoi, Hidetoshi Shimodaira
ACL (1)4
2025 SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches
abstract
Researchers and practitioners in natural language processing and computational linguistics frequently observe and analyze the real language usage in large-scale corpora. For that purpose, they often employ off-the-shelf pattern-matching tools, such as grep, and keyword-in-context concordancers, which is widely used in corpus linguistics for gathering examples. Nonetheless, these existing techniques rely on surface-level string matching, and thus they suffer from the major limitation of not being able to handle orthographic variations and paraphrasing---notable and common phenomena in any natural language. In addition, existing continuous approaches such as dense vector search tend to be overly coarse, often retrieving texts that are unrelated but share similar topics. Given these challenges, we propose a novel algorithm that achieves soft (or semantic) yet efficient pattern matching by relaxing a surface-level matching with word embeddings. Our algorithm is highly scalable with respect to the size of the corpus text utilizing inverted indexes. We have prepared an efficient implementation, and we provide an accessible web tool. Our experiments demonstrate that the proposed method (i) can execute searches on billion-scale corpora in less than a second, which is comparable in speed to surface-level string matching and dense vector search; (ii) can extract harmful instances that semantically match queries from a large set of English and Japanese Wikipedia articles; and (iii) can be effectively applied to corpus-linguistic analyses of Latin, a language with highly diverse inflections.
Hiroyuki Deguchi 0002, Go Kamoda, Yusuke Matsushita 0002, Chihiro Taguchi, Kohei Suenaga, Masaki Waga, Sho Yokoi
ICLR7
2025 TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
abstract
Causal language models have demonstrated remarkable capabilities, but their size poses significant challenges for deployment in resource-constrained environments. Knowledge distillation, a widely-used technique for transferring knowledge from a large teacher model to a small student model, presents a promising approach for model compression. A significant remaining issue lies in the major differences between teacher and student models, namely the substantial capacity gap, mode averaging, and mode collapse, which pose barriers during distillation. To address these issues, we introduce $\textit{Temporally Adaptive Interpolated Distillation (TAID)}$, a novel knowledge distillation approach that dynamically interpolates student and teacher distributions through an adaptive intermediate distribution, gradually shifting from the student's initial distribution towards the teacher's distribution. We provide a theoretical analysis demonstrating TAID's ability to prevent mode collapse and empirically show its effectiveness in addressing the capacity gap while balancing mode averaging and mode collapse. Our comprehensive experiments demonstrate TAID's superior performance across various model sizes and architectures in both instruction tuning and pre-training scenarios. Furthermore, we showcase TAID's practical impact by developing two state-of-the-art compact foundation models: $\texttt{TAID-LLM-1.5B}$ for language tasks and $\texttt{TAID-VLM-2B}$ for vision-language tasks. These results demonstrate TAID's effectiveness in creating high-performing and efficient models, advancing the development of more accessible AI technologies.
Makoto Shing, Kou Misaki, Sho Yokoi, Takuya Akiba
ICLR4
2024 Analyzing Feed-Forward Blocks in Transformers through the Lens of Attention Maps
abstract
Transformers are ubiquitous in wide tasks. Interpreting their internals is a pivotal goal. Nevertheless, their particular components, feed-forward (FF) blocks, have typically been less analyzed despite their substantial parameter amounts. We analyze the input contextualization effects of FF blocks by rendering them in the attention maps as a human-friendly visualization scheme. Our experiments with both masked- and causal-language models reveal that FF networks modify the input contextualization to emphasize specific types of linguistic compositions. In addition, FF and its surrounding components tend to cancel out each other's effects, suggesting potential redundancy in the processing of the Transformer layer.
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro Inui
ICLR3
2024 Subspace Representations for Soft Set Operations and Sentence Similarities
abstract
Yoichi Ishibashi, Sho Yokoi, Katsuhito Sudoh, Satoshi Nakamura. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yoichi Ishibashi, Sho Yokoi, Katsuhito Sudoh, Satoshi Nakamura 0001
NAACL-HLT2
2024 Zipfian Whitening
abstract
The word embedding space in neural models is skewed, and correcting this can improve task performance. We point out that most approaches for modeling, correcting, and measuring the symmetry of an embedding space implicitly assume that the word frequencies are *uniform*; in reality, word frequencies follow a highly non-uniform distribution, known as *Zipf's law*. Surprisingly, simply performing PCA whitening weighted by the empirical word frequency that follows Zipf's law significantly improves task performance, surpassing established baselines. From a theoretical perspective, both our approach and existing methods can be clearly categorized: word representations are distributed according to an exponential family with either uniform or Zipfian base measures. By adopting the latter approach, we can naturally emphasize informative low-frequency words in terms of their vector norm, which becomes evident from the information-geometric perspective (Oyama et al., EMNLP 2023), and in terms of the loss functions for imbalanced classification (Menon et al. ICLR 2021). Additionally, our theory corroborates that popular natural language processing methods, such as skip-gram negative sampling (Mikolov et al., NIPS 2013), WhiteningBERT (Huang et al., Findings of EMNLP 2021), and headless language models (Godey et al., ICLR 2024), work well just because their word embeddings encode the empirical word frequency into the underlying probabilistic model.
Sho Yokoi, Han Bao 0002, Hiroto Kurita, Hidetoshi Shimodaira
NeurIPS1
2023 Unbalanced Optimal Transport for Unbalanced Word Alignment
abstract
Monolingual word alignment is crucial to model semantic interactions between sentences.In particular, null alignment, a phenomenon in which words have no corresponding counterparts, is pervasive and critical in handling semantically divergent sentences.Identification of null alignment is useful on its own to reason about the semantic similarity of sentences by indicating there exists information inequality.To achieve unbalanced word alignment that values both alignment and null alignment, this study shows that the family of optimal transport (OT), i.e., balanced, partial, and unbalanced OT, are natural and powerful approaches even without tailor-made techniques.Our extensive experiments covering unsupervised and supervised settings indicate that our generic OT-based alignment methods are competitive against the state-of-the-arts specially designed for word alignment, remarkably on challenging datasets with high null alignment frequencies.
Yuki Arase, Han Bao 0002, Sho Yokoi
ACL (1)3
2023 Norm of Word Embedding Encodes Information Gain
abstract
Distributed representations of words encode lexical semantic information, but what type of information is encoded and how?Focusing on the skip-gram with negative-sampling method, we found that the squared norm of static word embedding encodes the information gain conveyed by the word; the information gain is defined by the Kullback-Leibler divergence of the co-occurrence distribution of the word to the unigram distribution.Our findings are explained by the theoretical framework of the exponential family of probability distributions and confirmed through precise experiments that remove spurious correlations arising from word frequency.This theory also extends to contextualized word embeddings in language models or any neural networks with the softmax output layer.We also demonstrate that both the KL divergence and the squared norm of embedding provide a useful metric of the informativeness of a word in tasks such as keyword extraction, proper-noun discrimination, and hypernym discrimination.
Momose Oyama, Sho Yokoi, Hidetoshi Shimodaira
EMNLP2
2021 Incorporating Residual and Normalization Layers into Analysis of Masked Language Models
abstract
Transformer architecture has become ubiquitous in the natural language processing field.To interpret the Transformer-based models, their attention patterns have been extensively analyzed.However, the Transformer architecture is not only composed of the multihead attention; other components can also contribute to Transformers' progressive performance.In this study, we extended the scope of the analysis of Transformers from solely the attention patterns to the whole attention block, i.e., multi-head attention, residual connection, and layer normalization.Our analysis of Transformer-based masked language models shows that the token-to-token interaction performed via attention has less impact on the intermediate representations than previously assumed.These results provide new intuitive explanations of existing reports; for example, discarding the learned attention patterns tends not to adversely affect the performance.The codes of our experiments are publicly available.
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro Inui
EMNLP (1)3
2021 Evaluation of Similarity-based Explanations
Kazuaki Hanawa, Sho Yokoi, Satoshi Hara 0001, Kentaro Inui
ICLR2
2021 Instance-Based Neural Dependency Parsing
abstract
Abstract Interpretable rationales for model predictions are crucial in practical applications. We develop neural models that possess an interpretable inference process for dependency parsing. Our models adopt instance-based inference, where dependency edges are extracted and labeled by comparing them to edges in a training set. The training edges are explicitly used for the predictions; thus, it is easy to grasp the contribution of each edge to the predictions. Our experiments show that our instance-based models achieve competitive accuracy with standard neural models and have the reasonable plausibility of instance-based explanations.
Hiroki Ouchi, Jun Suzuki 0001, Sosuke Kobayashi, Sho Yokoi, Tatsuki Kuribayashi, Masashi Yoshikawa, Kentaro Inui
Trans. Assoc. Comput. Linguistics4
2020 Instance-Based Learning of Span Representations: A Case Study through Named Entity Recognition
abstract
Interpretable rationales for model predictions play a critical role in practical applications.In this study, we develop models possessing interpretable inference process for structured prediction.Specifically, we present a method of instance-based learning that learns similarities between spans.At inference time, each span is assigned a class label based on its similar spans in the training set, where it is easy to understand how much each training instance contributes to the predictions.Through empirical analysis on named entity recognition, we demonstrate that our method enables to build models that have high interpretability without sacrificing performance.
Hiroki Ouchi, Jun Suzuki 0001, Sosuke Kobayashi, Sho Yokoi, Tatsuki Kuribayashi, Ryuto Konno, Kentaro Inui
ACL4
2020 Modeling Event Salience in Narratives via Barthes' Cardinal Functions
abstract
Events in a narrative differ in salience: some are more important to the story than others.Estimating event salience is useful for tasks such as story generation, and as a tool for text analysis in narratology and folkloristics.To compute event salience without any annotations, we adopt Barthes' definition of event salience and propose several unsupervised methods that require only a pre-trained language model.Evaluating the proposed methods on folktales with event salience annotation, we show that the proposed methods outperform baseline methods and find fine-tuning a language model on narrative texts is a key factor in improving the proposed methods.
Takaki Otake, Sho Yokoi, Naoya Inoue, Tatsuki Kuribayashi, Kentaro Inui
COLING2
2020 Filtering Noisy Dialogue Corpora by Connectivity and Content Relatedness
abstract
Large-scale dialogue datasets have recently become available for training neural dialogue agents. However, these datasets have been reported to contain a non-negligible number of unacceptable utterance pairs. In this paper, we propose a method for scoring the quality of utterance pairs in terms of their connectivity and relatedness. The proposed scoring method is designed based on findings widely shared in the dialogue and linguistics research communities. We demonstrate that it has a relatively good correlation with the human judgment of dialogue quality. Furthermore, the method is applied to filter out potentially unacceptable utterance pairs from a large-scale noisy dialogue corpus to ensure its quality. We experimentally confirm that training data filtered by the proposed method improves the quality of neural dialogue agents in response generation.
Reina Akama, Sho Yokoi, Jun Suzuki 0001, Kentaro Inui
EMNLP (1)2
2020 Attention is Not Only a Weight: Analyzing Transformers with Vector Norms
abstract
Attention is a key component of Transformers, which have recently achieved considerable success in natural language processing. Hence, attention is being extensively studied to investigate various linguistic capabilities of Transformers, focusing on analyzing the parallels between attention weights and specific linguistic phenomena. This paper shows that attention weights alone are only one of the two factors that determine the output of attention and proposes a norm-based analysis that incorporates the second factor, the norm of the transformed input vectors. The findings of our norm-based analyses of BERT and a Transformer-based neural machine translation system include the following: (i) contrary to previous studies, BERT pays poor attention to special tokens, and (ii) reasonable word alignment can be extracted from attention mechanisms of Transformer. These findings provide insights into the inner workings of Transformers.
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro Inui
EMNLP (1)3
2020 Word Rotator's Distance
abstract
A key principle in assessing textual similarity is measuring the degree of semantic overlap between two texts by considering the word alignment.Such alignment-based approaches are intuitive and interpretable; however, they are empirically inferior to the simple cosine similarity between general-purpose sentence vectors.To address this issue, we focus on and demonstrate the fact that the norm of word vectors is a good proxy for word importance, and their angle is a good proxy for word similarity.Alignment-based approaches do not distinguish them, whereas sentence-vector approaches automatically use the norm as the word importance.Accordingly, we propose a method that first decouples word vectors into their norm and direction, and then computes alignment-based similarity using earth mover's distance (i.e., optimal transport cost), which we refer to as word rotator's distance.Besides, we find how to "grow" the norm and direction of word vectors (vector converter), which is a new systematic approach derived from sentence-vector estimation methods.On several textual similarity datasets, the combination of these simple proposed methods outperformed not only alignment-based approaches but also strong baselines.1
Sho Yokoi, Reina Akama, Jun Suzuki 0001, Kentaro Inui
EMNLP (1)1
2018 Pointwise HSIC: A Linear-Time Kernelized Co-occurrence Norm for Sparse Linguistic Expressions
abstract
In this paper, we propose a new kernel-based co-occurrence measure that can be applied to sparse linguistic expressions (e.g., sentences) with a very short learning time, as an alternative to pointwise mutual information (PMI).As well as deriving PMI from mutual information, we derive this new measure from the Hilbert-Schmidt independence criterion (HSIC); thus, we call the new measure the pointwise HSIC (PHSIC).PHSIC can be interpreted as a smoothed variant of PMI that allows various similarity metrics (e.g., sentence embeddings) to be plugged in as kernels.Moreover, PHSIC can be estimated by simple and fast (linear in the size of the data) matrix calculations regardless of whether we use linear or nonlinear kernels.Empirically, in a dialogue response selection task, PHSIC is learned thousands of times faster than an RNNbased PMI while outperforming PMI in accuracy.In addition, we also demonstrate that PH-SIC is beneficial as a criterion of a data selection task for machine translation owing to its ability to give high (low) scores to a consistent (inconsistent) pair with other pairs.
Sho Yokoi, Sosuke Kobayashi, Kenji Fukumizu, Jun Suzuki 0001, Kentaro Inui
EMNLP1
2017 Learning Co-Substructures by Kernel Dependence Maximization
abstract
Modeling associations between items in a dataset is a problem that is frequently encountered in data and knowledge mining research. Most previous studies have simply applied a predefined fixed pattern for extracting the substructure of each item pair and then analyzed the associations between these substructures. Using such fixed patterns may not, however, capture the significant association. We, therefore, propose the novel machine learning task of extracting a strongly associated substructure pair (co-substructure) from each input item pair. We call this task dependent co-substructure extraction (DCSE), and formalize it as a dependence maximization problem. Then, we discuss critical issues with this task: the data sparsity problem and a huge search space. To address the data sparsity problem, we adopt the Hilbert--Schmidt independence criterion as an objective function. To improve search efficiency, we adopt the Metropolis--Hastings algorithm. We report the results of empirical evaluations, in which the proposed method is applied for acquiring and predicting narrative event pairs, an active task in the field of natural language processing.
Sho Yokoi, Daichi Mochihashi, Naoaki Okazaki, Kentaro Inui
IJCAI1
2016 Link Prediction by Incidence Matrix Factorization
abstract
Link prediction suffers from the data sparsity problem. This paper presents and validates our hypothesis that, for sparse networks, incidence matrix factorization (IMF) could perform better than adjacency matrix factorization (AMF), which has been used in many previous studies. A key observation supporting the hypothesis is that IMF models a partially-observed graph more accurately than AMF. A technical challenge for validating our hypothesis is that, unlike AMF approach, there does not exist an obvious method to make predictions using a factorized incidence matrix. To this end, we newly develop an optimization-based link prediction method adopting IMF. We have conducted thorough experiments using synthetic and real-world datasets to investigate the relationship between the sparsity of a network and the performance of the aforementioned two methods. The experimental results show that IMF performs better than AMF as networks become sparser, which strongly validates our hypothesis.
Sho Yokoi, Hiroshi Kajino, Hisashi Kashima
ECAI1