EDBT 2026 Demo / reviewers in the wild / expert
Hiroto Kurita
dblp:02/4775
· DBLP profile ↗
3ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Representation and self-supervised learning · 91% Language models and text generation · 9% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
second-order statistics |
1.0 | 1 | 2026 | Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings · ACL (1) 2026 |
Machine learning › Representation and self-supervised learning
text embedding |
1.0 | 1 | 2026 | Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings · ACL (1) 2026 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.8 | 1 | 2024 | Zipfian Whitening · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.3 | 1 | 2026 | Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings · ACL (1) 2026 |
Natural language and speech › Language models and text generation › text representation
text encoder |
0.3 | 1 | 2026 | Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
mean pooling · 1.0contrastive fine-tuning · 1.0information geometry · 0.8PCA whitening · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Why Mean Pooling Works: Quantifying Second-Order Collapse in Text EmbeddingsabstractFor constructing text embeddings, mean pooling, which averages token embeddings, is the standard approach.This paper examines whether mean pooling actually works well in real text encoders.First, we note that mean pooling can collapse information beyond the first-order statistics of the token embeddings, such as second-order statistics that capture their spatial structure, potentially mapping distinct token embedding distributions to similar text embeddings.Motivated by this concern, we propose a simple metric to quantify such a collapse induced by mean pooling.Then, using this metric, we empirically measure how often this collapse arises in actual models and texts, and find that mean pooling works well in modern text encoders.In particular, this collapse is less likely to arise in contrastive fine-tuned text encoders than in their pretrained backbone models.We also find that the robustness of these text encoders to collapse stems from the concentration of token embeddings within each text.In addition, we find that robustness to this collapse, as quantified by our proposed metric, correlates with downstream task performance.Overall, our findings help explain why modern text encoders remain effective despite relying on seemingly coarse mean pooling. Tomomasa Hara, Hiroto Kurita, Masaaki Imaizumi, Kentaro Inui, Sho Yokoi |
ACL (1) | 2 |
| 2024 | Zipfian WhiteningabstractThe word embedding space in neural models is skewed, and correcting this can improve task performance.
We point out that most approaches for modeling, correcting, and measuring the symmetry of an embedding space implicitly assume that the word frequencies are *uniform*; in reality, word frequencies follow a highly non-uniform distribution, known as *Zipf's law*.
Surprisingly, simply performing PCA whitening weighted by the empirical word frequency that follows Zipf's law significantly improves task performance, surpassing established baselines.
From a theoretical perspective, both our approach and existing methods can be clearly categorized: word representations are distributed according to an exponential family with either uniform or Zipfian base measures.
By adopting the latter approach, we can naturally emphasize informative low-frequency words in terms of their vector norm, which becomes evident from the information-geometric perspective (Oyama et al., EMNLP 2023), and in terms of the loss functions for imbalanced classification (Menon et al. ICLR 2021).
Additionally, our theory corroborates that popular natural language processing methods, such as skip-gram negative sampling (Mikolov et al., NIPS 2013), WhiteningBERT (Huang et al., Findings of EMNLP 2021), and headless language models (Godey et al., ICLR 2024), work well just because their word embeddings encode the empirical word frequency into the underlying probabilistic model. Sho Yokoi, Han Bao 0002, Hiroto Kurita, Hidetoshi Shimodaira |
NeurIPS | 3 |
| 2007 | Efficient Query Processing for Large XML Data in Distributed EnvironmentsabstractWe propose an efficient distributed query processing method for large XML data by partitioning and distributing XML data to multiple computation nodes. There are several steps involved in this method; however, we focused particularly on XML data partitioning and dynamic relocation of partitioned XML data in our research. Since the efficiency of query processing depends on both XML data size and its structure, these factors should be considered when XML data is partitioned. Each partitioned XML data is distributed to computation nodes so that the CPU load can be balanced. In addition, it is important to take account of the query workload among each of the computation nodes because it is closely related to the query processing cost in distributed environments. In case of load skew among computation nodes, partitioned XML data should be relocated to balance the CPU load. Thus, we implemented an algorithm for relocating partitioned XML data based on the CPU load of query processing. From our experiments, we found that there is a performance advantage in our approach for executing distributed query processing of large XML data. Hiroto Kurita, Kenji Hatano, Jun Miyazaki, Shunsuke Uemura |
AINA | 1 |