VLDB 2026 Research / reviewers in the wild / expert
Kumiko Tanaka-Ishii
dblp:42/2790 · also Kumiko Tanaka
· DBLP profile ↗
45ranked-venue papers
16as first author
7since 2021 · last 2026
0000-0003-1752-3951ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 12 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Language models and text generation · 40% Information extraction and text analysis · 28% Representation and self-supervised learning · 17% | |
| Databases, data mining, and information retrieval
4 papers |
Information retrieval · 60% Data mining · 40% | |
| Theoretical computer science
3 papers |
Information theory · 74% Coding theory · 26% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational finance and economics · 100% |
Topics — the 30 heaviest of 37, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
document retrieval |
1.6 | 2 | 2025 | Information-Theoretic Generative Clustering of Documents · AAAI 2025 Bottleneck-Minimal Indexing for Generative Document Retrieval · ICML 2024 |
Information retrieval › retrieval models
generative retrieval |
1.6 | 2 | 2025 | Information-Theoretic Generative Clustering of Documents · AAAI 2025 Bottleneck-Minimal Indexing for Generative Document Retrieval · ICML 2024 |
Natural language and speech › Language models and text generation
large language model evaluation |
1.3 | 2 | 2026 | Repeated Sequences Reveal Gaps between Large Language Models and Natural Language · ACL (1) 2026 Taylor's law for Human Linguistic Sequences · ACL (1) 2018 |
Data mining
clustering |
1.0 | 2 | 2025 | Information-Theoretic Generative Clustering of Documents · AAAI 2025 Multilingual Spectral Clustering Using Document Similarity Propagation · EMNLP 2009 |
Natural language and speech › Language models and text generation › neural language model
autoregressive language model |
0.9 | 1 | 2025 | Correlation Dimension of Autoregressive Large Language Models · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › word representation
contextualized word representation |
0.9 | 1 | 2025 | A New Formulation of Zipf's Meaning-Frequency Law through Contextual Diversity · ACL (1) 2025 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | Correlation Dimension of Autoregressive Large Language Models · NeurIPS 2025 |
Natural language and speech › Information extraction and text analysis
lexical semantics |
0.9 | 1 | 2025 | A New Formulation of Zipf's Meaning-Frequency Law through Contextual Diversity · ACL (1) 2025 |
Data mining › clustering
document clustering |
0.9 | 1 | 2025 | Information-Theoretic Generative Clustering of Documents · AAAI 2025 |
Data mining › clustering › model-based clustering
generative clustering |
0.9 | 1 | 2025 | Information-Theoretic Generative Clustering of Documents · AAAI 2025 |
Information retrieval
indexing |
0.8 | 1 | 2024 | Bottleneck-Minimal Indexing for Generative Document Retrieval · ICML 2024 |
Natural language and speech › Information extraction and text analysis › word sense disambiguation
polysemy |
0.6 | 1 | 2022 | FIRE: Semantic Field of Words Represented as Non-Linear Functions · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.6 | 1 | 2022 | FIRE: Semantic Field of Words Represented as Non-Linear Functions · NeurIPS 2022 |
Natural language and speech › Information extraction and text analysis
word sense disambiguation |
0.6 | 1 | 2022 | FIRE: Semantic Field of Words Represented as Non-Linear Functions · NeurIPS 2022 |
Computational finance and economics › portfolio management
portfolio optimization |
0.4 | 1 | 2020 | Stock Embeddings Acquired from News Articles and Price History, and an Application to Portfolio Optimization · ACL 2020 |
Information theory › estimation theory
entropy estimation |
0.3 | 1 | 2026 | Repeated Sequences Reveal Gaps between Large Language Models and Natural Language · ACL (1) 2026 |
Information theory › information measures › entropy › generalized entropy
rényi entropy |
0.3 | 1 | 2026 | Repeated Sequences Reveal Gaps between Large Language Models and Natural Language · ACL (1) 2026 |
Natural language and speech › Language models and text generation
hallucination detection |
0.3 | 1 | 2025 | Correlation Dimension of Autoregressive Large Language Models · NeurIPS 2025 |
Coding theory › source coding
rate-distortion theory |
0.2 | 1 | 2024 | Bottleneck-Minimal Indexing for Generative Document Retrieval · ICML 2024 |
Natural language and speech › Information extraction and text analysis
text segmentation |
0.2 | 2 | 2012 | Text Segmentation by Language Using Minimum Description Length · ACL (1) 2012 Unsupervised Segmentation of Chinese Text by Use of Branching Entropy · ACL 2006 |
Natural language and speech › Speech recognition and synthesis › speech analysis
language identification |
0.1 | 1 | 2012 | Text Segmentation by Language Using Minimum Description Length · ACL (1) 2012 |
Interaction techniques and input
text entry |
0.1 | 2 | 2009 | Kansuke: A logograph look-up interface based on a few modified stroke prototypes · ACM Trans. Comput. Hum. Interact. 2009 Acquiring Vocabulary for Predictive Text Entry through Dynamic Reuse of a Small User Corpus · ACL 2003 |
Natural language and speech › Machine translation
multimodal machine translation |
0.1 | 1 | 2011 | picoTrans: Using Pictures as Input for Machine Translation on Mobile Devices · IJCAI 2011 |
Information retrieval › cross-language information retrieval
cross-lingual document similarity |
0.1 | 1 | 2009 | Multilingual Spectral Clustering Using Document Similarity Propagation · EMNLP 2009 |
Data mining › clustering
spectral clustering |
0.1 | 1 | 2009 | Multilingual Spectral Clustering Using Document Similarity Propagation · EMNLP 2009 |
Interaction techniques and input › pen input
stroke-based input |
0.1 | 1 | 2009 | Kansuke: A logograph look-up interface based on a few modified stroke prototypes · ACM Trans. Comput. Hum. Interact. 2009 |
Natural language and speech › Information extraction and text analysis › word segmentation
unsupervised word segmentation |
0.1 | 1 | 2006 | Unsupervised Segmentation of Chinese Text by Use of Branching Entropy · ACL 2006 |
Information theory › information measures
entropy |
0.1 | 1 | 2006 | Unsupervised Segmentation of Chinese Text by Use of Branching Entropy · ACL 2006 |
Information retrieval › search engines
search result analysis |
0.1 | 1 | 2005 | A multilingual usage consultation tool based on internet searching: more than a search engine, less than QA · WWW 2005 |
Interaction techniques and input › text entry
predictive text entry |
0.0 | 1 | 2003 | Acquiring Vocabulary for Predictive Text Entry through Dynamic Reuse of a Small User Corpus · ACL 2003 |
Methods — techniques the papers use, named apart from their topics
repeated subsequence analysis · 2.0power-law modeling · 2.0neural autoregressive model · 1.5mutual information · 1.5information-theoretic analysis · 1.5masked language model · 0.9large language model · 0.9importance sampling · 0.9fractal geometry · 0.9correlation dimension · 0.9contextualized word vectors · 0.9autoregressive language model · 0.9KL divergence · 0.9nonlinear function representation · 0.6neural network · 0.4deep learning · 0.4taylor's law · 0.3power-law analysis · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Repeated Sequences Reveal Gaps between Large Language Models and Natural LanguageabstractEvaluating whether large language models (LLMs) capture the structure of natural language beyond local fluency remains an open challenge.Existing evaluation methods, largely based on task performance or short-context behavior, provide limited insight into the longrange statistical organization of generated text.We propose a complementary evaluation framework based on repeated subsequences.By analyzing their distribution across scales and relating it to higher-order Rényi entropies, we probe how texts reuse previously established structure under finite-length conditions.Experiments on human-written texts and length-matched GPTgenerated texts show that, while power-law models can describe restricted ranges of block length, the observed entropy growth is often equally or better characterized by logarithmicpower forms.Across datasets, natural language exhibits stable entropy-growth patterns over accessible ranges, with consistent average behavior despite variability across individual texts.In contrast, GPT-generated texts show systematic and statistically significant shifts in estimated exponents with model size.These results demonstrate that repeated-subsequence entropy provides a quantitative structural diagnostic that reveals systematic differences in long-range organization, distinguishing natural language from state-of-the-art LLM outputs beyond surfacelevel fluency. Kumiko Tanaka-Ishii |
ACL (1) | 1 |
| 2025 | Information-Theoretic Generative Clustering of DocumentsabstractWe present *generative clustering* (GC) for clustering a set of documents, X, by using texts Y generated by large language models (LLMs) instead of by clustering the original documents X. Because LLMs provide probability distributions, the similarity between two documents can be rigorously defined in an information-theoretic manner by the KL divergence. We also propose a natural, novel clustering algorithm by using importance sampling. We show that GC outperforms any previous clustering method, often by a large margin. Furthermore, we show an application to generative document retrieval in which documents are indexed via hierarchical clustering and our method improves the retrieval accuracy. Kumiko Tanaka-Ishii |
AAAI | 2 |
| 2025 | A New Formulation of Zipf's Meaning-Frequency Law through Contextual DiversityabstractThis paper proposes formulating Zipf’s meaning-frequency law, the power law between word frequency and the number of meanings, as a relationship between word frequency and contextual diversity. The proposed formulation quantifies meaning counts as contextual diversity, which is based on the directions of contextualized word vectors obtained from a Language Model (LM). This formulation gives a new interpretation to the law and also enables us to examine it for a wider variety of words and corpora than previous studies have explored. In addition, this paper shows that the law becomes unobservable when the size of the LM used is small and that autoregressive LMs require much more parameters than masked LMs to be able to observe the law. Ryo Nagata, Kumiko Tanaka-Ishii |
ACL (1) | 2 |
| 2025 | Correlation Dimension of Autoregressive Large Language ModelsabstractLarge language models (LLMs) have achieved remarkable progress in natural
language generation, yet they continue to display puzzling behaviors—such as
repetition and incoherence—even when exhibiting low perplexity. This
highlights a key limitation of conventional evaluation metrics, which
emphasize local prediction accuracy while overlooking long-range structural
complexity. We introduce correlation dimension, a fractal-geometric measure
of self-similarity, to quantify the epistemological complexity of text as
perceived by a language model. This measure captures the hierarchical
recurrence structure of language, bridging local and global properties in a
unified framework. Through extensive experiments, we show that correlation
dimension (1) reveals three distinct phases during pretraining, (2) reflects
context-dependent complexity, (3) indicates a model's tendency toward
hallucination, and (4) reliably detects multiple forms of degeneration in
generated text. The method is computationally efficient, robust to model
quantization (down to 4-bit precision), broadly applicable across
autoregressive architectures (e.g., Transformer and Mamba), and provides
fresh insight into the generative dynamics of LLMs. Kumiko Tanaka-Ishii |
NeurIPS | 2 |
| 2024 | Bottleneck-Minimal Indexing for Generative Document RetrievalabstractWe apply an information-theoretic perspective to reconsider generative document retrieval (GDR), in which a document $x \in \mathcal{X}$ is indexed by $t \in \mathcal{T}$, and a neural autoregressive model is trained to map queries $\mathcal{Q}$ to $\mathcal{T}$. GDR can be considered to involve information transmission from documents $\mathcal{X}$ to queries $\mathcal{Q}$, with the requirement to transmit more bits via the indexes $\mathcal{T}$. By applying Shannon’s rate-distortion theory, the optimality of indexing can be analyzed in terms of the mutual information, and the design of the indexes $\mathcal{T}$ can then be regarded as a bottleneck in GDR. After reformulating GDR from this perspective, we empirically quantify the bottleneck underlying GDR. Finally, using the NQ320K and MARCO datasets, we evaluate our proposed bottleneck-minimal indexing method in comparison with various previous indexing methods, and we show that it outperforms those methods. Lixin Xiu, Kumiko Tanaka-Ishii |
ICML | 3 |
| 2022 | FIRE: Semantic Field of Words Represented as Non-Linear FunctionsabstractState-of-the-art word embeddings presume a linear vector space, but this approach does not easily incorporate the nonlinearity that is necessary to represent polysemy. We thus propose a novel semantic FIeld REepresentation, called FIRE, which is a $D$-dimensional field in which every word is represented as a set of its locations and a nonlinear function covering the field. The strength of a word's relation to another word at a certain location is measured as the function value at that location. With FIRE, compositionality is represented via functional additivity, whereas polysemy is represented via the set of points and the function's multimodality. By implementing FIRE for English and comparing it with previous representation methods via word and sentence similarity tasks, we show that FIRE produces comparable or even better results. In an evaluation of polysemy to predict the number of word senses, FIRE greatly outperformed BERT and Word2vec, providing evidence of how FIRE represents polysemy. The code is available at https://github.com/kduxin/firelang. Kumiko Tanaka-Ishii |
NeurIPS | 2 |
| 2022 | Stock portfolio selection balancing variance and tail risk via stock vector representation acquired from price data and texts
Kumiko Tanaka-Ishii |
Knowl. Based Syst. | 2 |
| 2020 | Stock Embeddings Acquired from News Articles and Price History, and an Application to Portfolio OptimizationabstractPrevious works that integrated news articles to better process stock prices used a variety of neural networks to predict price movements.The textual and price information were both encoded in the neural network, and it is therefore difficult to apply this approach in situations other than the original framework of the notoriously hard problem of price prediction.In contrast, this paper presents a method to encode the influence of news articles through a vector representation of stocks called a stock embedding.The stock embedding is acquired with a deep learning framework using both news articles and price history.Because the embedding takes the operational form of a vector, it is applicable to other financial problems besides price prediction.As one example application, we show the results of portfolio optimization using Reuters & Bloomberg headlines, producing a capital gain 2.8 times larger than that obtained with a baseline method using only stock price data.This suggests that the proposed stock embedding can leverage textual financial semantics to solve financial prediction problems. Kumiko Tanaka-Ishii |
ACL | 2 |
| 2019 | Evaluating Computational Language Models with Scaling Properties of Natural LanguageabstractIn this article, we evaluate computational models of natural language with respect to the universal statistical behaviors of natural language. Statistical mechanical analyses have revealed that natural language text is characterized by scaling properties, which quantify the global structure in the vocabulary population and the long memory of a text. We study whether five scaling properties (given by Zipf’s law, Heaps’ law, Ebeling’s method, Taylor’s law, and long-range correlation analysis) can serve for evaluation of computational models. Specifically, we test n-gram language models, a probabilistic context-free grammar, language models based on Simon/Pitman-Yor processes, neural language models, and generative adversarial networks for text generation. Our analysis reveals that language models based on recurrent neural networks with a gating mechanism (i.e., long short-term memory; a gated recurrent unit; and quasi-recurrent neural networks) are the only computational models that can reproduce the long memory behavior of natural language. Furthermore, through comparison with recently proposed model-based evaluation methods, we find that the exponent of Taylor’s law is a good indicator of model quality. Shuntaro Takahashi, Kumiko Tanaka-Ishii |
Comput. Linguistics | 2 |
| 2018 | Taylor's law for Human Linguistic SequencesabstractTaylor's law describes the fluctuation characteristics underlying a system in which the variance of an event within a time span grows by a power law with respect to the mean.Although Taylor's law has been applied in many natural and social systems, its application for language has been scarce.This article describes a new quantification of Taylor's law in natural language and reports an analysis of over 1100 texts across 14 languages.The Taylor exponents of written natural language texts were found to exhibit almost the same value.The exponent was also compared for other language-related data, such as the child-directed speech, music, and programming language code.The results show how the Taylor exponent serves to quantify the fundamental structural complexity underlying linguistic time series.The article also shows the applicability of these findings in evaluating language models. Tatsuru Kobayashi, Kumiko Tanaka-Ishii |
ACL (1) | 2 |
| 2018 | Extraction of templates from phrases using Sequence Binary Decision DiagramsabstractAbstract The extraction of templates such as ‘regard X as Y’ from a set of related phrases requires the identification of their internal structures. This paper presents an unsupervised approach for extracting templates on-the-fly from only tagged text by using a novel relaxed variant of the Sequence Binary Decision Diagram (SeqBDD). A SeqBDD can compress a set of sequences into a graphical structure equivalent to a minimal deterministic finite state automata, but more compact and better suited to the task of template extraction. The main contribution of this paper is a relaxed form of the SeqBDD construction algorithm that enables it to form general representations from a small amount of data. The process of compression of shared structures in the text during Relaxed SeqBDD construction, naturally induces the templates we wish to extract. Experiments show that the method is capable of high-quality extraction on tasks based on verb+preposition templates from corpora and phrasal templates from short messages from social media. Daiki Hirano, Kumiko Tanaka-Ishii, Andrew M. Finch |
Nat. Lang. Eng. | 2 |
| 2017 | Component Awareness in Convolutional Neural NetworksabstractIn this work, we investigate the ability of Convolutional Neural Networks (CNN) to infer the presence of components that comprise an image. In recent years, CNNs have achieved powerful results in classification, detection, and segmentation. However, these models learn from instance-level supervision of the detected object. In this paper, we determine if CNNs can detect objects using image-level weakly supervised labels without localization. To demonstrate that a CNN can infer awareness of objects, we evaluate a CNN's classification ability with a database constructed of Chinese characters with only character-level labeled components. We show that the CNN is able to achieve a high accuracy in identifying the presence of these components without specific knowledge of the component. Furthermore, we verify that the CNN is deducing the knowledge of the target component by comparing the results to an experiment with the component removed. This research is important for applications with large amounts of data without robust annotation such as Chinese character recognition. Brian Kenji Iwana, Letao Zhou, Kumiko Tanaka-Ishii, Seiichi Uchida |
ICDAR | 3 |
| 2017 | Inducing a Bilingual Lexicon from Short Parallel Multiword SequencesabstractThis article proposes a technique for mining bilingual lexicons from pairs of parallel short word sequences. The technique builds a generative model from a corpus of training data consisting of such pairs. The model is a hierarchical nonparametric Bayesian model that directly induces a bilingual lexicon while training. The model learns in an unsupervised manner and is designed to exploit characteristics of the language pairs being mined. The proposed model is capable of utilizing commonly used word-pair frequency information and additionally can employ the internal character alignments within the words themselves. It is thereby capable of mining transliterations and can use reliably aligned transliteration pairs to support the mining of other words in their context. The model is also capable of performing word reordering and word deletion during the alignment process, and it is furthermore capable of operating in the absence of full segmentation information. In this work, we study two mining tasks based on English-Japanese and English-Chinese language pairs, and compare the proposed approach to baselines based on a simpler models that use only word-pair frequency information. Our results show that the proposed method is able to mine bilingual word pairs at higher levels of precision and recall than the baselines. Andrew M. Finch, Taisuke Harada, Kumiko Tanaka-Ishii, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2015 | Computational Constancy Measures of Texts - Yule's K and Rényi's EntropyabstractThis article presents a mathematical and empirical verification of computational constancy measures for natural language text. A constancy measure characterizes a given text by having an invariant value for any size larger than a certain amount. The study of such measures has a 70-year history dating back to Yule's K, with the original intended application of author identification. We examine various measures proposed since Yule and reconsider reports made so far, thus overviewing the study of constancy measures. We then explain how K is essentially equivalent to an approximation of the second-order Rényi entropy, thus indicating its signification within language science. We then empirically examine constancy measure candidates within this new, broader context. The approximated higher-order entropy exhibits stable convergence across different languages and kinds of text. We also show, however, that it cannot identify authors, contrary to Yule's intention. Lastly, we apply K to two unknown scripts, the Voynich manuscript and Rongorongo, and show how the results support previous hypotheses about these scripts. Kumiko Tanaka-Ishii, Shunsuke Aihara |
Comput. Linguistics | 1 |
| 2013 | picoTrans: An intelligent icon-driven interface for cross-lingual communicationabstractpicoTrans is a prototype system that introduces a novel icon-based paradigm for cross-lingual communication on mobile devices. Our approach marries a machine translation system with the popular picture book. Users interact with picoTrans by pointing at pictures as if it were a picture book; the system generates natural language from these icons and the user is able to interact with the icon sequence to refine the meaning of the words that are generated. When users are satisfied that the sentence generated represents what they wish to express, they tap a translate button and picoTrans displays the translation. Structuring the process of communication in this way has many advantages. First, tapping icons is a very natural method of user input on mobile devices; typing is cumbersome and speech input errorful. Second, the sequence of icons which is annotated both with pictures and bilingually with words is meaningful to both users, and it opens up a second channel of communication between them that conveys the gist of what is being expressed. We performed a number of evaluations of picoTrans to determine: its coverage of a corpus of in-domain sentences; the input efficiency in terms of the number of key presses required relative to text entry; and users' overall impressions of using the system compared to using a picture book. Our results show that we are able to cover 74% of the expressions in our test corpus using a 2000-icon set; we believe that this icon set size is realistic for a mobile device. We also found that picoTrans requires fewer key presses than typing the input and that the system is able to predict the correct, intended natural language sentence from the icon sequence most of the time, making user interaction with the icon sequence often unnecessary. In the user evaluation, we found that in general users prefer using picoTrans and are able to communicate more rapidly and expressively. Furthermore, users had more confidence that they were able to communicate effectively using picoTrans. Andrew M. Finch, Kumiko Tanaka-Ishii, Keiji Yasuda, Eiichiro Sumita |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2012 | Text Segmentation by Language Using Minimum Description Length
Hiroshi Yamaguchi, Kumiko Tanaka-Ishii |
ACL (1) | 2 |
| 2011 | picoTrans: Using Pictures as Input for Machine Translation on Mobile DevicesabstractIn this paper we present a novel user interface that integrates two popular approaches to language translation for travelers allowing multimodal communication between the parties involved: the picture-book, in which the user simply points to multiple picture icons representing what they want to say, and the statistical machine translation (SMT) system that can translate arbitrary word sequences. Our prototype system tightly couples both processes within a translation framework that inherits many of the the positive features of both approaches, while at the same time mitigating their main weaknesses. Our system differs from traditional approaches in that its mode of input is a sequence of pictures, rather than text or speech. Text in the source language is generated automatically, and is used as a detailed representation of the intended meaning. The picture sequence which not only provides a rapid method to communicate basic concepts but also gives a 'second opinion' on the machine transition output that catches machine translation errors and allows the users to retry the translation, avoiding misunderstandings. Andrew M. Finch, Kumiko Tanaka-Ishii, Eiichiro Sumita |
IJCAI | 3 |
| 2011 | Relational Lasso - An Improved Method Using the Relations Among Features -
Kotaro Kitagawa, Kumiko Tanaka-Ishii |
IJCNLP | 2 |
| 2011 | picoTrans: an icon-driven user interface for machine translation on mobile devicesabstractIn this paper we present a novel user interface that integrates two popular approaches to language translation for travelers allowing multimodal communication between the parties involved. In our approach we integrate the popular picture-book, in which the user simply points to multiple picture icons representing what they want to say, with a statistical machine translation system that can translate arbitrary word sequences. The simple pointing at pictures paradigm is used as the primary method of user input and the users can use the device as if it were a picture book. The application is then able to generate a complete sentence in the user's native language for what they wish to say from the sequence of picture icons chosen by the user. Once the user is satisfied that the sentence provided by the system adequately represents what they wish to convey, the application can automatically translate the sentence into the language of the other party, who can interpret the intended meaning of the first party by combining evidence from both modes of communication: the picture sequence, and the machine translation. The prototype system we have developed inherits many of the positive features of both approaches, while at the same time mitigating their main weaknesses. The user may combine the pictures in considerably more combinations than is possible with a picture book designed with combinations from only within the same page spread of the book in mind, making the application more expressive than a book. The machine translation system can contribute a detailed and precise translation which is supported by the picture-based mode which not only provides a rapid method to communicate basic concepts but also gives a 'second opinion' on the machine transition output that catches machine translation errors and allows the users to retry the sentence, avoiding misunderstandings. Andrew M. Finch, Kumiko Tanaka-Ishii, Eiichiro Sumita |
IUI | 3 |
| 2010 | YouBot: A Simple Framework for Building Virtual Networking Agents
Seiji Takegata, Kumiko Tanaka-Ishii |
SIGDIAL Conference | 2 |
| 2010 | Sorting Texts by ReadabilityabstractThis article presents a novel approach for readability assessment through sorting. A comparator that judges the relative readability between two texts is generated through machine learning, and a given set of texts is sorted by this comparator. Our proposal is advantageous because it solves the problem of a lack of training data, because the construction of the comparator only requires training data annotated with two reading levels. The proposed method is compared with regression methods and a state-of-the art classification method. Moreover, we present our application, called Terrace, which retrieves texts with readability similar to that of a given input text. Kumiko Tanaka-Ishii, Satoshi Tezuka, Hiroshi Terada |
Comput. Linguistics | 1 |
| 2009 | Multilingual Spectral Clustering Using Document Similarity Propagation
Dani Yogatama, Kumiko Tanaka-Ishii |
EMNLP | 2 |
| 2009 | Kansuke: A logograph look-up interface based on a few modified stroke prototypesabstractWe have developed a method that makes it easier for language novices to look up Japanese and Chinese logographs. Instead of using the arbitrary conventions of logographs, this method is based on three simple prototypes: horizontal, vertical, and other strokes. For example, the code for the logograph ⊞ ( ta , meaning rice field) is 3-3-0, indicating the logograph consists of three horizontal strokes and three vertical strokes. Such codes allow a novice to look up logographs even with no knowledge of the logographic conventions used by native speakers. To make the search easier, a complex logograph can be looked up via the components making up the logograph. We conducted a user evaluation of this system and found that novices could look up logographs with fewer failures with our system than with conventional methods. Kumiko Tanaka-Ishii, Julian Godon |
ACM Trans. Comput. Hum. Interact. | 1 |
| 2008 | Multilingual Text Entry using Automatic Language Detection
Yo Ehara, Kumiko Tanaka-Ishii |
IJCNLP | 2 |
| 2007 | Introduction to the Special Issue on Recent PACLING Meetings
Kiyoshi Kogure, Shun Ishizaki, Tsutomu Endo, Yoshihiko Nitta, Kumiko Tanaka-Ishii |
Comput. Intell. | 5 |
| 2007 | Multilingual phrase-based concordance generation in real-time
Kumiko Tanaka-Ishii, Yuichiro Ishii |
Inf. Retr. | 1 |
| 2007 | Word-based predictive text entry using adaptive language modelsabstractThe recent scaling down of mobile device form factors has increased the importance of predictive text entry. It is now also becoming an important communication tool for the disabled. Techniques related to predictive text entry software are discussed in a generalized, language-independent manner. The essence of predictive text entry is twofold, consisting of (1) the design of codes for text entry, and (2) the use of adaptive language models for decoding. Code design is examined in terms of the information-theoretical efficiency. Four adaptive language models are introduced and compared, and experimental results on text entry with these models are shown for English, Thai and Japanese. Kumiko Tanaka-Ishii |
Nat. Lang. Eng. | 1 |
| 2006 | Unsupervised Segmentation of Chinese Text by Use of Branching Entropy
Zhihui Jin, Kumiko Tanaka-Ishii |
ACL | 2 |
| 2005 | Entropy as an Indicator of Context Boundaries: An Experiment Using a Web Search Engine
Kumiko Tanaka-Ishii |
IJCNLP | 1 |
| 2005 | A multilingual usage consultation tool based on internet searching: more than a search engine, less than QAabstractWe present a usage consultation tool, based on Internet searching, for language learners. When a user enters a string of words for which he wants to find usages, the system sends this string as a query to a search engine and obtains search results about the string. The usages are extracted by performing statistical analysis on snippets and then fed back to the user.Unlike existing tools, this usage consultation tool is multi-lingual, so that usages can be obtained even in a language for which there are no well-established analytical methods. Our evaluation has revealed that usages can be obtained more effectively than by only using a search engine directly. Also, we have found that the resulting usage does not depend on the search engine for a prominent usage when the amount of data downloaded from the search engine is increased. Kumiko Tanaka-Ishii, Hiroshi Nakagawa |
WWW | 1 |
| 2004 | An Interactive Proofreading System for Inappropriately Selected Words on Using Predictive Text Entry
Hideya Iwasaki, Kumiko Tanaka-Ishii |
IJCNLP | 2 |
| 2004 | Introduction (Thematic Session: Natural Language Technology in Mobile Information Retrieval and Text Processing User Interfaces)
Michael Kuehn, Mun-Kew Leong, Kumiko Tanaka-Ishii |
IJCNLP | 3 |
| 2004 | Dit4dah: Predictive Pruning for Morse Code Text Entry
Kumiko Tanaka-Ishii, Ian Frank |
IJCNLP | 1 |
| 2004 | EMMA: a web-based report system for programming course--automated verification and enhanced feedbackabstractNo abstract available. Kumiko Tanaka-Ishii, Kazuhiko Kakehi 0001, Masato Takeichi |
ITiCSE | 1 |
| 2003 | Acquiring Vocabulary for Predictive Text Entry through Dynamic Reuse of a Small User CorpusabstractAs mobile computing and communications have become popular, predictive text entry systems have become an increasingly important technology. Existing methods still need refinement, though, with respect to personalization, especially how to acquire vocabulary not pre-registered in the system dictionary. In this paper, we report on an automatic method that dynamically obtains a user specific vocabulary from the user's unanalyzed documents. When a user makes an entry, the system dynamically extracts the corresponding chunks from the user text and suggests them along with words suggested by the dictionary. With our method, texts in a particular style or concerning a specific domain can be entered using a predictive text entry system. We verified that a large amount of words not registered in the dictionary can be entered using our method. Kumiko Tanaka-Ishii, Daichi Hayakawa, Masato Takeichi |
ACL | 1 |
| 2003 | Performance Competitions as Research Infrastructure: Large Scale Comparative Studies of Multi-Agent Teams
Gal A. Kaminka, Ian Frank, Katsuto Arai, Kumiko Tanaka-Ishii |
Auton. Agents Multi Agent Syst. | 4 |
| 2002 | Entering Text with a Four-Button Device
Kumiko Tanaka-Ishii, Yusuke Inutsuka, Masato Takeichi |
COLING | 1 |
| 2001 | Walkie-Talkie MIKE
Ian Frank, Kumiko Tanaka-Ishii, Hitoshi Matsubara, Eiichi Osawa |
RoboCup | 2 |
| 2000 | Multi-Agent Explanation Strategies in Real-Time DomainsabstractWe examine the benefits of using multiple agents to produce explanations. In particular, we identify the ability to construct prior plans as a key issue constraining the effectiveness of a single-agent approach. We describe an implemented system that uses multiple agents to tackle a problem for which prior planning is particularly impractical: real-time soccer commentary. Our commentary system demonstrates a number of the advantages of decomposing an explanation task among several agents. Most notably, it shows how individual agents can benefit from following different discourse strategies. Further, it illustrates that discourse issues such as controlling interruption, abbreviation, and maintaining consistency can also be decomposed: rather than considering them at the single level of one linear explanation they can also be tackled separately within each individual agent. We evaluate our system's output, and show that it closely compares to the speaking patterns of a human commentary team. Kumiko Tanaka-Ishii, Ian Frank |
ACL | 1 |
| 2000 | The Statistics Proxy Server
Ian Frank, Kumiko Tanaka-Ishii, Katsuto Arai, Hitoshi Matsubara |
RoboCup | 2 |
| 2000 | And the Fans Are Going Wild! SIG plus MIKE
Ian Frank, Kumiko Tanaka-Ishii, Hiroshi G. Okuno, Junichi Akita, Yukiko Nakagawa, Kazuaki Maeda, Kazuhiro Nakadai, Hiroaki Kitano |
RoboCup | 2 |
| 1999 | A Statistical Perspective on the RoboCup Simulator League: Progress and Prospects
Kumiko Tanaka-Ishii, Ian Frank, Itsuki Noda, Hitoshi Matsubara |
RoboCup | 1 |
| 1998 | Automatic Soccer Commentary and RoboCup
Hitoshi Matsubara, Itsuki Noda, Ian Frank, Hideyuki Nakashima, Kumiko Tanaka-Ishii, Kôiti Hasida |
RoboCup | 5 |
| 1996 | Extraction of Lexical Translations from Non-Aligned Corpora
Kumiko Tanaka-Ishii, Hideya Iwasaki |
COLING | 1 |
| 1994 | Construction of a Bilingual Dictionary Intermediated by a Third Language
Kumiko Tanaka-Ishii, Kyoji Umemura |
COLING | 1 |