Kumiko Tanaka-Ishii

dblp:42/2790 · also Kumiko Tanaka · DBLP profile ↗
← Back
45ranked-venue papers
16as first author
7since 2021 · last 2026
0000-0003-1752-3951ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 12 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Language models and text generation · 40% Information extraction and text analysis · 28% Representation and self-supervised learning · 17%
Databases, data mining, and information retrieval
4 papers
Information retrieval · 60% Data mining · 40%
Theoretical computer science
3 papers
Information theory · 74% Coding theory · 26%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational finance and economics · 100%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
document retrieval
1.622025
Information-Theoretic Generative Clustering of Documents · AAAI 2025
Bottleneck-Minimal Indexing for Generative Document Retrieval · ICML 2024
Information retrieval › retrieval models
generative retrieval
1.622025
Information-Theoretic Generative Clustering of Documents · AAAI 2025
Bottleneck-Minimal Indexing for Generative Document Retrieval · ICML 2024
Natural language and speech › Language models and text generation
large language model evaluation
1.322026
Repeated Sequences Reveal Gaps between Large Language Models and Natural Language · ACL (1) 2026
Taylor's law for Human Linguistic Sequences · ACL (1) 2018
Data mining
clustering
1.022025
Information-Theoretic Generative Clustering of Documents · AAAI 2025
Multilingual Spectral Clustering Using Document Similarity Propagation · EMNLP 2009
Natural language and speech › Language models and text generation › neural language model
autoregressive language model
0.912025
Correlation Dimension of Autoregressive Large Language Models · NeurIPS 2025
Machine learning › Representation and self-supervised learning › word representation
contextualized word representation
0.912025
A New Formulation of Zipf's Meaning-Frequency Law through Contextual Diversity · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model
0.912025
Correlation Dimension of Autoregressive Large Language Models · NeurIPS 2025
Natural language and speech › Information extraction and text analysis
lexical semantics
0.912025
A New Formulation of Zipf's Meaning-Frequency Law through Contextual Diversity · ACL (1) 2025
Data mining › clustering
document clustering
0.912025
Information-Theoretic Generative Clustering of Documents · AAAI 2025
Data mining › clustering › model-based clustering
generative clustering
0.912025
Information-Theoretic Generative Clustering of Documents · AAAI 2025
Information retrieval
indexing
0.812024
Bottleneck-Minimal Indexing for Generative Document Retrieval · ICML 2024
Natural language and speech › Information extraction and text analysis › word sense disambiguation
polysemy
0.612022
FIRE: Semantic Field of Words Represented as Non-Linear Functions · NeurIPS 2022
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.612022
FIRE: Semantic Field of Words Represented as Non-Linear Functions · NeurIPS 2022
Natural language and speech › Information extraction and text analysis
word sense disambiguation
0.612022
FIRE: Semantic Field of Words Represented as Non-Linear Functions · NeurIPS 2022
Computational finance and economics › portfolio management
portfolio optimization
0.412020
Stock Embeddings Acquired from News Articles and Price History, and an Application to Portfolio Optimization · ACL 2020
Information theory › estimation theory
entropy estimation
0.312026
Repeated Sequences Reveal Gaps between Large Language Models and Natural Language · ACL (1) 2026
Information theory › information measures › entropy › generalized entropy
rényi entropy
0.312026
Repeated Sequences Reveal Gaps between Large Language Models and Natural Language · ACL (1) 2026
Natural language and speech › Language models and text generation
hallucination detection
0.312025
Correlation Dimension of Autoregressive Large Language Models · NeurIPS 2025
Coding theory › source coding
rate-distortion theory
0.212024
Bottleneck-Minimal Indexing for Generative Document Retrieval · ICML 2024
Natural language and speech › Information extraction and text analysis
text segmentation
0.222012
Text Segmentation by Language Using Minimum Description Length · ACL (1) 2012
Unsupervised Segmentation of Chinese Text by Use of Branching Entropy · ACL 2006
Natural language and speech › Speech recognition and synthesis › speech analysis
language identification
0.112012
Text Segmentation by Language Using Minimum Description Length · ACL (1) 2012
Interaction techniques and input
text entry
0.122009
Kansuke: A logograph look-up interface based on a few modified stroke prototypes · ACM Trans. Comput. Hum. Interact. 2009
Acquiring Vocabulary for Predictive Text Entry through Dynamic Reuse of a Small User Corpus · ACL 2003
Natural language and speech › Machine translation
multimodal machine translation
0.112011
picoTrans: Using Pictures as Input for Machine Translation on Mobile Devices · IJCAI 2011
Information retrieval › cross-language information retrieval
cross-lingual document similarity
0.112009
Multilingual Spectral Clustering Using Document Similarity Propagation · EMNLP 2009
Data mining › clustering
spectral clustering
0.112009
Multilingual Spectral Clustering Using Document Similarity Propagation · EMNLP 2009
Interaction techniques and input › pen input
stroke-based input
0.112009
Kansuke: A logograph look-up interface based on a few modified stroke prototypes · ACM Trans. Comput. Hum. Interact. 2009
Natural language and speech › Information extraction and text analysis › word segmentation
unsupervised word segmentation
0.112006
Unsupervised Segmentation of Chinese Text by Use of Branching Entropy · ACL 2006
Information theory › information measures
entropy
0.112006
Unsupervised Segmentation of Chinese Text by Use of Branching Entropy · ACL 2006
Information retrieval › search engines
search result analysis
0.112005
A multilingual usage consultation tool based on internet searching: more than a search engine, less than QA · WWW 2005
Interaction techniques and input › text entry
predictive text entry
0.012003
Acquiring Vocabulary for Predictive Text Entry through Dynamic Reuse of a Small User Corpus · ACL 2003

Methods — techniques the papers use, named apart from their topics

repeated subsequence analysis · 2.0power-law modeling · 2.0neural autoregressive model · 1.5mutual information · 1.5information-theoretic analysis · 1.5masked language model · 0.9large language model · 0.9importance sampling · 0.9fractal geometry · 0.9correlation dimension · 0.9contextualized word vectors · 0.9autoregressive language model · 0.9KL divergence · 0.9nonlinear function representation · 0.6neural network · 0.4deep learning · 0.4taylor's law · 0.3power-law analysis · 0.3
YearPublicationVenuePosition
2026 Repeated Sequences Reveal Gaps between Large Language Models and Natural Language
abstract
Evaluating whether large language models (LLMs) capture the structure of natural language beyond local fluency remains an open challenge.Existing evaluation methods, largely based on task performance or short-context behavior, provide limited insight into the longrange statistical organization of generated text.We propose a complementary evaluation framework based on repeated subsequences.By analyzing their distribution across scales and relating it to higher-order Rényi entropies, we probe how texts reuse previously established structure under finite-length conditions.Experiments on human-written texts and length-matched GPTgenerated texts show that, while power-law models can describe restricted ranges of block length, the observed entropy growth is often equally or better characterized by logarithmicpower forms.Across datasets, natural language exhibits stable entropy-growth patterns over accessible ranges, with consistent average behavior despite variability across individual texts.In contrast, GPT-generated texts show systematic and statistically significant shifts in estimated exponents with model size.These results demonstrate that repeated-subsequence entropy provides a quantitative structural diagnostic that reveals systematic differences in long-range organization, distinguishing natural language from state-of-the-art LLM outputs beyond surfacelevel fluency.
Kumiko Tanaka-Ishii
ACL (1)1
2025 Information-Theoretic Generative Clustering of Documents
abstract
We present *generative clustering* (GC) for clustering a set of documents, X, by using texts Y generated by large language models (LLMs) instead of by clustering the original documents X. Because LLMs provide probability distributions, the similarity between two documents can be rigorously defined in an information-theoretic manner by the KL divergence. We also propose a natural, novel clustering algorithm by using importance sampling. We show that GC outperforms any previous clustering method, often by a large margin. Furthermore, we show an application to generative document retrieval in which documents are indexed via hierarchical clustering and our method improves the retrieval accuracy.
Kumiko Tanaka-Ishii
AAAI2
2025 A New Formulation of Zipf's Meaning-Frequency Law through Contextual Diversity
abstract
This paper proposes formulating Zipf’s meaning-frequency law, the power law between word frequency and the number of meanings, as a relationship between word frequency and contextual diversity. The proposed formulation quantifies meaning counts as contextual diversity, which is based on the directions of contextualized word vectors obtained from a Language Model (LM). This formulation gives a new interpretation to the law and also enables us to examine it for a wider variety of words and corpora than previous studies have explored. In addition, this paper shows that the law becomes unobservable when the size of the LM used is small and that autoregressive LMs require much more parameters than masked LMs to be able to observe the law.
Ryo Nagata, Kumiko Tanaka-Ishii
ACL (1)2
2025 Correlation Dimension of Autoregressive Large Language Models
abstract
Large language models (LLMs) have achieved remarkable progress in natural language generation, yet they continue to display puzzling behaviors—such as repetition and incoherence—even when exhibiting low perplexity. This highlights a key limitation of conventional evaluation metrics, which emphasize local prediction accuracy while overlooking long-range structural complexity. We introduce correlation dimension, a fractal-geometric measure of self-similarity, to quantify the epistemological complexity of text as perceived by a language model. This measure captures the hierarchical recurrence structure of language, bridging local and global properties in a unified framework. Through extensive experiments, we show that correlation dimension (1) reveals three distinct phases during pretraining, (2) reflects context-dependent complexity, (3) indicates a model's tendency toward hallucination, and (4) reliably detects multiple forms of degeneration in generated text. The method is computationally efficient, robust to model quantization (down to 4-bit precision), broadly applicable across autoregressive architectures (e.g., Transformer and Mamba), and provides fresh insight into the generative dynamics of LLMs.
Kumiko Tanaka-Ishii
NeurIPS2
2024 Bottleneck-Minimal Indexing for Generative Document Retrieval
abstract
We apply an information-theoretic perspective to reconsider generative document retrieval (GDR), in which a document $x \in \mathcal{X}$ is indexed by $t \in \mathcal{T}$, and a neural autoregressive model is trained to map queries $\mathcal{Q}$ to $\mathcal{T}$. GDR can be considered to involve information transmission from documents $\mathcal{X}$ to queries $\mathcal{Q}$, with the requirement to transmit more bits via the indexes $\mathcal{T}$. By applying Shannon’s rate-distortion theory, the optimality of indexing can be analyzed in terms of the mutual information, and the design of the indexes $\mathcal{T}$ can then be regarded as a bottleneck in GDR. After reformulating GDR from this perspective, we empirically quantify the bottleneck underlying GDR. Finally, using the NQ320K and MARCO datasets, we evaluate our proposed bottleneck-minimal indexing method in comparison with various previous indexing methods, and we show that it outperforms those methods.
Lixin Xiu, Kumiko Tanaka-Ishii
ICML3
2022 FIRE: Semantic Field of Words Represented as Non-Linear Functions
abstract
State-of-the-art word embeddings presume a linear vector space, but this approach does not easily incorporate the nonlinearity that is necessary to represent polysemy. We thus propose a novel semantic FIeld REepresentation, called FIRE, which is a $D$-dimensional field in which every word is represented as a set of its locations and a nonlinear function covering the field. The strength of a word's relation to another word at a certain location is measured as the function value at that location. With FIRE, compositionality is represented via functional additivity, whereas polysemy is represented via the set of points and the function's multimodality. By implementing FIRE for English and comparing it with previous representation methods via word and sentence similarity tasks, we show that FIRE produces comparable or even better results. In an evaluation of polysemy to predict the number of word senses, FIRE greatly outperformed BERT and Word2vec, providing evidence of how FIRE represents polysemy. The code is available at https://github.com/kduxin/firelang.
Kumiko Tanaka-Ishii
NeurIPS2
2022 Stock portfolio selection balancing variance and tail risk via stock vector representation acquired from price data and texts
Kumiko Tanaka-Ishii
Knowl. Based Syst.2
2020 Stock Embeddings Acquired from News Articles and Price History, and an Application to Portfolio Optimization
abstract
Previous works that integrated news articles to better process stock prices used a variety of neural networks to predict price movements.The textual and price information were both encoded in the neural network, and it is therefore difficult to apply this approach in situations other than the original framework of the notoriously hard problem of price prediction.In contrast, this paper presents a method to encode the influence of news articles through a vector representation of stocks called a stock embedding.The stock embedding is acquired with a deep learning framework using both news articles and price history.Because the embedding takes the operational form of a vector, it is applicable to other financial problems besides price prediction.As one example application, we show the results of portfolio optimization using Reuters & Bloomberg headlines, producing a capital gain 2.8 times larger than that obtained with a baseline method using only stock price data.This suggests that the proposed stock embedding can leverage textual financial semantics to solve financial prediction problems.
Kumiko Tanaka-Ishii
ACL2
2019 Evaluating Computational Language Models with Scaling Properties of Natural Language
abstract
In this article, we evaluate computational models of natural language with respect to the universal statistical behaviors of natural language. Statistical mechanical analyses have revealed that natural language text is characterized by scaling properties, which quantify the global structure in the vocabulary population and the long memory of a text. We study whether five scaling properties (given by Zipf’s law, Heaps’ law, Ebeling’s method, Taylor’s law, and long-range correlation analysis) can serve for evaluation of computational models. Specifically, we test n-gram language models, a probabilistic context-free grammar, language models based on Simon/Pitman-Yor processes, neural language models, and generative adversarial networks for text generation. Our analysis reveals that language models based on recurrent neural networks with a gating mechanism (i.e., long short-term memory; a gated recurrent unit; and quasi-recurrent neural networks) are the only computational models that can reproduce the long memory behavior of natural language. Furthermore, through comparison with recently proposed model-based evaluation methods, we find that the exponent of Taylor’s law is a good indicator of model quality.
Shuntaro Takahashi, Kumiko Tanaka-Ishii
Comput. Linguistics2
2018 Taylor's law for Human Linguistic Sequences
abstract
Taylor's law describes the fluctuation characteristics underlying a system in which the variance of an event within a time span grows by a power law with respect to the mean.Although Taylor's law has been applied in many natural and social systems, its application for language has been scarce.This article describes a new quantification of Taylor's law in natural language and reports an analysis of over 1100 texts across 14 languages.The Taylor exponents of written natural language texts were found to exhibit almost the same value.The exponent was also compared for other language-related data, such as the child-directed speech, music, and programming language code.The results show how the Taylor exponent serves to quantify the fundamental structural complexity underlying linguistic time series.The article also shows the applicability of these findings in evaluating language models.
Tatsuru Kobayashi, Kumiko Tanaka-Ishii
ACL (1)2
2018 Extraction of templates from phrases using Sequence Binary Decision Diagrams
abstract
Abstract The extraction of templates such as ‘regard X as Y’ from a set of related phrases requires the identification of their internal structures. This paper presents an unsupervised approach for extracting templates on-the-fly from only tagged text by using a novel relaxed variant of the Sequence Binary Decision Diagram (SeqBDD). A SeqBDD can compress a set of sequences into a graphical structure equivalent to a minimal deterministic finite state automata, but more compact and better suited to the task of template extraction. The main contribution of this paper is a relaxed form of the SeqBDD construction algorithm that enables it to form general representations from a small amount of data. The process of compression of shared structures in the text during Relaxed SeqBDD construction, naturally induces the templates we wish to extract. Experiments show that the method is capable of high-quality extraction on tasks based on verb+preposition templates from corpora and phrasal templates from short messages from social media.
Daiki Hirano, Kumiko Tanaka-Ishii, Andrew M. Finch
Nat. Lang. Eng.2
2017 Component Awareness in Convolutional Neural Networks
abstract
In this work, we investigate the ability of Convolutional Neural Networks (CNN) to infer the presence of components that comprise an image. In recent years, CNNs have achieved powerful results in classification, detection, and segmentation. However, these models learn from instance-level supervision of the detected object. In this paper, we determine if CNNs can detect objects using image-level weakly supervised labels without localization. To demonstrate that a CNN can infer awareness of objects, we evaluate a CNN's classification ability with a database constructed of Chinese characters with only character-level labeled components. We show that the CNN is able to achieve a high accuracy in identifying the presence of these components without specific knowledge of the component. Furthermore, we verify that the CNN is deducing the knowledge of the target component by comparing the results to an experiment with the component removed. This research is important for applications with large amounts of data without robust annotation such as Chinese character recognition.
Brian Kenji Iwana, Letao Zhou, Kumiko Tanaka-Ishii, Seiichi Uchida
ICDAR3
2017 Inducing a Bilingual Lexicon from Short Parallel Multiword Sequences
abstract
This article proposes a technique for mining bilingual lexicons from pairs of parallel short word sequences. The technique builds a generative model from a corpus of training data consisting of such pairs. The model is a hierarchical nonparametric Bayesian model that directly induces a bilingual lexicon while training. The model learns in an unsupervised manner and is designed to exploit characteristics of the language pairs being mined. The proposed model is capable of utilizing commonly used word-pair frequency information and additionally can employ the internal character alignments within the words themselves. It is thereby capable of mining transliterations and can use reliably aligned transliteration pairs to support the mining of other words in their context. The model is also capable of performing word reordering and word deletion during the alignment process, and it is furthermore capable of operating in the absence of full segmentation information. In this work, we study two mining tasks based on English-Japanese and English-Chinese language pairs, and compare the proposed approach to baselines based on a simpler models that use only word-pair frequency information. Our results show that the proposed method is able to mine bilingual word pairs at higher levels of precision and recall than the baselines.
Andrew M. Finch, Taisuke Harada, Kumiko Tanaka-Ishii, Eiichiro Sumita
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2015 Computational Constancy Measures of Texts - Yule's K and Rényi's Entropy
abstract
This article presents a mathematical and empirical verification of computational constancy measures for natural language text. A constancy measure characterizes a given text by having an invariant value for any size larger than a certain amount. The study of such measures has a 70-year history dating back to Yule's K, with the original intended application of author identification. We examine various measures proposed since Yule and reconsider reports made so far, thus overviewing the study of constancy measures. We then explain how K is essentially equivalent to an approximation of the second-order Rényi entropy, thus indicating its signification within language science. We then empirically examine constancy measure candidates within this new, broader context. The approximated higher-order entropy exhibits stable convergence across different languages and kinds of text. We also show, however, that it cannot identify authors, contrary to Yule's intention. Lastly, we apply K to two unknown scripts, the Voynich manuscript and Rongorongo, and show how the results support previous hypotheses about these scripts.
Kumiko Tanaka-Ishii, Shunsuke Aihara
Comput. Linguistics1
2013 picoTrans: An intelligent icon-driven interface for cross-lingual communication
abstract
picoTrans is a prototype system that introduces a novel icon-based paradigm for cross-lingual communication on mobile devices. Our approach marries a machine translation system with the popular picture book. Users interact with picoTrans by pointing at pictures as if it were a picture book; the system generates natural language from these icons and the user is able to interact with the icon sequence to refine the meaning of the words that are generated. When users are satisfied that the sentence generated represents what they wish to express, they tap a translate button and picoTrans displays the translation. Structuring the process of communication in this way has many advantages. First, tapping icons is a very natural method of user input on mobile devices; typing is cumbersome and speech input errorful. Second, the sequence of icons which is annotated both with pictures and bilingually with words is meaningful to both users, and it opens up a second channel of communication between them that conveys the gist of what is being expressed. We performed a number of evaluations of picoTrans to determine: its coverage of a corpus of in-domain sentences; the input efficiency in terms of the number of key presses required relative to text entry; and users' overall impressions of using the system compared to using a picture book. Our results show that we are able to cover 74% of the expressions in our test corpus using a 2000-icon set; we believe that this icon set size is realistic for a mobile device. We also found that picoTrans requires fewer key presses than typing the input and that the system is able to predict the correct, intended natural language sentence from the icon sequence most of the time, making user interaction with the icon sequence often unnecessary. In the user evaluation, we found that in general users prefer using picoTrans and are able to communicate more rapidly and expressively. Furthermore, users had more confidence that they were able to communicate effectively using picoTrans.
Andrew M. Finch, Kumiko Tanaka-Ishii, Keiji Yasuda, Eiichiro Sumita
ACM Trans. Interact. Intell. Syst.3
2012 Text Segmentation by Language Using Minimum Description Length
Hiroshi Yamaguchi, Kumiko Tanaka-Ishii
ACL (1)2
2011 picoTrans: Using Pictures as Input for Machine Translation on Mobile Devices
abstract
In this paper we present a novel user interface that integrates two popular approaches to language translation for travelers allowing multimodal communication between the parties involved: the picture-book, in which the user simply points to multiple picture icons representing what they want to say, and the statistical machine translation (SMT) system that can translate arbitrary word sequences. Our prototype system tightly couples both processes within a translation framework that inherits many of the the positive features of both approaches, while at the same time mitigating their main weaknesses. Our system differs from traditional approaches in that its mode of input is a sequence of pictures, rather than text or speech. Text in the source language is generated automatically, and is used as a detailed representation of the intended meaning. The picture sequence which not only provides a rapid method to communicate basic concepts but also gives a 'second opinion' on the machine transition output that catches machine translation errors and allows the users to retry the translation, avoiding misunderstandings.
Andrew M. Finch, Kumiko Tanaka-Ishii, Eiichiro Sumita
IJCAI3
2011 Relational Lasso - An Improved Method Using the Relations Among Features -
Kotaro Kitagawa, Kumiko Tanaka-Ishii
IJCNLP2
2011 picoTrans: an icon-driven user interface for machine translation on mobile devices
abstract
In this paper we present a novel user interface that integrates two popular approaches to language translation for travelers allowing multimodal communication between the parties involved. In our approach we integrate the popular picture-book, in which the user simply points to multiple picture icons representing what they want to say, with a statistical machine translation system that can translate arbitrary word sequences. The simple pointing at pictures paradigm is used as the primary method of user input and the users can use the device as if it were a picture book. The application is then able to generate a complete sentence in the user's native language for what they wish to say from the sequence of picture icons chosen by the user. Once the user is satisfied that the sentence provided by the system adequately represents what they wish to convey, the application can automatically translate the sentence into the language of the other party, who can interpret the intended meaning of the first party by combining evidence from both modes of communication: the picture sequence, and the machine translation. The prototype system we have developed inherits many of the positive features of both approaches, while at the same time mitigating their main weaknesses. The user may combine the pictures in considerably more combinations than is possible with a picture book designed with combinations from only within the same page spread of the book in mind, making the application more expressive than a book. The machine translation system can contribute a detailed and precise translation which is supported by the picture-based mode which not only provides a rapid method to communicate basic concepts but also gives a 'second opinion' on the machine transition output that catches machine translation errors and allows the users to retry the sentence, avoiding misunderstandings.
Andrew M. Finch, Kumiko Tanaka-Ishii, Eiichiro Sumita
IUI3
2010 YouBot: A Simple Framework for Building Virtual Networking Agents
Seiji Takegata, Kumiko Tanaka-Ishii
SIGDIAL Conference2
2010 Sorting Texts by Readability
abstract
This article presents a novel approach for readability assessment through sorting. A comparator that judges the relative readability between two texts is generated through machine learning, and a given set of texts is sorted by this comparator. Our proposal is advantageous because it solves the problem of a lack of training data, because the construction of the comparator only requires training data annotated with two reading levels. The proposed method is compared with regression methods and a state-of-the art classification method. Moreover, we present our application, called Terrace, which retrieves texts with readability similar to that of a given input text.
Kumiko Tanaka-Ishii, Satoshi Tezuka, Hiroshi Terada
Comput. Linguistics1
2009 Multilingual Spectral Clustering Using Document Similarity Propagation
Dani Yogatama, Kumiko Tanaka-Ishii
EMNLP2
2009 Kansuke: A logograph look-up interface based on a few modified stroke prototypes
abstract
We have developed a method that makes it easier for language novices to look up Japanese and Chinese logographs. Instead of using the arbitrary conventions of logographs, this method is based on three simple prototypes: horizontal, vertical, and other strokes. For example, the code for the logograph ⊞ ( ta , meaning rice field) is 3-3-0, indicating the logograph consists of three horizontal strokes and three vertical strokes. Such codes allow a novice to look up logographs even with no knowledge of the logographic conventions used by native speakers. To make the search easier, a complex logograph can be looked up via the components making up the logograph. We conducted a user evaluation of this system and found that novices could look up logographs with fewer failures with our system than with conventional methods.
Kumiko Tanaka-Ishii, Julian Godon
ACM Trans. Comput. Hum. Interact.1
2008 Multilingual Text Entry using Automatic Language Detection
Yo Ehara, Kumiko Tanaka-Ishii
IJCNLP2
2007 Introduction to the Special Issue on Recent PACLING Meetings
Kiyoshi Kogure, Shun Ishizaki, Tsutomu Endo, Yoshihiko Nitta, Kumiko Tanaka-Ishii
Comput. Intell.5
2007 Multilingual phrase-based concordance generation in real-time
Kumiko Tanaka-Ishii, Yuichiro Ishii
Inf. Retr.1
2007 Word-based predictive text entry using adaptive language models
abstract
The recent scaling down of mobile device form factors has increased the importance of predictive text entry. It is now also becoming an important communication tool for the disabled. Techniques related to predictive text entry software are discussed in a generalized, language-independent manner. The essence of predictive text entry is twofold, consisting of (1) the design of codes for text entry, and (2) the use of adaptive language models for decoding. Code design is examined in terms of the information-theoretical efficiency. Four adaptive language models are introduced and compared, and experimental results on text entry with these models are shown for English, Thai and Japanese.
Kumiko Tanaka-Ishii
Nat. Lang. Eng.1
2006 Unsupervised Segmentation of Chinese Text by Use of Branching Entropy
Zhihui Jin, Kumiko Tanaka-Ishii
ACL2
2005 Entropy as an Indicator of Context Boundaries: An Experiment Using a Web Search Engine
Kumiko Tanaka-Ishii
IJCNLP1
2005 A multilingual usage consultation tool based on internet searching: more than a search engine, less than QA
abstract
We present a usage consultation tool, based on Internet searching, for language learners. When a user enters a string of words for which he wants to find usages, the system sends this string as a query to a search engine and obtains search results about the string. The usages are extracted by performing statistical analysis on snippets and then fed back to the user.Unlike existing tools, this usage consultation tool is multi-lingual, so that usages can be obtained even in a language for which there are no well-established analytical methods. Our evaluation has revealed that usages can be obtained more effectively than by only using a search engine directly. Also, we have found that the resulting usage does not depend on the search engine for a prominent usage when the amount of data downloaded from the search engine is increased.
Kumiko Tanaka-Ishii, Hiroshi Nakagawa
WWW1
2004 An Interactive Proofreading System for Inappropriately Selected Words on Using Predictive Text Entry
Hideya Iwasaki, Kumiko Tanaka-Ishii
IJCNLP2
2004 Introduction (Thematic Session: Natural Language Technology in Mobile Information Retrieval and Text Processing User Interfaces)
Michael Kuehn, Mun-Kew Leong, Kumiko Tanaka-Ishii
IJCNLP3
2004 Dit4dah: Predictive Pruning for Morse Code Text Entry
Kumiko Tanaka-Ishii, Ian Frank
IJCNLP1
2004 EMMA: a web-based report system for programming course--automated verification and enhanced feedback
abstract
No abstract available.
Kumiko Tanaka-Ishii, Kazuhiko Kakehi 0001, Masato Takeichi
ITiCSE1
2003 Acquiring Vocabulary for Predictive Text Entry through Dynamic Reuse of a Small User Corpus
abstract
As mobile computing and communications have become popular, predictive text entry systems have become an increasingly important technology. Existing methods still need refinement, though, with respect to personalization, especially how to acquire vocabulary not pre-registered in the system dictionary. In this paper, we report on an automatic method that dynamically obtains a user specific vocabulary from the user's unanalyzed documents. When a user makes an entry, the system dynamically extracts the corresponding chunks from the user text and suggests them along with words suggested by the dictionary. With our method, texts in a particular style or concerning a specific domain can be entered using a predictive text entry system. We verified that a large amount of words not registered in the dictionary can be entered using our method.
Kumiko Tanaka-Ishii, Daichi Hayakawa, Masato Takeichi
ACL1
2003 Performance Competitions as Research Infrastructure: Large Scale Comparative Studies of Multi-Agent Teams
Gal A. Kaminka, Ian Frank, Katsuto Arai, Kumiko Tanaka-Ishii
Auton. Agents Multi Agent Syst.4
2002 Entering Text with a Four-Button Device
Kumiko Tanaka-Ishii, Yusuke Inutsuka, Masato Takeichi
COLING1
2001 Walkie-Talkie MIKE
Ian Frank, Kumiko Tanaka-Ishii, Hitoshi Matsubara, Eiichi Osawa
RoboCup2
2000 Multi-Agent Explanation Strategies in Real-Time Domains
abstract
We examine the benefits of using multiple agents to produce explanations. In particular, we identify the ability to construct prior plans as a key issue constraining the effectiveness of a single-agent approach. We describe an implemented system that uses multiple agents to tackle a problem for which prior planning is particularly impractical: real-time soccer commentary. Our commentary system demonstrates a number of the advantages of decomposing an explanation task among several agents. Most notably, it shows how individual agents can benefit from following different discourse strategies. Further, it illustrates that discourse issues such as controlling interruption, abbreviation, and maintaining consistency can also be decomposed: rather than considering them at the single level of one linear explanation they can also be tackled separately within each individual agent. We evaluate our system's output, and show that it closely compares to the speaking patterns of a human commentary team.
Kumiko Tanaka-Ishii, Ian Frank
ACL1
2000 The Statistics Proxy Server
Ian Frank, Kumiko Tanaka-Ishii, Katsuto Arai, Hitoshi Matsubara
RoboCup2
2000 And the Fans Are Going Wild! SIG plus MIKE
Ian Frank, Kumiko Tanaka-Ishii, Hiroshi G. Okuno, Junichi Akita, Yukiko Nakagawa, Kazuaki Maeda, Kazuhiro Nakadai, Hiroaki Kitano
RoboCup2
1999 A Statistical Perspective on the RoboCup Simulator League: Progress and Prospects
Kumiko Tanaka-Ishii, Ian Frank, Itsuki Noda, Hitoshi Matsubara
RoboCup1
1998 Automatic Soccer Commentary and RoboCup
Hitoshi Matsubara, Itsuki Noda, Ian Frank, Hideyuki Nakashima, Kumiko Tanaka-Ishii, Kôiti Hasida
RoboCup5
1996 Extraction of Lexical Translations from Non-Aligned Corpora
Kumiko Tanaka-Ishii, Hideya Iwasaki
COLING1
1994 Construction of a Bilingual Dictionary Intermediated by a Third Language
Kumiko Tanaka-Ishii, Kyoji Umemura
COLING1