Hideki Tanaka

dblp:55/3834 · DBLP profile ↗
← Back
35ranked-venue papers
6as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Registering Source Tokens to Target Language Spaces in Multilingual Neural Machine Translation
abstract
Zhi Qu, Yiran Wang, Jiannan Mao, Chenchen Ding, Hideki Tanaka, Masao Utiyama, Taro Watanabe. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhi Qu 0001, Yiran Wang 0006, Jiannan Mao, Chenchen Ding, Hideki Tanaka, Masao Utiyama, Taro Watanabe
ACL (1)5
2025 PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
abstract
This work introduces PrahokBART, a compact pre-trained sequence-to-sequence model trained from scratch for Khmer using carefully curated Khmer and English corpora. We focus on improving the pre-training corpus quality and addressing the linguistic issues of Khmer, which are ignored in existing multilingual models, by incorporating linguistic components such as word segmentation and normalization. We evaluate PrahokBART on three generative tasks: machine translation, text summarization, and headline generation, where our results demonstrate that it outperforms mBART50, a strong multilingual pre-trained model. Additionally, our analysis provides insights into the impact of each linguistic module and evaluates how effectively our model handles space during text generation, which is crucial for the naturalness of texts in Khmer.
Hour Kaing, Raj Dabre, Haiyue Song, Van-Hien Tran, Hideki Tanaka, Masao Utiyama
COLING5
2025 Personality Trait Analysis System Using Dialogue Simulation with Autonomous Agents
abstract
In personality trait analysis, self-evaluation through questionnaire methods is widely used due to ease of implementation and analysis. However, respondents can easily provide false evaluation scores based on their intentions. This study proposes a method to obtain truthful evaluation scores from dialogue records through dialogue simulation with autonomous agents for questionnaire items. Since dialogue requires immediate responses and detailed explanations, we hypothesized this method would be less prone to deception compared to traditional questionnaire methods. We developed a comprehensive system using Large Language Models that handles everything from setting dialogue situations based on questionnaire items to conducting dialogues with autonomous agents, and automatically evaluating dialogue records. In preliminary experiments, we found that evaluation scores from the proposed method when participants were prompted to lie were close to truthful questionnaire responses when participants were prompted to answer truthfully. These findings suggest the effectiveness of psychological measurement through dialogue with autonomous agents.
Kei Hyodo, Tomonori Kubota, Satoshi Sato, Hideki Tanaka, Kohei Ogawa
HAI4
2025 Tikzero: Zero-Shot Text-Guided Graphics Program Synthesis
Jonas Belouadi, Eddy Ilg, Margret Keuper, Hideki Tanaka, Masao Utiyama, Raj Dabre, Steffen Eger, Simone Paolo Ponzetto
ICCV4
2024 NGLUEni: Benchmarking and Adapting Pretrained Language Models for Nguni Languages
abstract
The Nguni languages have over 20 million home language speakers in South Africa. There has been considerable growth in the datasets for Nguni languages, but so far no analysis of the performance of NLP models for these languages has been reported across languages and tasks. In this paper we study pretrained language models for the 4 Nguni languages - isiXhosa, isiZulu, isiNdebele, and Siswati. We compile publicly available datasets for natural language understanding and generation, spanning 6 tasks and 11 datasets. This benchmark, which we call NGLUEni, is the first centralised evaluation suite for the Nguni languages, allowing us to systematically evaluate the Nguni-language capabilities of pretrained language models (PLMs). Besides evaluating existing PLMs, we develop new PLMs for the Nguni languages through multilingual adaptive finetuning. Our models, Nguni-XLMR and Nguni-ByT5, outperform their base models and large-scale adapted models, showing that performance gains are obtainable through limited language group-based adaptation. We also perform experiments on cross-lingual transfer and machine translation. Our models achieve notable cross-lingual transfer improvements in the lower resourced Nguni languages (isiNdebele and Siswati). To facilitate future use of NGLUEni as a standardised evaluation suite for the Nguni languages, we create a web portal to access the collection of datasets and publicly release our models.
Francois Meyer, Haiyue Song, Abhisek Chakrabarty, Jan Buys, Raj Dabre, Hideki Tanaka
LREC/COLING6
2024 SubMerge: Merging Equivalent Subword Tokenizations for Subword Regularized Models in Neural Machine Translation
abstract
Subword regularized models leverage multiple subword tokenizations of one target sentence during training. However, selecting one tokenization during inference leads to the underutilization of knowledge learned about multiple tokenizations.We propose the SubMerge algorithm to rescue the ignored Subword tokenizations through merging equivalent ones during inference.SubMerge is a nested search algorithm where the outer beam search treats the word as the minimal unit, and the inner beam search provides a list of word candidates and their probabilities, merging equivalent subword tokenizations. SubMerge estimates the probability of the next word more precisely, providing better guidance during inference.Experimental results on six low-resource to high-resource machine translation datasets show that SubMerge utilizes a greater proportion of a model’s probability weight during decoding (lower word perplexities for hypotheses). It also improves BLEU and chrF++ scores for many translation directions, most reliably for low-resource scenarios. We investigate the effect of different beam sizes, training set sizes, dropout rates, and whether it is effective on non-regularized models.
Haiyue Song, Francois Meyer, Raj Dabre, Hideki Tanaka, Chenhui Chu, Sadao Kurohashi
EAMT (1)4
2023 Subset Retrieval Nearest Neighbor Machine Translation
abstract
Hiroyuki Deguchi, Taro Watanabe, Yusuke Matsui, Masao Utiyama, Hideki Tanaka, Eiichiro Sumita. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Hiroyuki Deguchi 0002, Taro Watanabe, Yusuke Matsui 0001, Masao Utiyama, Hideki Tanaka, Eiichiro Sumita
ACL (1)5
2023 A Study on the Effectiveness of Large Language Models for Translation with Markup
abstract
In this paper we evaluate the utility of large language models (LLMs) for translation of text with markup in which the most important and challenging aspect is to correctly transfer markup tags while ensuring that the content, both, inside and outside tags is correctly translated. While LLMs have been shown to be effective for plain text translation, their effectiveness for structured document translation is not well understood. To this end, we experiment with BLOOM and BLOOMZ, which are open-source multilingual LLMs, using zero, one and few-shot prompting, and compare with a domain-specific in-house NMT system using a detag-and-project approach for markup tags. We observe that LLMs with in-context learning exhibit poorer translation quality compared to the domain-specific NMT system, however, they are effective in transferring markup tags, especially the large BLOOM model (176 billion parameters). This is further confirmed by our human evaluation which also reveals the types of errors of the different tag transfer techniques. While LLM-based approaches come with the risk of losing, hallucinating and corrupting tags, they excel at placing them correctly in the translation.
Raj Dabre, Bianka Buschbeck-Wolf, Miriam Exel, Hideki Tanaka
MTSummit (1)4
2023 Improving Embedding Transfer for Low-Resource Machine Translation
abstract
Low-resource machine translation (LRMT) poses a substantial challenge due to the scarcity of parallel training data. This paper introduces a new method to improve the transfer of the embedding layer from the Parent model to the Child model in LRMT, utilizing trained token embeddings in the Parent model’s high-resource vocabulary. Our approach involves projecting all tokens into a shared semantic space and measuring the semantic similarity between tokens in the low-resource and high-resource languages. These measures are then utilized to initialize token representations in the Child model’s low-resource vocabulary. We evaluated our approach on three benchmark datasets of low-resource language pairs: Myanmar-English, Indonesian-English, and Turkish-English. The experimental results demonstrate that our method outperforms previous methods regarding translation quality. Additionally, our approach is computationally efficient, leading to reduced training time compared to prior works.
Van-Hien Tran, Chenchen Ding, Hideki Tanaka, Masao Utiyama
MTSummit (1)3
2023 Improving Zero-Shot Dependency Parsing by Unsupervised Learning
Jiannan Mao, Chenchen Ding, Hour Kaing, Hideki Tanaka, Masao Utiyama, Tadahiro Matsumoto
PACLIC4
2023 Refining History for Future-Aware Neural Machine Translation
abstract
Neural machine translation uses a decoder to generate target words auto-regressively by predicting the next target word conditioned on a given source sentence and its previously predicted target words, i.e, its translation history, which suffers from two limitations: 1) the prediction of next word depends heavily on the quality of its history information. Moreover, the discrepancy between training and inference exacerbates this limitation; 2) this left-to-right decoding way cannot make full use of the target-side future information, which leads to the issue of unbalanced outputs. On the one hand, we alleviate the first limitation with a history-refining module, which learns to examine the quality of each history word by assigning it a confidence score. The confidence score is further used as a gate to control the amount of its word embedding flowing to the decoder. On the other hand, we attack the second limitation with a future-foreseeing module, which learns the distribution of future translation at each decoding time step. More importantly, we further propose refining history for future-aware NMT since the two modules can be closely incorporated as they focus on different kinds of context. Experimental results on various translation tasks with different scaled datasets, including WMT English$\leftrightarrow${German, French, Romanian}, show that our proposed approach achieves significant improvements over strong Transformer-based NMT baselines.
Xinglin Lyu, Junhui Li 0001, Min Zhang 0005, Chenchen Ding, Hideki Tanaka, Masao Utiyama
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 FeatureBART: Feature Based Sequence-to-Sequence Pre-Training for Low-Resource NMT
abstract
In this paper we present FeatureBART, a linguistically motivated sequence-to-sequence monolingual pre-training strategy in which syntactic features such as lemma, part-of-speech and dependency labels are incorporated into the span prediction based pre-training framework (BART). These automatically extracted features are incorporated via approaches such as concatenation and relevance mechanisms, among which the latter is known to be better than the former. When used for low-resource NMT as a downstream task, we show that these feature based models give large improvements in bilingual settings and modest ones in multilingual settings over their counterparts that do not use features.
Abhisek Chakrabarty, Raj Dabre, Chenchen Ding, Hideki Tanaka, Masao Utiyama, Eiichiro Sumita
COLING4
2020 Content-Equivalent Translated Parallel News Corpus and Extension of Domain Adaptation for NMT
abstract
In this paper, we deal with two problems in Japanese-English machine translation of news articles. The first problem is the quality of parallel corpora. Neural machine translation (NMT) systems suffer degraded performance when trained with noisy data. Because there is no clean Japanese-English parallel data for news articles, we build a novel parallel news corpus consisting of Japanese news articles translated into English in a content-equivalent manner. This is the first content-equivalent Japanese-English news corpus translated specifically for training NMT systems. The second problem involves the domain-adaptation technique. NMT systems suffer degraded performance when trained with mixed data having different features, such as noisy data and clean data. Though the existing methods try to overcome this problem by using tags for distinguishing the differences between corpora, it is not sufficient. We thus extend a domain-adaptation method using multi-tags to train an NMT model effectively with the clean corpus and existing parallel news corpora with some types of noise. Experimental results show that our corpus increases the translation quality, and that our domain-adaptation method is more effective for learning with the multiple types of corpora than existing domain-adaptation methods are.
Hideya Mino, Hideki Tanaka, Hitoshi Ito, Isao Goto, Ichiro Yamada, Takenobu Tokunaga
LREC2
2015 Japanese news simplification: tak design, data set construction, and analysis of simplified text
Isao Goto, Hideki Tanaka, Tadashi Kumano
MTSummit2
2012 Measuring the Similarity between TV Programs using Semantic Relations
Ichiro Yamada, Masaru Miyazaki, Hideki Sumiyoshi, Atsushi Matsui, Hironori Furumiya, Hideki Tanaka
COLING6
2011 A VLSI Spiking Neural Network with Symmetric STDP and Associative Memory Operation
Frank L. Maldonado Huayaney, Hideki Tanaka, Takayuki Matsuo, Takashi Morie, Kazuyuki Aihara
ICONIP (3)2
2010 Relevant TV program retrieval using broadcast summaries
abstract
On-demand services for TV program, which provide users with past programs on demand, are becoming popular. It is therefore necessary to find a means of efficiently searching for programs that users want to view, from huge program archives. This paper proposes an automatic method of retrieving programs related to the one being viewed by the user. To that end, we compute similarity between program summaries and closed captions obtained from broadcasting by weighting significant words such as compound words and named entities. Additionally our method provides inter-program relationship labels to indicate why the results of relevant programs were chosen. The results of an evaluation showed that the method recommended relevant programs with higher accuracy than baseline methods and indicated appropriate relationship labels for related programs.
Jun Goto, Hideki Sumiyoshi, Masaru Miyazaki, Hideki Tanaka, Masahiro Shibata, Akiko Aizawa
IUI4
2009 Opinion classification with tree kernel SVM using linguistic modality analysis
abstract
We propose a method for classifying opinions which captures the role of linguistic modalities in the sentence. We use features than simple bag-of-words or opinion-holding predicates. The method is based on a machine learning and utilizes opinion-holding predicates and linguistic modalities as features. Two different detectors help to classify the opinions: the opinion-holding predicate detector and the modality detector. An opinion in the target is first parsed into a dependency structure, and then the opinion-holding predicates and modalities stick onto the leaf nodes of the dependency tree. The whole tree is regarded as input features of the opinion, and it becomes the input of tree kernel support vector machines. We have applied method to opinions in Japanese about television programs, and have confirmed the effectiveness of the method against conventional bag-of-words features, or against simple opinion-holding predicates features
Takeshi S. Kobayakawa, Tadashi Kumano, Hideki Tanaka, Naoaki Okazaki, Jin-Dong Kim, Jun'ichi Tsujii
CIKM3
2004 Back Transliteration from Japanese to English using Target English Context
Isao Goto, Naoto Katoh, Terumasa Ehara, Hideki Tanaka
COLING4
2004 Example-Based Machine Translation Without Saying Inferable Predicate
Eiji Aramaki, Sadao Kurohashi, Hideki Kashioka, Hideki Tanaka
IJCNLP4
2004 Acquiring Bilingual Named Entity Translations from Content-Aligned Corpora
Tadashi Kumano, Hideki Kashioka, Hideki Tanaka, Takahiro Fukusima
IJCNLP3
2003 The Word Is Mightier Than the Count: Accumulating Translation Resources from Parsed Parallel Corpora
Stephen Nightingale, Hideki Tanaka
CICLing2
2003 Building a parallel corpus for monologues with clause alignment
abstract
Many studies have been reported in the domain of speech-to-speech machine translation systems for travel conversation use. Therefore, a large number of travel domain corpora have become available in recent years. From a wider viewpoint, speech-to-speech systems are required for many purposes other than travel conversation. One of these is monologues (e.g., TV news, lectures, technical presentations). However, in monologues, sentences tend to be long and complicated, which often causes problems for parsing and translation. Therefore, we need a suitable translation unit, rather than the sentence. We propose the clause as a unit for translation. To develop a speech-to-speech machine translation system for monologues based on the clause as the translation unit, we need a monologue parallel corpus with clause alignment. In this paper, we describe how to build a Japanese-English monologue parallel corpus with clauses aligned, and discuss the features of this corpus.
Hideki Kashioka, Takehiko Maruyama, Hideki Tanaka
MTSummit3
2002 Speech to speech translation system for monologues-data driven approach
Hideki Tanaka, Stephen Nightingale, Hideki Kashioka, Kenji Matsumoto, Masamichi Nishiwaki, Tadashi Kumano, Takehiko Maruyama
INTERSPEECH1
2002 Automatic Alignment of Japanese and English Newspaper Articles using an MT System and a Bilingual Company Name Dictionary
Kenji Matsumoto, Hideki Tanaka
LREC2
2000 Progressive 2-pass decoder for real-time broadcast news captioning
abstract
This paper describes a 2-pass decoder that progressively outputs the latest available results used for real-time closed captioning of Japanese broadcast news. The decoder practically eliminates the disadvantage of multiple-pass decoders that delay a decision until the end of a sentence. During the first pass of search the proposed decoder periodically executes the second pass that rescores partial N-best word sequences up to that time. If the rescored best word sequence has words in common with the previous one, that part is regarded as likely to be correct and is decided to be a part of the final result. This method is not theoretically optimal but makes a quick response with a negligible increase in word errors. In a recognition experiment on Japanese broadcast news, the decoder worked with an average decision delay of 554 msec for each word and degraded word accuracy only by 0.22%.
Toru Imai, Akio Kobayashi, Shoei Sato, Hideki Tanaka, Akio Ando
ICASSP4
2000 Selective training of HMMs by using two-stage clustering
Shoei Sato, Toru Imai, Hideki Tanaka, Akio Ando
INTERSPEECH3
1999 An Efficient Statistical Speech Act Type Tagging System for Speech Translation Systems
abstract
This paper describes a new efficient speech act type tagging system. This system covers the tasks of (1) segmenting a turn into the optimal number of speech act units (SA units), and (2) assigning a speech act type tag (SA tag) to each SA unit. Our method is based on a theoretically clear statistical model that integrates linguistic, acoustic and situational information. We report tagging experiments on Japanese and English dialogue corpora manually labeled with SA tags. We then discuss the performance difference between the two languages. We also report on some translation experiments on positive response expressions using SA tags.
Hideki Tanaka, Akio Yokoo
ACL1
1999 An efficient document clustering algorithm and its application to a document browser
Hideki Tanaka, Tadashi Kumano, Noriyoshi Uratani, Terumasa Ehara
Inf. Process. Manag.1
1998 Planning Dialogue Contributions With New Information
Kristiina Jokinen, Hideki Tanaka, Akio Yokoo
INLG2
1997 Super low power 8-bit CPU with pass-transistor logic
abstract
A very low power 8-bit CPU core has been designed based on an original pass-transistor logic family, the SPL (single-rail pass-transistor logic) and SPHL (single-rail pass-transistor and holders logic). The instruction set and external timings are compatible with the Zilog Z80. The average supply current is 740 /spl mu/A at 3 V with a 10 MHz-clock, equivalent to 26% of that of the commercial CMOS Z80 CPU cores using the same design rules.
Kazuo Taki, Bu-Yeol Lee, Hideki Tanaka, Kenzo Konishi
ASP-DAC3
1996 Decision Tree Learning Algorithm with Structured Attributes: Application to Verbal Case Frame Acquisition
Hideki Tanaka
COLING1
1994 Verbal Case Frame Acquisition From A Bilingual Corpus: Gradual Knowledge Acquisition
Hideki Tanaka
COLING1
1992 A Method Of Translating English Delexical Structures Into Japanese
Hideki Tanaka, Teruaki Aizawa, Yeun-Bae Kim, Nobuko Hatada
COLING1
1990 A Machine Translation System for Foreign News in Satellite Broadcasting
Teruaki Aizawa, Terumasa Ehara, Noriyoshi Uratani, Hideki Tanaka, Naoto Katoh, Sumio Nakase, Norikazu Aruga, Takeo Matsuda
COLING4