Katsuhito Sudoh

dblp:66/2380 · DBLP profile ↗
← Back
46ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-2122-9846ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 6 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
Anum Afzal, Yuki Saito 0001, Hiroya Takamura, Katsuhito Sudoh, Shinnosuke Takamichi, Graham Neubig, Florian Matthes, Tatsuya Ishigaki
LREC4
2025 Measuring Time Delay Tolerance in Third-Person Live Commentary for Super Smash Bros. Ultimate
abstract
This study proposes a methodology for measuring the acceptable delay tolerance for third-person game commentary. Third-person game commentary refers to commentary delivered by someone other than the player, with the role of helping viewers better understand the game and enhancing the viewing experience. With the recent advancement of AI, there has been increasing interest in automating such commentary using video understanding and audio generation. However, automating this process using video understanding and audio generation introduces delays, potentially affecting the naturalness of the commentary. In this context, since the extent to which such delays are acceptable to viewers remains unclear, we address this issue. The tolerance is modeled using an unnormalized Gaussian function. Through experiments on Super Smash Bros. Ultimate with 727 participants, we found that the average acceptable delay for this game is 3.71 seconds, with variations depending on different viewer attributes and gameplay contexts.
Ryosuke Matsushita, Ryosuke Sakai, Koki Fukuda, Shinnosuke Takamichi, Kota Iura, Yuki Saito 0001, Graham Neubig, Katsuhito Sudoh, Hiroya Takamura, Tatsuya Ishigaki
CoG8
2024 Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation
abstract
Human evaluation is a critical component in machine translation system development and has received much attention in text translation research. However, little prior work exists on the topic of human evaluation for speech translation, which adds additional challenges such as noisy data and segmentation mismatches. We take the first steps to fill this gap by conducting a comprehensive human evaluation of the results of several shared tasks from the last International Workshop on Spoken Language Translation (IWSLT 2023). We propose an effective evaluation strategy based on automatic resegmentation and direct assessment with segment context. Our analysis revealed that: 1) the proposed evaluation strategy is robust and scores well-correlated with other types of human judgements; 2) automatic metrics are usually, but not always, well-correlated with direct assessment scores; and 3) COMET as a slightly stronger automatic metric than chrF, despite the segmentation noise introduced by the resegmentation step systems. We release the collected human-annotated data in order to encourage further investigation.
Matthias Sperber, Ondrej Bojar, Barry Haddow, Dávid Javorský, Xutai Ma, Matteo Negri, Jan Niehues, Peter Polak, Elizabeth Salesky, Katsuhito Sudoh, Marco Turchi
LREC/COLING10
2024 NAIST-SIC-Aligned: An Aligned English-Japanese Simultaneous Interpretation Corpus
abstract
It remains a question that how simultaneous interpretation (SI) data affects simultaneous machine translation (SiMT). Research has been limited due to the lack of a large-scale training corpus. In this work, we aim to fill in the gap by introducing NAIST-SIC-Aligned, which is an automatically-aligned parallel English-Japanese SI dataset. Starting with a non-aligned corpus NAIST-SIC, we propose a two-stage alignment approach to make the corpus parallel and thus suitable for model training. The first stage is coarse alignment where we perform a many-to-many mapping between source and target sentences, and the second stage is fine-grained alignment where we perform intra- and inter-sentence filtering to improve the quality of aligned pairs. To ensure the quality of the corpus, each step has been validated either quantitatively or qualitatively. This is the first open-sourced large-scale parallel SI dataset in the literature. We also manually curated a small test set for evaluation purposes. Our results show that models trained with SI data lead to significant improvement in translation quality and latency over baselines. We hope our work advances research on SI corpora construction and SiMT. Our data will be released upon the paper’s acceptance.
Jinming Zhao, Katsuhito Sudoh, Satoshi Nakamura 0001, Yuka Ko, Kosuke Doi, Ryo Fukuda
LREC/COLING2
2024 LLMs Are Zero-Shot Context-Aware Simultaneous Translators
abstract
The advent of transformers has fueled progress in machine translation.More recently large language models (LLMs) have come to the spotlight thanks to their generality and strong performance in a wide range of language tasks, including translation.Here we show that open-source LLMs perform on par with or better than some state-of-the-art baselines in simultaneous machine translation (SiMT) tasks, zero-shot.We also demonstrate that injection of minimal background information, which is easy with an LLM, brings further performance gains, especially on challenging technical subject-matter.This highlights LLMs' potential for building next generation of massively multilingual, context-aware and terminologically accurate SiMT systems that require no resource-intensive training or fine-tuning.The code is available at https://github.com/RomanKoshkin/toLLMatch.
Roman Koshkin, Katsuhito Sudoh, Satoshi Nakamura 0001
EMNLP2
2024 Subspace Representations for Soft Set Operations and Sentence Similarities
abstract
Yoichi Ishibashi, Sho Yokoi, Katsuhito Sudoh, Satoshi Nakamura. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yoichi Ishibashi, Sho Yokoi, Katsuhito Sudoh, Satoshi Nakamura 0001
NAACL-HLT3
2024 Improving Speech Translation Accuracy and Time Efficiency With Fine-Tuned wav2vec 2.0-Based Speech Segmentation
abstract
Speech translation (ST) automatically converts utterances in a source language into text in another language. Splitting continuous speech into shorter segments, known as speech segmentation, plays an important role in ST. Recent segmentation methods trained to mimic the segmentation of ST corpora have surpassed traditional approaches. Tsiamas et al. [1] proposed a segmentation frame classifier (SFC) based on a pre-trained speech encoder called wav2vec 2.0. Their method, named SHAS, retains 95-98% of the BLEU score for ST corpus segmentation. However, the segments generated by SHAS are very different from ST corpus segmentation and tend to be longer with multiple combined utterances. This is due to SHAS's reliance on length heuristics, i.e., it splits speech into segments of easily translatable length without fully considering the potential for ST improvement by splitting them into even shorter segments. Longer segments often degrade translation quality and ST's time efficiency. In this study, we extended SHAS to improve ST translation accuracy and efficiency by splitting speech into shorter segments that correspond to sentences. We introduced a simple segmentation avlgorithm using the moving average of SFC predictions without relying on length heuristics and explored wav2vec 2.0 fine-tuning for improved speech segmentation prediction. Our experimental results reveal that our speech segmentation method significantly improved the quality and the time efficiency of speech translation compared to SHAS.
Ryo Fukuda, Katsuhito Sudoh, Satoshi Nakamura 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Evaluating the Robustness of Discrete Prompts
abstract
Discrete prompts have been used for finetuning Pre-trained Language Models for diverse NLP tasks.In particular, automatic methods that generate discrete prompts from a small set of training instances have reported superior performance.However, a closer look at the learnt prompts reveals that they contain noisy and counter-intuitive lexical constructs that would not be encountered in manuallywritten prompts.This raises an important yet understudied question regarding the robustness of automatically learnt discrete prompts when used in downstream tasks.To address this question, we conduct a systematic study of the robustness of discrete prompts by applying carefully designed perturbations into an application using AutoPrompt and then measure their performance in two Natural Language Inference (NLI) datasets.Our experimental results show that although the discrete prompt-based method remains relatively robust against perturbations to NLI inputs, they are highly sensitive to other types of perturbations such as shuffling and deletion of prompt tokens.Moreover, they generalize poorly across different NLI datasets.We hope our findings will inspire future work on robust discrete prompt learning.1
Yoichi Ishibashi, Danushka Bollegala, Katsuhito Sudoh, Satoshi Nakamura 0001
EACL3
2023 Average Token Delay: A Latency Metric for Simultaneous Translation
Yasumasa Kano, Katsuhito Sudoh, Satoshi Nakamura 0001
INTERSPEECH2
2023 Reflective action selection based on positive-unlabeled learning and causality detection model
abstract
Task-oriented dialogue systems need to take appropriate actions not only for clear user requests but also for ambiguous and vague ones. In this study, “ambiguous” denotes that although users have potential requests, they failed to clearly define and verbalize their content and conditions which can be associated with system actions. For such ambiguous requests, taking reflective actions is one plausible choice for such systems. In our study, “reflective” denotes taking actions that satisfy user requests before the users themselves clarify their demands. We constructed such a reflective dialogue agent by collecting a corpus that includes pairs of ambiguous user requests and corresponding reflective system actions on sightseeing navigation with a smartphone. Since annotating every possible combination of user requests and system actions is impossible, this study built a corpus where one reflective action is annotated to one ambiguous user request. To train an action selection model on such incomplete training data in which only one action is associated with a request, we applied the positive/unlabeled (PU) learning method, which assumes that only part of the data is labeled with positive examples. In addition, we enhanced the action selection by extracting and distilling knowledge that corresponds to causality from the training data using a causality detection model. The experimental results show that both the PU learning method and the causality detection model improved the performances of the reflective action selection compared to the conventional positive/negative (PN) learning method.
Shohei Tanaka, Koichiro Yoshino, Katsuhito Sudoh, Satoshi Nakamura 0001
Comput. Speech Lang.3
2022 Speech Segmentation Optimization using Segmented Bilingual Speech Corpus for End-to-end Speech Translation
abstract
Speech segmentation, which splits long speech into short segments, is essential for speech translation (ST).Popular VAD tools like WebRTC VAD 1 have generally relied on pause-based segmentation.Unfortunately, pauses in speech do not necessarily match sentence boundaries, and sentences can be connected by a very short pause that is difficult to detect by VAD.In this study, we propose a speech segmentation method using a binary classification model trained using a segmented bilingual speech corpus.We also propose a hybrid method that combines VAD and the above speech segmentation method.Experimental results reveal that the proposed method is more suitable for cascade and end-to-end ST systems than conventional segmentation methods.The hybrid approach further improves the translation performance.
Ryo Fukuda, Katsuhito Sudoh, Satoshi Nakamura 0001
INTERSPEECH2
2022 VTTS: Visual-Text To Speech
abstract
This paper proposes a visual-text to speech (vTTS) method, a method for synthesizing speech directly from visual text (i.e., text as an image). vTTS can use visual features in visual text that should be important for speech synthesis such as emphasis and radicals (components in Chinese characters), but they are not available in conventional TTS using discrete symbols as the input. The proposed vTTS method extracts visual features with a convolutional neural network and then generates acoustic features with a non-autoregressive model in an end-to-end manner. Experimental results show that 1) the vTTS method is capable of generating speech with naturalness comparable to or better than a conventional TTS, 2) it can transfer emphasis and emotion attributes in visual text to speech without additional labels and architectures, and 3) it can synthesize more natural and intelligible speech from unseen and rare characters than conventional TTS.
Yoshifumi Nakano, Takaaki Saeki, Shinnosuke Takamichi, Katsuhito Sudoh, Hiroshi Saruwatari
SLT4
2021 ASR Posterior-Based Loss for Multi-Task End-to-End Speech Translation
Yuka Ko, Katsuhito Sudoh, Sakriani Sakti, Satoshi Nakamura 0001
Interspeech2
2021 Transcribing Paralinguistic Acoustic Cues to Target Language Text in Transformer-Based Speech-to-Text Translation
Hirotaka Tokuyama, Sakriani Sakti, Katsuhito Sudoh, Satoshi Nakamura 0001
Interspeech3
2021 ARTA: Collection and Classification of Ambiguous Requests and Thoughtful Actions
abstract
Human-assisting systems such as dialogue systems must take thoughtful, appropriate actions not only for clear and unambiguous user requests, but also for ambiguous user requests, even if the users themselves are not aware of their potential requirements.To construct such a dialogue agent, we collected a corpus and developed a model that classifies ambiguous user requests into corresponding system actions.In order to collect a high-quality corpus, we asked workers to input antecedent user requests whose pre-defined actions could be regarded as thoughtful.Although multiple actions could be identified as thoughtful for a single user request, annotating all combinations of user requests and system actions is impractical.For this reason, we fully annotated only the test data and left the annotation of the training data incomplete.In order to train the classification model on such training data, we applied the positive/unlabeled (PU) learning method, which assumes that only a part of the data is labeled with positive examples.The experimental results show that the PU learning method achieved better performance than the general positive/negative (PN) learning method to classify thoughtful actions given an ambiguous user request.
Shohei Tanaka, Koichiro Yoshino, Katsuhito Sudoh, Satoshi Nakamura 0001
SIGDIAL3
2021 Towards Tokenization and Part-of-Speech Tagging for Khmer: Data and Discussion
abstract
As a highly analytic language, Khmer has considerable ambiguities in tokenization and part-of-speech (POS) tagging processing. This topic is investigated in this study. Specifically, a 20,000-sentence Khmer corpus with manual tokenization and POS-tagging annotation is released after a series of work over the last 4 years. This is the largest morphologically annotated Khmer dataset as of 2020, when this article was prepared. Based on the annotated data, experiments were conducted to establish a comprehensive benchmark on the automatic processing of tokenization and POS-tagging for Khmer. Specifically, a support vector machine, a conditional random field (CRF) , a long short-term memory (LSTM) -based recurrent neural network, and an integrated LSTM-CRF model have been investigated and discussed. As a primary conclusion, processing at morpheme-level is satisfactory for the provided data. However, it is intrinsically difficult to identify further grammatical constituents of compounds or phrases because of the complex analytic features of the language. Syntactic annotation and automatic parsing for Khmer will be scheduled in the near future.
Hour Kaing, Chenchen Ding, Masao Utiyama, Eiichiro Sumita, Sam Sethserey, Sopheap Seng, Katsuhito Sudoh, Satoshi Nakamura 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.7
2020 Automatic Machine Translation Evaluation using Source Language Inputs and Cross-lingual Language Model
abstract
We propose an automatic evaluation method of machine translation that uses source language sentences regarded as additional pseudo references.The proposed method evaluates a translation hypothesis in a regression model.The model takes the paired source, reference, and hypothesis sentence all together as an input.A pretrained large scale cross-lingual language model encodes the input to sentence-pair vectors, and the model predicts a human evaluation score with those vectors.Our experiments show that our proposed method using Crosslingual Language Model (XLM) trained with a translation language modeling (TLM) objective achieves a higher correlation with human judgments than a baseline method that uses only hypothesis and reference sentences.Additionally, using source sentences in our proposed method is confirmed to improve the evaluation performance.
Kosuke Takahashi, Katsuhito Sudoh, Satoshi Nakamura 0001
ACL2
2020 Incorporating Noisy Length Constraints into Transformer with Length-aware Positional Encodings
abstract
Neural Machine Translation often suffers from an under-translation problem due to its limited modeling of output sequence lengths.In this work, we propose a novel approach to training a Transformer model using length constraints based on length-aware positional encoding (PE).Since length constraints with exact target sentence lengths degrade translation performance, we add random noise within a certain window size to the length constraints in the PE during the training.In the inference step, we predict the output lengths using input sequences and a BERTbased length prediction model.Experimental results in an ASPEC English-to-Japanese translation showed the proposed method produced translations with lengths close to the reference ones and outperformed a vanilla Transformer by 3.22 points in BLEU on short sentences within ten subwords.The average translation results using our length prediction model were also better than another baseline method using input lengths for the length constraints.The proposed noise injection improved robustness for length prediction errors, especially within the window size.
Yui Oka, Katsuki Chousa, Katsuhito Sudoh, Satoshi Nakamura 0001
COLING3
2020 Improving Spoken Language Understanding by Wisdom of Crowds
abstract
Spoken language understanding (SLU), which converts user requests in natural language to machine-interpretable expressions, is becoming an essential task.The lack of training data is an important problem, especially for new system tasks, because existing SLU systems are based on statistical approaches.In this paper, we proposed to use two sources of the "wisdom of crowds," crowdsourcing and knowledge community website, for improving the SLU system.We firstly collected paraphrasing variations for new system tasks through crowdsourcing as seed data, and then augmented them using similar questions from a knowledge community website.We investigated the effects of the proposed data augmentation method in SLU task, even with small seed data.In particular, the proposed architecture augmented more than 120,000 samples to improve SLU accuracies.
Koichiro Yoshino, Kana Ikeuchi, Katsuhito Sudoh, Satoshi Nakamura 0001
COLING3
2020 Multi-Source Neural Machine Translation With Missing Data
abstract
Machine translation is rife with ambiguities in word ordering and word choice, and even with the advent of machine-learning methods that learn to resolve this ambiguity based on statistics from large corpora, mistakes are frequent. Multi-source translation is an approach that attempts to resolve these ambiguities by exploiting multiple inputs (e.g. sentences in three different languages) to increase translation accuracy. These methods are trained on multilingual corpora, which include the multiple source languages and the target language, and then at test time uses information from both source languages while generating the target. While there are many of these multilingual corpora, such as multilingual translations of TED talks or European parliament proceedings, in practice, many multilingual corpora are not complete due to the difficulty to provide translations in all of the relevant languages. Existing studies on multi-source translation did not explicitly handle such situations, and thus are only applicable to complete corpora that have all of the languages of interest, severely limiting their practical applicability. In this article, we examine approaches for multi-source neural machine translation (NMT) that can learn from and translate such incomplete corpora. Specifically, we propose methods to deal with incomplete corpora at both training time and test time. For training time, we examine two methods: (1) a simple method that simply replaces missing source translations with a special NULL symbol, and (2) a data augmentation approach that fills in incomplete parts with source translations created from multi-source NMT. For test-time, we examine methods that use multi-source translation even when only a single source is provided by first translating into an additional auxiliary language using standard NMT, then using multi-source translation on the original source and this generated auxiliary language sentence. Extensive experiments demonstrate that the proposed training-time and test-time methods both significantly improve translation performance.
Yuta Nishimura, Katsuhito Sudoh, Graham Neubig, Satoshi Nakamura 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Learning to Rank for Coordination Detection
Rumeng Li, Hiroyuki Shindo, Katsuhito Sudoh, Masaaki Nagata
CICLing (1)4
2016 Exploring Text Links for Coherent Multi-Document Summarization
abstract
Summarization aims to represent source documents by a shortened passage. Existing methods focus on the extraction of key information, but often neglect coherence. Hence the generated summaries suffer from a lack of readability. To address this problem, we have developed a graph-based method by exploring the links between text to produce coherent summaries. Our approach involves finding a sequence of sentences that best represent the key information in a coherent way. In contrast to the previous methods that focus only on salience, the proposed method addresses both coherence and informativeness based on textual linkages. We conduct experiments on the DUC2004 summarization task data set. A performance comparison reveals that the summaries generated by the proposed system achieve comparable results in terms of the ROUGE metric, and show improvements in readability by human evaluation.
Masaaki Nishino, Tsutomu Hirao, Katsuhito Sudoh, Masaaki Nagata
COLING4
2015 Enhanced Word Embeddings from a Hierarchical Neural Language Model
abstract
This paper proposes a neural language model to capture the interaction of text units of different levels, i.e.., documents, paragraphs, sentences, words in an hierarchical structure. At each paralleled level, the model incorporates Markov property while each higher-level unit hierarchically influences its containing units. Such an architecture enables the learned word embeddings to encode both global and local information. We evaluate the learned word embeddings and experiments demonstrate the effectiveness of our model.
Katsuhito Sudoh, Masaaki Nagata
CIKM2
2015 Empty Category Detection With Joint Context-Label Embeddings
abstract
This paper presents a novel technique for empty category (EC) detection using distributed word representations.A joint model is learned from the labeled data to map both the distributed representations of the contexts of ECs and EC types to a low dimensional space.In the testing phase, the context of possible EC positions will be projected into the same space for empty category detection.Experiments on Chinese Treebank prove the effectiveness of the proposed method.We improve the precision by about 6 points on a subset of Chinese Treebank, which is a new state-ofthe-art performance on CTB.
Katsuhito Sudoh, Masaaki Nagata
HLT-NAACL2
2015 Rating Entities and Aspects Using a Hierarchical Model
Katsuhito Sudoh, Masaaki Nagata
PAKDD (2)2
2015 Summarization Based on Task-Oriented Discourse Parsing
abstract
Previous research demonstrates that discourse relations can help generate high-quality summaries. Existing studies usually adopt existing discourse parsers directly with no modifications, hence cannot take full advantage of discourse parsing. This paper describes a new single document summarization system. In contrast to previous work, we train a discourse parser specially for summarization by using summaries. The training data are dynamically changed during the training phase to enable the parser to grab the text units that are important for summaries. A special tree-based summary extraction algorithm is designed to work with the new parser. The proposed system enables us to combine discourse parsing and summarization in a unified scheme. Experiments on both the RST-DT and DUC2001 datasets show the effectiveness of the proposed method.
Yasuhisa Yoshida, Tsutomu Hirao, Katsuhito Sudoh, Masaaki Nagata
IEEE ACM Trans. Audio Speech Lang. Process.4
2013 Shift-Reduce Word Reordering for Machine Translation
abstract
This paper presents a novel word reordering model that employs a shift-reduce parser for inversion transduction grammars.Our model uses rich syntax parsing features for word reordering and runs in linear time.We apply it to postordering of phrase-based machine translation (PBMT) for Japanese-to-English patent tasks.Our experimental results show that our method achieves a significant improvement of +3.1 BLEU scores against 30.15BLEU scores of the baseline PBMT system.
Katsuhiko Hayashi 0001, Katsuhito Sudoh, Hajime Tsukada, Jun Suzuki 0001, Masaaki Nagata
EMNLP2
2013 Noise-Aware Character Alignment for Bootstrapping Statistical Machine Transliteration from Bilingual Corpora
abstract
This paper proposes a novel noise-aware character alignment method for bootstrapping statistical machine transliteration from automatically extracted phrase pairs.The model is an extension of a Bayesian many-to-many alignment method for distinguishing nontransliteration (noise) parts in phrase pairs.It worked effectively in the experiments of bootstrapping Japanese-to-English statistical machine transliteration in patent domain using patent bilingual corpora.
Katsuhito Sudoh, Shinsuke Mori, Masaaki Nagata
EMNLP1
2013 Two-Stage Pre-ordering for Japanese-to-English Statistical Machine Translation
Sho Hoshino, Yusuke Miyao, Katsuhito Sudoh, Masaaki Nagata
IJCNLP3
2013 Effects of Parsing Errors on Pre-Reordering Performance for Chinese-to-Japanese SMT
Pascual Martínez-Gómez, Yusuke Miyao, Katsuhito Sudoh, Masaaki Nagata
PACLIC4
2013 Syntax-Based Post-Ordering for Efficient Japanese-to-English Translation
abstract
This article proposes a novel reordering method for efficient two-step Japanese-to-English statistical machine translation (SMT) that isolates reordering from SMT and solves it after lexical translation. This reordering problem, called post-ordering , is solved as an SMT problem from Head-Final English (HFE) to English. HFE is syntax-based reordered English that is very successfully used for reordering with English-to-Japanese SMT. The proposed method incorporates its advantage into the reverse direction, Japanese-to-English, and solves the post-ordering problem by accurate syntax-based SMT with target language syntax. Two-step SMT with the proposed post-ordering empirically reduces the decoding time of the accurate but slow syntax-based SMT by its good approximation using intermediate HFE. The proposed method improves the decoding speed of syntax-based SMT decoding by about six times with comparable translation accuracy in Japanese-to-English patent translation experiments.
Katsuhito Sudoh, Xianchao Wu, Kevin Duh, Hajime Tsukada, Masaaki Nagata
ACM Trans. Asian Lang. Inf. Process.1
2012 Learning to Translate with Multiple Objectives
Kevin Duh, Katsuhito Sudoh, Xianchao Wu, Hajime Tsukada, Masaaki Nagata
ACL (1)2
2012 HPSG-Based Preprocessing for English-to-Japanese Translation
abstract
Japanese sentences have completely different word orders from corresponding English sentences. Typical phrase-based statistical machine translation (SMT) systems such as Moses search for the best word permutation within a given distance limit (distortion limit). For English-to-Japanese translation, we need a large distance limit to obtain acceptable translations, and the number of translation candidates is extremely large. Therefore, SMT systems often fail to find acceptable translations within a limited time. To solve this problem, some researchers use rule-based preprocessing approaches, which reorder English words just like Japanese by using dozens of rules. Our idea is based on the following two observations: (1) Japanese is a typical head-final language, and (2) we can detect heads of English sentences by a head-driven phrase structure grammar (HPSG) parser. The main contributions of this article are twofold: First, we demonstrate how off-the-shelf, state-of-the-art HPSG parser enables us to write the reordering rules in an abstract level and can easily improve the quality of English-to-Japanese translation. Second, we also show that syntactic heads achieve better results than semantic heads. The proposed method outperforms the best system of NTCIR-7 PATMT EJ task.
Hideki Isozaki, Katsuhito Sudoh, Hajime Tsukada, Kevin Duh
ACM Trans. Asian Lang. Inf. Process.2
2011 Generalized Minimum Bayes Risk System Combination
Kevin Duh, Katsuhito Sudoh, Xianchao Wu, Hajime Tsukada, Masaaki Nagata
IJCNLP2
2011 Extracting Pre-ordering Rules from Predicate-Argument Structures
Xianchao Wu, Katsuhito Sudoh, Kevin Duh, Hajime Tsukada, Masaaki Nagata
IJCNLP2
2011 Alignment Inference and Bayesian Adaptation for Machine Translation
Kevin Duh, Katsuhito Sudoh, Tomoharu Iwata, Hajime Tsukada
MTSummit2
2011 Post-ordering in Statistical Machine Translation
Katsuhito Sudoh, Xianchao Wu, Kevin Duh, Hajime Tsukada, Masaaki Nagata
MTSummit1
2011 Extracting Pre-ordering Rules from Chunk-based Dependency Trees for Japanese-to-English Translation
Xianchao Wu, Katsuhito Sudoh, Kevin Duh, Hajime Tsukada, Masaaki Nagata
MTSummit2
2010 Hierarchical Phrase-based Machine Translation with Word-based Reordering Model
Katsuhiko Hayashi 0001, Hajime Tsukada, Katsuhito Sudoh, Kevin Duh, Seiichi Yamamoto
COLING3
2010 Automatic Evaluation of Translation Quality for Distant Language Pairs
Hideki Isozaki, Tsutomu Hirao, Kevin Duh, Katsuhito Sudoh, Hajime Tsukada
EMNLP4
2006 Incorporating Speech Recognition Confidence into Discriminative Named Entity Recognition of Speech Data
abstract
This paper proposes a named entity recognition (NER) method for speech recognition results that uses confidence on automatic speech recognition (ASR) as a feature.The ASR confidence feature indicates whether each word has been correctly recognized.The NER model is trained using ASR results with named entity (NE) labels as well as the corresponding transcriptions with NE labels.In experiments using support vector machines (SVMs) and speech data from Japanese newspaper articles, the proposed method outperformed a simple application of textbased NER to ASR results in NER Fmeasure by improving precision.These results show that the proposed method is effective in NER for noisy inputs.
Katsuhito Sudoh, Hajime Tsukada, Hideki Isozaki
ACL1
2006 Discriminative named entity recognition of speech data using speech recognition confidence
Katsuhito Sudoh, Hajime Tsukada, Hideki Isozaki
INTERSPEECH1
2006 Incorporating discourse features into confidence scoring of intention recognition results in spoken dialogue systems
Ryuichiro Higashinaka, Katsuhito Sudoh, Mikio Nakano
Speech Commun.2
2005 Incorporating Discourse Features into Confidence Scoring of Intention Recognition Results in Spoken Dialogue Systems
abstract
The paper proposes a method for the confidence scoring of intention recognition results in spoken dialogue systems. To achieve tasks, a spoken dialogue system has to recognize user intentions. However, because of speech recognition errors and ambiguity in user utterances, it sometimes has difficulty recognizing them correctly. Confidence scoring allows errors to be detected in intention recognition results and has proved useful for dialogue management. Conventional methods use the features obtained from speech recognition results for single utterances for confidence scoring. However, this may be insufficient since the intention recognition result is a result of discourse processing. We propose incorporating discourse features for a more accurate confidence scoring of intention recognition results. Experimental results show that incorporating discourse features significantly improves the confidence scoring.
Ryuichiro Higashinaka, Katsuhito Sudoh, Mikio Nakano
ICASSP (1)2
2005 Tightly integrated spoken language understanding using word-to-concept translation
Katsuhito Sudoh, Hajime Tsukada
INTERSPEECH1
2005 Post-dialogue confidence scoring for unsupervised statistical language model training
Katsuhito Sudoh, Mikio Nakano
Speech Commun.1