EDBT 2026 Demo / reviewers in the wild / expert
Katsuhito Sudoh
dblp:66/2380
· DBLP profile ↗
46ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-2122-9846ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 6 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
Anum Afzal, Yuki Saito 0001, Hiroya Takamura, Katsuhito Sudoh, Shinnosuke Takamichi, Graham Neubig, Florian Matthes, Tatsuya Ishigaki |
LREC | 4 |
| 2025 | Measuring Time Delay Tolerance in Third-Person Live Commentary for Super Smash Bros. UltimateabstractThis study proposes a methodology for measuring the acceptable delay tolerance for third-person game commentary. Third-person game commentary refers to commentary delivered by someone other than the player, with the role of helping viewers better understand the game and enhancing the viewing experience. With the recent advancement of AI, there has been increasing interest in automating such commentary using video understanding and audio generation. However, automating this process using video understanding and audio generation introduces delays, potentially affecting the naturalness of the commentary. In this context, since the extent to which such delays are acceptable to viewers remains unclear, we address this issue. The tolerance is modeled using an unnormalized Gaussian function. Through experiments on Super Smash Bros. Ultimate with 727 participants, we found that the average acceptable delay for this game is 3.71 seconds, with variations depending on different viewer attributes and gameplay contexts. Ryosuke Matsushita, Ryosuke Sakai, Koki Fukuda, Shinnosuke Takamichi, Kota Iura, Yuki Saito 0001, Graham Neubig, Katsuhito Sudoh, Hiroya Takamura, Tatsuya Ishigaki |
CoG | 8 |
| 2024 | Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and SegmentationabstractHuman evaluation is a critical component in machine translation system development and has received much attention in text translation research. However, little prior work exists on the topic of human evaluation for speech translation, which adds additional challenges such as noisy data and segmentation mismatches. We take the first steps to fill this gap by conducting a comprehensive human evaluation of the results of several shared tasks from the last International Workshop on Spoken Language Translation (IWSLT 2023). We propose an effective evaluation strategy based on automatic resegmentation and direct assessment with segment context. Our analysis revealed that: 1) the proposed evaluation strategy is robust and scores well-correlated with other types of human judgements; 2) automatic metrics are usually, but not always, well-correlated with direct assessment scores; and 3) COMET as a slightly stronger automatic metric than chrF, despite the segmentation noise introduced by the resegmentation step systems. We release the collected human-annotated data in order to encourage further investigation. Matthias Sperber, Ondrej Bojar, Barry Haddow, Dávid Javorský, Xutai Ma, Matteo Negri, Jan Niehues, Peter Polak, Elizabeth Salesky, Katsuhito Sudoh, Marco Turchi |
LREC/COLING | 10 |
| 2024 | NAIST-SIC-Aligned: An Aligned English-Japanese Simultaneous Interpretation CorpusabstractIt remains a question that how simultaneous interpretation (SI) data affects simultaneous machine translation (SiMT). Research has been limited due to the lack of a large-scale training corpus. In this work, we aim to fill in the gap by introducing NAIST-SIC-Aligned, which is an automatically-aligned parallel English-Japanese SI dataset. Starting with a non-aligned corpus NAIST-SIC, we propose a two-stage alignment approach to make the corpus parallel and thus suitable for model training. The first stage is coarse alignment where we perform a many-to-many mapping between source and target sentences, and the second stage is fine-grained alignment where we perform intra- and inter-sentence filtering to improve the quality of aligned pairs. To ensure the quality of the corpus, each step has been validated either quantitatively or qualitatively. This is the first open-sourced large-scale parallel SI dataset in the literature. We also manually curated a small test set for evaluation purposes. Our results show that models trained with SI data lead to significant improvement in translation quality and latency over baselines. We hope our work advances research on SI corpora construction and SiMT. Our data will be released upon the paper’s acceptance. Jinming Zhao, Katsuhito Sudoh, Satoshi Nakamura 0001, Yuka Ko, Kosuke Doi, Ryo Fukuda |
LREC/COLING | 2 |
| 2024 | LLMs Are Zero-Shot Context-Aware Simultaneous TranslatorsabstractThe advent of transformers has fueled progress in machine translation.More recently large language models (LLMs) have come to the spotlight thanks to their generality and strong performance in a wide range of language tasks, including translation.Here we show that open-source LLMs perform on par with or better than some state-of-the-art baselines in simultaneous machine translation (SiMT) tasks, zero-shot.We also demonstrate that injection of minimal background information, which is easy with an LLM, brings further performance gains, especially on challenging technical subject-matter.This highlights LLMs' potential for building next generation of massively multilingual, context-aware and terminologically accurate SiMT systems that require no resource-intensive training or fine-tuning.The code is available at https://github.com/RomanKoshkin/toLLMatch. Roman Koshkin, Katsuhito Sudoh, Satoshi Nakamura 0001 |
EMNLP | 2 |
| 2024 | Subspace Representations for Soft Set Operations and Sentence SimilaritiesabstractYoichi Ishibashi, Sho Yokoi, Katsuhito Sudoh, Satoshi Nakamura. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yoichi Ishibashi, Sho Yokoi, Katsuhito Sudoh, Satoshi Nakamura 0001 |
NAACL-HLT | 3 |
| 2024 | Improving Speech Translation Accuracy and Time Efficiency With Fine-Tuned wav2vec 2.0-Based Speech SegmentationabstractSpeech translation (ST) automatically converts utterances in a source language into text in another language. Splitting continuous speech into shorter segments, known as speech segmentation, plays an important role in ST. Recent segmentation methods trained to mimic the segmentation of ST corpora have surpassed traditional approaches. Tsiamas et al. [1] proposed a segmentation frame classifier (SFC) based on a pre-trained speech encoder called wav2vec 2.0. Their method, named SHAS, retains 95-98% of the BLEU score for ST corpus segmentation. However, the segments generated by SHAS are very different from ST corpus segmentation and tend to be longer with multiple combined utterances. This is due to SHAS's reliance on length heuristics, i.e., it splits speech into segments of easily translatable length without fully considering the potential for ST improvement by splitting them into even shorter segments. Longer segments often degrade translation quality and ST's time efficiency. In this study, we extended SHAS to improve ST translation accuracy and efficiency by splitting speech into shorter segments that correspond to sentences. We introduced a simple segmentation avlgorithm using the moving average of SFC predictions without relying on length heuristics and explored wav2vec 2.0 fine-tuning for improved speech segmentation prediction. Our experimental results reveal that our speech segmentation method significantly improved the quality and the time efficiency of speech translation compared to SHAS. Ryo Fukuda, Katsuhito Sudoh, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Evaluating the Robustness of Discrete PromptsabstractDiscrete prompts have been used for finetuning Pre-trained Language Models for diverse NLP tasks.In particular, automatic methods that generate discrete prompts from a small set of training instances have reported superior performance.However, a closer look at the learnt prompts reveals that they contain noisy and counter-intuitive lexical constructs that would not be encountered in manuallywritten prompts.This raises an important yet understudied question regarding the robustness of automatically learnt discrete prompts when used in downstream tasks.To address this question, we conduct a systematic study of the robustness of discrete prompts by applying carefully designed perturbations into an application using AutoPrompt and then measure their performance in two Natural Language Inference (NLI) datasets.Our experimental results show that although the discrete prompt-based method remains relatively robust against perturbations to NLI inputs, they are highly sensitive to other types of perturbations such as shuffling and deletion of prompt tokens.Moreover, they generalize poorly across different NLI datasets.We hope our findings will inspire future work on robust discrete prompt learning.1 Yoichi Ishibashi, Danushka Bollegala, Katsuhito Sudoh, Satoshi Nakamura 0001 |
EACL | 3 |
| 2023 | Average Token Delay: A Latency Metric for Simultaneous Translation
Yasumasa Kano, Katsuhito Sudoh, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2023 | Reflective action selection based on positive-unlabeled learning and causality detection modelabstractTask-oriented dialogue systems need to take appropriate actions not only for clear user requests but also for ambiguous and vague ones. In this study, “ambiguous” denotes that although users have potential requests, they failed to clearly define and verbalize their content and conditions which can be associated with system actions. For such ambiguous requests, taking reflective actions is one plausible choice for such systems. In our study, “reflective” denotes taking actions that satisfy user requests before the users themselves clarify their demands. We constructed such a reflective dialogue agent by collecting a corpus that includes pairs of ambiguous user requests and corresponding reflective system actions on sightseeing navigation with a smartphone. Since annotating every possible combination of user requests and system actions is impossible, this study built a corpus where one reflective action is annotated to one ambiguous user request. To train an action selection model on such incomplete training data in which only one action is associated with a request, we applied the positive/unlabeled (PU) learning method, which assumes that only part of the data is labeled with positive examples. In addition, we enhanced the action selection by extracting and distilling knowledge that corresponds to causality from the training data using a causality detection model. The experimental results show that both the PU learning method and the causality detection model improved the performances of the reflective action selection compared to the conventional positive/negative (PN) learning method. Shohei Tanaka, Koichiro Yoshino, Katsuhito Sudoh, Satoshi Nakamura 0001 |
Comput. Speech Lang. | 3 |
| 2022 | Speech Segmentation Optimization using Segmented Bilingual Speech Corpus for End-to-end Speech TranslationabstractSpeech segmentation, which splits long speech into short segments, is essential for speech translation (ST).Popular VAD tools like WebRTC VAD 1 have generally relied on pause-based segmentation.Unfortunately, pauses in speech do not necessarily match sentence boundaries, and sentences can be connected by a very short pause that is difficult to detect by VAD.In this study, we propose a speech segmentation method using a binary classification model trained using a segmented bilingual speech corpus.We also propose a hybrid method that combines VAD and the above speech segmentation method.Experimental results reveal that the proposed method is more suitable for cascade and end-to-end ST systems than conventional segmentation methods.The hybrid approach further improves the translation performance. Ryo Fukuda, Katsuhito Sudoh, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2022 | VTTS: Visual-Text To SpeechabstractThis paper proposes a visual-text to speech (vTTS) method, a method for synthesizing speech directly from visual text (i.e., text as an image). vTTS can use visual features in visual text that should be important for speech synthesis such as emphasis and radicals (components in Chinese characters), but they are not available in conventional TTS using discrete symbols as the input. The proposed vTTS method extracts visual features with a convolutional neural network and then generates acoustic features with a non-autoregressive model in an end-to-end manner. Experimental results show that 1) the vTTS method is capable of generating speech with naturalness comparable to or better than a conventional TTS, 2) it can transfer emphasis and emotion attributes in visual text to speech without additional labels and architectures, and 3) it can synthesize more natural and intelligible speech from unseen and rare characters than conventional TTS. Yoshifumi Nakano, Takaaki Saeki, Shinnosuke Takamichi, Katsuhito Sudoh, Hiroshi Saruwatari |
SLT | 4 |
| 2021 | ASR Posterior-Based Loss for Multi-Task End-to-End Speech Translation
Yuka Ko, Katsuhito Sudoh, Sakriani Sakti, Satoshi Nakamura 0001 |
Interspeech | 2 |
| 2021 | Transcribing Paralinguistic Acoustic Cues to Target Language Text in Transformer-Based Speech-to-Text Translation
Hirotaka Tokuyama, Sakriani Sakti, Katsuhito Sudoh, Satoshi Nakamura 0001 |
Interspeech | 3 |
| 2021 | ARTA: Collection and Classification of Ambiguous Requests and Thoughtful ActionsabstractHuman-assisting systems such as dialogue systems must take thoughtful, appropriate actions not only for clear and unambiguous user requests, but also for ambiguous user requests, even if the users themselves are not aware of their potential requirements.To construct such a dialogue agent, we collected a corpus and developed a model that classifies ambiguous user requests into corresponding system actions.In order to collect a high-quality corpus, we asked workers to input antecedent user requests whose pre-defined actions could be regarded as thoughtful.Although multiple actions could be identified as thoughtful for a single user request, annotating all combinations of user requests and system actions is impractical.For this reason, we fully annotated only the test data and left the annotation of the training data incomplete.In order to train the classification model on such training data, we applied the positive/unlabeled (PU) learning method, which assumes that only a part of the data is labeled with positive examples.The experimental results show that the PU learning method achieved better performance than the general positive/negative (PN) learning method to classify thoughtful actions given an ambiguous user request. Shohei Tanaka, Koichiro Yoshino, Katsuhito Sudoh, Satoshi Nakamura 0001 |
SIGDIAL | 3 |
| 2021 | Towards Tokenization and Part-of-Speech Tagging for Khmer: Data and DiscussionabstractAs a highly analytic language, Khmer has considerable ambiguities in tokenization and part-of-speech (POS) tagging processing. This topic is investigated in this study. Specifically, a 20,000-sentence Khmer corpus with manual tokenization and POS-tagging annotation is released after a series of work over the last 4 years. This is the largest morphologically annotated Khmer dataset as of 2020, when this article was prepared. Based on the annotated data, experiments were conducted to establish a comprehensive benchmark on the automatic processing of tokenization and POS-tagging for Khmer. Specifically, a support vector machine, a conditional random field (CRF) , a long short-term memory (LSTM) -based recurrent neural network, and an integrated LSTM-CRF model have been investigated and discussed. As a primary conclusion, processing at morpheme-level is satisfactory for the provided data. However, it is intrinsically difficult to identify further grammatical constituents of compounds or phrases because of the complex analytic features of the language. Syntactic annotation and automatic parsing for Khmer will be scheduled in the near future. Hour Kaing, Chenchen Ding, Masao Utiyama, Eiichiro Sumita, Sam Sethserey, Sopheap Seng, Katsuhito Sudoh, Satoshi Nakamura 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 7 |
| 2020 | Automatic Machine Translation Evaluation using Source Language Inputs and Cross-lingual Language ModelabstractWe propose an automatic evaluation method of machine translation that uses source language sentences regarded as additional pseudo references.The proposed method evaluates a translation hypothesis in a regression model.The model takes the paired source, reference, and hypothesis sentence all together as an input.A pretrained large scale cross-lingual language model encodes the input to sentence-pair vectors, and the model predicts a human evaluation score with those vectors.Our experiments show that our proposed method using Crosslingual Language Model (XLM) trained with a translation language modeling (TLM) objective achieves a higher correlation with human judgments than a baseline method that uses only hypothesis and reference sentences.Additionally, using source sentences in our proposed method is confirmed to improve the evaluation performance. Kosuke Takahashi, Katsuhito Sudoh, Satoshi Nakamura 0001 |
ACL | 2 |
| 2020 | Incorporating Noisy Length Constraints into Transformer with Length-aware Positional EncodingsabstractNeural Machine Translation often suffers from an under-translation problem due to its limited modeling of output sequence lengths.In this work, we propose a novel approach to training a Transformer model using length constraints based on length-aware positional encoding (PE).Since length constraints with exact target sentence lengths degrade translation performance, we add random noise within a certain window size to the length constraints in the PE during the training.In the inference step, we predict the output lengths using input sequences and a BERTbased length prediction model.Experimental results in an ASPEC English-to-Japanese translation showed the proposed method produced translations with lengths close to the reference ones and outperformed a vanilla Transformer by 3.22 points in BLEU on short sentences within ten subwords.The average translation results using our length prediction model were also better than another baseline method using input lengths for the length constraints.The proposed noise injection improved robustness for length prediction errors, especially within the window size. Yui Oka, Katsuki Chousa, Katsuhito Sudoh, Satoshi Nakamura 0001 |
COLING | 3 |
| 2020 | Improving Spoken Language Understanding by Wisdom of CrowdsabstractSpoken language understanding (SLU), which converts user requests in natural language to machine-interpretable expressions, is becoming an essential task.The lack of training data is an important problem, especially for new system tasks, because existing SLU systems are based on statistical approaches.In this paper, we proposed to use two sources of the "wisdom of crowds," crowdsourcing and knowledge community website, for improving the SLU system.We firstly collected paraphrasing variations for new system tasks through crowdsourcing as seed data, and then augmented them using similar questions from a knowledge community website.We investigated the effects of the proposed data augmentation method in SLU task, even with small seed data.In particular, the proposed architecture augmented more than 120,000 samples to improve SLU accuracies. Koichiro Yoshino, Kana Ikeuchi, Katsuhito Sudoh, Satoshi Nakamura 0001 |
COLING | 3 |
| 2020 | Multi-Source Neural Machine Translation With Missing DataabstractMachine translation is rife with ambiguities in word ordering and word choice, and even with the advent of machine-learning methods that learn to resolve this ambiguity based on statistics from large corpora, mistakes are frequent. Multi-source translation is an approach that attempts to resolve these ambiguities by exploiting multiple inputs (e.g. sentences in three different languages) to increase translation accuracy. These methods are trained on multilingual corpora, which include the multiple source languages and the target language, and then at test time uses information from both source languages while generating the target. While there are many of these multilingual corpora, such as multilingual translations of TED talks or European parliament proceedings, in practice, many multilingual corpora are not complete due to the difficulty to provide translations in all of the relevant languages. Existing studies on multi-source translation did not explicitly handle such situations, and thus are only applicable to complete corpora that have all of the languages of interest, severely limiting their practical applicability. In this article, we examine approaches for multi-source neural machine translation (NMT) that can learn from and translate such incomplete corpora. Specifically, we propose methods to deal with incomplete corpora at both training time and test time. For training time, we examine two methods: (1) a simple method that simply replaces missing source translations with a special NULL symbol, and (2) a data augmentation approach that fills in incomplete parts with source translations created from multi-source NMT. For test-time, we examine methods that use multi-source translation even when only a single source is provided by first translating into an additional auxiliary language using standard NMT, then using multi-source translation on the original source and this generated auxiliary language sentence. Extensive experiments demonstrate that the proposed training-time and test-time methods both significantly improve translation performance. Yuta Nishimura, Katsuhito Sudoh, Graham Neubig, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Learning to Rank for Coordination Detection
Rumeng Li, Hiroyuki Shindo, Katsuhito Sudoh, Masaaki Nagata |
CICLing (1) | 4 |
| 2016 | Exploring Text Links for Coherent Multi-Document SummarizationabstractSummarization aims to represent source documents by a shortened passage. Existing methods focus on the extraction of key information, but often neglect coherence. Hence the generated summaries suffer from a lack of readability. To address this problem, we have developed a graph-based method by exploring the links between text to produce coherent summaries. Our approach involves finding a sequence of sentences that best represent the key information in a coherent way. In contrast to the previous methods that focus only on salience, the proposed method addresses both coherence and informativeness based on textual linkages. We conduct experiments on the DUC2004 summarization task data set. A performance comparison reveals that the summaries generated by the proposed system achieve comparable results in terms of the ROUGE metric, and show improvements in readability by human evaluation. Masaaki Nishino, Tsutomu Hirao, Katsuhito Sudoh, Masaaki Nagata |
COLING | 4 |
| 2015 | Enhanced Word Embeddings from a Hierarchical Neural Language ModelabstractThis paper proposes a neural language model to capture the interaction of text units of different levels, i.e.., documents, paragraphs, sentences, words in an hierarchical structure. At each paralleled level, the model incorporates Markov property while each higher-level unit hierarchically influences its containing units. Such an architecture enables the learned word embeddings to encode both global and local information. We evaluate the learned word embeddings and experiments demonstrate the effectiveness of our model. Katsuhito Sudoh, Masaaki Nagata |
CIKM | 2 |
| 2015 | Empty Category Detection With Joint Context-Label EmbeddingsabstractThis paper presents a novel technique for empty category (EC) detection using distributed word representations.A joint model is learned from the labeled data to map both the distributed representations of the contexts of ECs and EC types to a low dimensional space.In the testing phase, the context of possible EC positions will be projected into the same space for empty category detection.Experiments on Chinese Treebank prove the effectiveness of the proposed method.We improve the precision by about 6 points on a subset of Chinese Treebank, which is a new state-ofthe-art performance on CTB. Katsuhito Sudoh, Masaaki Nagata |
HLT-NAACL | 2 |
| 2015 | Rating Entities and Aspects Using a Hierarchical Model
Katsuhito Sudoh, Masaaki Nagata |
PAKDD (2) | 2 |
| 2015 | Summarization Based on Task-Oriented Discourse ParsingabstractPrevious research demonstrates that discourse relations can help generate high-quality summaries. Existing studies usually adopt existing discourse parsers directly with no modifications, hence cannot take full advantage of discourse parsing. This paper describes a new single document summarization system. In contrast to previous work, we train a discourse parser specially for summarization by using summaries. The training data are dynamically changed during the training phase to enable the parser to grab the text units that are important for summaries. A special tree-based summary extraction algorithm is designed to work with the new parser. The proposed system enables us to combine discourse parsing and summarization in a unified scheme. Experiments on both the RST-DT and DUC2001 datasets show the effectiveness of the proposed method. Yasuhisa Yoshida, Tsutomu Hirao, Katsuhito Sudoh, Masaaki Nagata |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2013 | Shift-Reduce Word Reordering for Machine TranslationabstractThis paper presents a novel word reordering model that employs a shift-reduce parser for inversion transduction grammars.Our model uses rich syntax parsing features for word reordering and runs in linear time.We apply it to postordering of phrase-based machine translation (PBMT) for Japanese-to-English patent tasks.Our experimental results show that our method achieves a significant improvement of +3.1 BLEU scores against 30.15BLEU scores of the baseline PBMT system. Katsuhiko Hayashi 0001, Katsuhito Sudoh, Hajime Tsukada, Jun Suzuki 0001, Masaaki Nagata |
EMNLP | 2 |
| 2013 | Noise-Aware Character Alignment for Bootstrapping Statistical Machine Transliteration from Bilingual CorporaabstractThis paper proposes a novel noise-aware character alignment method for bootstrapping statistical machine transliteration from automatically extracted phrase pairs.The model is an extension of a Bayesian many-to-many alignment method for distinguishing nontransliteration (noise) parts in phrase pairs.It worked effectively in the experiments of bootstrapping Japanese-to-English statistical machine transliteration in patent domain using patent bilingual corpora. Katsuhito Sudoh, Shinsuke Mori, Masaaki Nagata |
EMNLP | 1 |
| 2013 | Two-Stage Pre-ordering for Japanese-to-English Statistical Machine Translation
Sho Hoshino, Yusuke Miyao, Katsuhito Sudoh, Masaaki Nagata |
IJCNLP | 3 |
| 2013 | Effects of Parsing Errors on Pre-Reordering Performance for Chinese-to-Japanese SMT
Pascual Martínez-Gómez, Yusuke Miyao, Katsuhito Sudoh, Masaaki Nagata |
PACLIC | 4 |
| 2013 | Syntax-Based Post-Ordering for Efficient Japanese-to-English TranslationabstractThis article proposes a novel reordering method for efficient two-step Japanese-to-English statistical machine translation (SMT) that isolates reordering from SMT and solves it after lexical translation. This reordering problem, called post-ordering , is solved as an SMT problem from Head-Final English (HFE) to English. HFE is syntax-based reordered English that is very successfully used for reordering with English-to-Japanese SMT. The proposed method incorporates its advantage into the reverse direction, Japanese-to-English, and solves the post-ordering problem by accurate syntax-based SMT with target language syntax. Two-step SMT with the proposed post-ordering empirically reduces the decoding time of the accurate but slow syntax-based SMT by its good approximation using intermediate HFE. The proposed method improves the decoding speed of syntax-based SMT decoding by about six times with comparable translation accuracy in Japanese-to-English patent translation experiments. Katsuhito Sudoh, Xianchao Wu, Kevin Duh, Hajime Tsukada, Masaaki Nagata |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2012 | Learning to Translate with Multiple Objectives
Kevin Duh, Katsuhito Sudoh, Xianchao Wu, Hajime Tsukada, Masaaki Nagata |
ACL (1) | 2 |
| 2012 | HPSG-Based Preprocessing for English-to-Japanese TranslationabstractJapanese sentences have completely different word orders from corresponding English sentences. Typical phrase-based statistical machine translation (SMT) systems such as Moses search for the best word permutation within a given distance limit (distortion limit). For English-to-Japanese translation, we need a large distance limit to obtain acceptable translations, and the number of translation candidates is extremely large. Therefore, SMT systems often fail to find acceptable translations within a limited time. To solve this problem, some researchers use rule-based preprocessing approaches, which reorder English words just like Japanese by using dozens of rules. Our idea is based on the following two observations: (1) Japanese is a typical head-final language, and (2) we can detect heads of English sentences by a head-driven phrase structure grammar (HPSG) parser. The main contributions of this article are twofold: First, we demonstrate how off-the-shelf, state-of-the-art HPSG parser enables us to write the reordering rules in an abstract level and can easily improve the quality of English-to-Japanese translation. Second, we also show that syntactic heads achieve better results than semantic heads. The proposed method outperforms the best system of NTCIR-7 PATMT EJ task. Hideki Isozaki, Katsuhito Sudoh, Hajime Tsukada, Kevin Duh |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2011 | Generalized Minimum Bayes Risk System Combination
Kevin Duh, Katsuhito Sudoh, Xianchao Wu, Hajime Tsukada, Masaaki Nagata |
IJCNLP | 2 |
| 2011 | Extracting Pre-ordering Rules from Predicate-Argument Structures
Xianchao Wu, Katsuhito Sudoh, Kevin Duh, Hajime Tsukada, Masaaki Nagata |
IJCNLP | 2 |
| 2011 | Alignment Inference and Bayesian Adaptation for Machine Translation
Kevin Duh, Katsuhito Sudoh, Tomoharu Iwata, Hajime Tsukada |
MTSummit | 2 |
| 2011 | Post-ordering in Statistical Machine Translation
Katsuhito Sudoh, Xianchao Wu, Kevin Duh, Hajime Tsukada, Masaaki Nagata |
MTSummit | 1 |
| 2011 | Extracting Pre-ordering Rules from Chunk-based Dependency Trees for Japanese-to-English Translation
Xianchao Wu, Katsuhito Sudoh, Kevin Duh, Hajime Tsukada, Masaaki Nagata |
MTSummit | 2 |
| 2010 | Hierarchical Phrase-based Machine Translation with Word-based Reordering Model
Katsuhiko Hayashi 0001, Hajime Tsukada, Katsuhito Sudoh, Kevin Duh, Seiichi Yamamoto |
COLING | 3 |
| 2010 | Automatic Evaluation of Translation Quality for Distant Language Pairs
Hideki Isozaki, Tsutomu Hirao, Kevin Duh, Katsuhito Sudoh, Hajime Tsukada |
EMNLP | 4 |
| 2006 | Incorporating Speech Recognition Confidence into Discriminative Named Entity Recognition of Speech DataabstractThis paper proposes a named entity recognition (NER) method for speech recognition results that uses confidence on automatic speech recognition (ASR) as a feature.The ASR confidence feature indicates whether each word has been correctly recognized.The NER model is trained using ASR results with named entity (NE) labels as well as the corresponding transcriptions with NE labels.In experiments using support vector machines (SVMs) and speech data from Japanese newspaper articles, the proposed method outperformed a simple application of textbased NER to ASR results in NER Fmeasure by improving precision.These results show that the proposed method is effective in NER for noisy inputs. Katsuhito Sudoh, Hajime Tsukada, Hideki Isozaki |
ACL | 1 |
| 2006 | Discriminative named entity recognition of speech data using speech recognition confidence
Katsuhito Sudoh, Hajime Tsukada, Hideki Isozaki |
INTERSPEECH | 1 |
| 2006 | Incorporating discourse features into confidence scoring of intention recognition results in spoken dialogue systems
Ryuichiro Higashinaka, Katsuhito Sudoh, Mikio Nakano |
Speech Commun. | 2 |
| 2005 | Incorporating Discourse Features into Confidence Scoring of Intention Recognition Results in Spoken Dialogue SystemsabstractThe paper proposes a method for the confidence scoring of intention recognition results in spoken dialogue systems. To achieve tasks, a spoken dialogue system has to recognize user intentions. However, because of speech recognition errors and ambiguity in user utterances, it sometimes has difficulty recognizing them correctly. Confidence scoring allows errors to be detected in intention recognition results and has proved useful for dialogue management. Conventional methods use the features obtained from speech recognition results for single utterances for confidence scoring. However, this may be insufficient since the intention recognition result is a result of discourse processing. We propose incorporating discourse features for a more accurate confidence scoring of intention recognition results. Experimental results show that incorporating discourse features significantly improves the confidence scoring. Ryuichiro Higashinaka, Katsuhito Sudoh, Mikio Nakano |
ICASSP (1) | 2 |
| 2005 | Tightly integrated spoken language understanding using word-to-concept translation
Katsuhito Sudoh, Hajime Tsukada |
INTERSPEECH | 1 |
| 2005 | Post-dialogue confidence scoring for unsupervised statistical language model training
Katsuhito Sudoh, Mikio Nakano |
Speech Commun. | 1 |