VLDB 2026 Research / reviewers in the wild / expert
Andrew M. Finch
dblp:02/2354
· DBLP profile ↗
46ranked-venue papers
14as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 13 first-authorGraphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-authorHuman-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Machine translation · 65% Transfer learning and domain adaptation · 9% Language models and text generation · 8% | |
| Human-computer interaction and pervasive computing
1 paper |
Interaction techniques and input · 100% |
Topics — the 16 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
statistical machine translation |
0.4 | 2 | 2015 | Leave-one-out Word Alignment without Garbage Collector Effects · EMNLP 2015 Refining Word Segmentation Using a Manually Aligned Corpus for Statistical Machine Translation · EMNLP 2014 |
Natural language and speech › Machine translation › statistical machine translation
word alignment |
0.4 | 2 | 2015 | Leave-one-out Word Alignment without Garbage Collector Effects · EMNLP 2015 Refining Word Segmentation Using a Manually Aligned Corpus for Statistical Machine Translation · EMNLP 2014 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.3 | 1 | 2018 | Sentence Selection and Weighting for Neural Machine Translation Domain Adaptation · IEEE ACM Trans. Audio Speech Lang. Process. 2018 |
Natural language and speech › Machine translation
neural machine translation |
0.3 | 1 | 2018 | Sentence Selection and Weighting for Neural Machine Translation Domain Adaptation · IEEE ACM Trans. Audio Speech Lang. Process. 2018 |
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation |
0.3 | 1 | 2017 | Translation Quality Estimation Using Only Bilingual Corpora · IEEE ACM Trans. Audio Speech Lang. Process. 2017 |
Natural language and speech › Machine translation › machine translation evaluation › translation quality estimation
word-level quality estimation |
0.3 | 1 | 2017 | Translation Quality Estimation Using Only Bilingual Corpora · IEEE ACM Trans. Audio Speech Lang. Process. 2017 |
Natural language and speech › Speech recognition and synthesis › pronunciation modeling
grapheme-to-phoneme conversion |
0.2 | 1 | 2016 | Agreement on Target-Bidirectional LSTMs for Sequence-to-Sequence Learning · AAAI 2016 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence-to-sequence learning |
0.2 | 1 | 2016 | Agreement on Target-Bidirectional LSTMs for Sequence-to-Sequence Learning · AAAI 2016 |
Natural language and speech › Machine translation
transliteration |
0.2 | 1 | 2016 | Agreement on Target-Bidirectional LSTMs for Sequence-to-Sequence Learning · AAAI 2016 |
Natural language and speech › Machine translation › statistical machine translation › word alignment
unsupervised word alignment |
0.2 | 1 | 2015 | Leave-one-out Word Alignment without Garbage Collector Effects · EMNLP 2015 |
Natural language and speech › Machine translation
multimodal machine translation |
0.1 | 1 | 2011 | picoTrans: Using Pictures as Input for Machine Translation on Mobile Devices · IJCAI 2011 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology
ontological knowledge |
0.1 | 1 | 2006 | Using Lexical Dependency and Ontological Knowledge to Improve a Detailed Syntactic and Semantic Tagger of English · ACL 2006 |
Interaction techniques and input › input modality
multimodal input |
0.0 | 1 | 2011 | picoTrans: Using Pictures as Input for Machine Translation on Mobile Devices · IJCAI 2011 |
Natural language and speech › Information extraction and text analysis › sequence labeling
part-of-speech tagging |
0.0 | 1 | 1999 | Applying Extrasentential Context To Maximum Entropy Based Tagging With A Large Semantic And Syntactic Tagset · EMNLP 1999 |
Natural language and speech › Information extraction and text analysis
sequence labeling |
0.0 | 1 | 1999 | Applying Extrasentential Context To Maximum Entropy Based Tagging With A Large Semantic And Syntactic Tagset · EMNLP 1999 |
Mathematical optimization
relaxation |
0.0 | 1 | 1996 | Softening Discrete Relaxation · NIPS 1996 |
Methods — techniques the papers use, named apart from their topics
sentence embedding similarity · 0.3dynamic training · 0.3regularized training objective · 0.3maximum marginal likelihood estimation · 0.3sequence-to-sequence learning · 0.2long short-term memory · 0.2approximate search · 0.2leave-one-out · 0.2hierarchical phrase-based decoding · 0.2expectation-maximization · 0.2statistical machine translation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Scalable Multilingual Frontend for TTSabstractThis paper describes progress towards making a Neural Text-to-Speech (TTS) Frontend that works for many languages and can be easily extended to new languages. We take a Machine Translation (MT) inspired approach to constructing the frontend, and model both text normalization and pronunciation on a sentence level by building and using sequence-to-sequence (S2S) models. We experimented with training normalization and pronunciation as separate S2S models and with training a single S2S model combining both functions. For our language-independent approach to pronunciation we do not use a lexicon. Instead all pronunciations, including context-based pronunciations, are captured in the S2S model. We also present a language-independent chunking and splicing technique that allows us to process arbitrary-length sentences. Models for 18 languages were trained and evaluated. Many of the accuracy measurements are above 99%. We also evaluated the models in the context of end-to-end synthesis against our current production system. Alistair Conkie, Andrew M. Finch |
ICASSP | 2 |
| 2020 | Agreement on Target-Bidirectional Recurrent Neural Networks for Sequence-to-Sequence LearningabstractRecurrent neural networks are extremely appealing for sequence-to-sequence learning tasks. Despite their great success, they typically suffer from a shortcoming: they are prone to generate unbalanced targets with good prefixes but bad suffixes, and thus performance suffers when dealing with long sequences. We propose a simple yet effective approach to overcome this shortcoming. Our approach relies on the agreement between a pair of target-directional RNNs, which generates more balanced targets. In addition, we develop two efficient approximate search methods for agreement that are empirically shown to be almost optimal in terms of either sequence level or non-sequence level metrics. Extensive experiments were performed on three standard sequence-to-sequence transduction tasks: machine transliteration, grapheme-to-phoneme transformation and machine translation. The results show that the proposed approach achieves consistent and substantial improvements, compared to many state-of-the-art systems. Lemao Liu, Andrew M. Finch, Masao Utiyama, Eiichiro Sumita |
J. Artif. Intell. Res. | 2 |
| 2018 | Extraction of templates from phrases using Sequence Binary Decision DiagramsabstractAbstract The extraction of templates such as ‘regard X as Y’ from a set of related phrases requires the identification of their internal structures. This paper presents an unsupervised approach for extracting templates on-the-fly from only tagged text by using a novel relaxed variant of the Sequence Binary Decision Diagram (SeqBDD). A SeqBDD can compress a set of sequences into a graphical structure equivalent to a minimal deterministic finite state automata, but more compact and better suited to the task of template extraction. The main contribution of this paper is a relaxed form of the SeqBDD construction algorithm that enables it to form general representations from a small amount of data. The process of compression of shared structures in the text during Relaxed SeqBDD construction, naturally induces the templates we wish to extract. Experiments show that the method is capable of high-quality extraction on tasks based on verb+preposition templates from corpora and phrasal templates from short messages from social media. Daiki Hirano, Kumiko Tanaka-Ishii, Andrew M. Finch |
Nat. Lang. Eng. | 3 |
| 2018 | Sentence Selection and Weighting for Neural Machine Translation Domain AdaptationabstractNeural machine translation (NMT) has been prominent in many machine translation tasks. However, in some domain-specific tasks, only the corpora from similar domains can improve translation performance. If out-of-domain corpora are directly added into the in-domain corpus, the translation performance may even degrade. Therefore, domain adaptation techniques are essential to solve the NMT domain problem. Most existing methods for domain adaptation are designed for the conventional phrase-based machine translation. For NMT domain adaptation, there have been only a few studies on topics such as fine tuning, domain tags, and domain features. In this paper, we have four goals for sentence level NMT domain adaptation. First, the NMT's internal sentence embedding is exploited and the sentence embedding similarity is used to select out-of-domain sentences that are close to the in-domain corpus. Second, we propose three sentence weighting methods, i.e., sentence weighting, domain weighting, and batch weighting, to balance the data distribution during NMT training. Third, in addition, we propose dynamic training methods to adjust the sentence selection and weighting during NMT training. Fourth, to solve the multidomain problem in a real-world NMT scenario where the domain distributions of training and testing data often mismatch, we proposed a multidomain sentence weighting method to balance the domain distributions of training data and match the domain distributions of training and testing data. The proposed methods are evaluated in international workshop on spoken language translation (IWSLT) English-to-French/German tasks and a multidomain English-to-French task. Empirical results show that the sentence selection and weighting methods can significantly improve the NMT performance, outperforming the existing baselines. Rui Wang 0015, Masao Utiyama, Andrew M. Finch, Lemao Liu, Kehai Chen, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | A Target Attention Model for Neural Machine Translation
Hideya Mino, Andrew M. Finch, Eiichiro Sumita |
MTSummit (1) | 2 |
| 2017 | Inducing a Bilingual Lexicon from Short Parallel Multiword SequencesabstractThis article proposes a technique for mining bilingual lexicons from pairs of parallel short word sequences. The technique builds a generative model from a corpus of training data consisting of such pairs. The model is a hierarchical nonparametric Bayesian model that directly induces a bilingual lexicon while training. The model learns in an unsupervised manner and is designed to exploit characteristics of the language pairs being mined. The proposed model is capable of utilizing commonly used word-pair frequency information and additionally can employ the internal character alignments within the words themselves. It is thereby capable of mining transliterations and can use reliably aligned transliteration pairs to support the mining of other words in their context. The model is also capable of performing word reordering and word deletion during the alignment process, and it is furthermore capable of operating in the absence of full segmentation information. In this work, we study two mining tasks based on English-Japanese and English-Chinese language pairs, and compare the proposed approach to baselines based on a simpler models that use only word-pair frequency information. Our results show that the proposed method is able to mine bilingual word pairs at higher levels of precision and recall than the baselines. Andrew M. Finch, Taisuke Harada, Kumiko Tanaka-Ishii, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2017 | Translation Quality Estimation Using Only Bilingual CorporaabstractIn computer-aided translation scenarios, quality estimation of machine translation hypotheses plays a critical role. Existing methods for word-level translation quality estimation (TQE) rely on the availability of manually annotated TQE training data obtained via direct annotation or postediting. However, due to the cost of human labor, such data are either limited in size or is only available for few tasks in practice. To avoid the reliance on such annotated TQE data, this paper proposes an approach to train word-level TQE models using bilingual corpora, which are typically used in machine translation training and is relatively easier to access. We formalize the training of our proposed method under the framework of maximum marginal likelihood estimation. To avoid degenerated solutions, we propose a novel regularized training objective whose optimization is achieved by an efficient approximation. Extensive experiments on both written and spoken language datasets empirically show that our approach yields comparable performance to the standard training on annotated data. Lemao Liu, Atsushi Fujita, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2016 | Agreement on Target-Bidirectional LSTMs for Sequence-to-Sequence LearningabstractRecurrent neural networks, particularly the long short- term memory networks, are extremely appealing for sequence-to-sequence learning tasks. Despite their great success, they typically suffer from a fundamental short- coming: they are prone to generate unbalanced targets with good prefixes but bad suffixes, and thus perfor- mance suffers when dealing with long sequences. We propose a simple yet effective approach to overcome this shortcoming. Our approach relies on the agreement between a pair of target-directional LSTMs, which generates more balanced targets. In addition, we develop two efficient approximate search methods for agreement that are empirically shown to be almost optimal in terms of sequence-level losses. Extensive experiments were performed on two standard sequence-to-sequence trans- duction tasks: machine transliteration and grapheme-to- phoneme transformation. The results show that the proposed approach achieves consistent and substantial im- provements, compared to six state-of-the-art systems. In particular, our approach outperforms the best reported error rates by a margin (up to 9% relative gains) on the grapheme-to-phoneme task. Lemao Liu, Andrew M. Finch, Masao Utiyama, Eiichiro Sumita |
AAAI | 2 |
| 2016 | Neural Machine Translation with Supervised AttentionabstractThe attention mechanism is appealing for neural machine translation, since it is able to dynamically encode a source sentence by generating a alignment between a target word and source words. Unfortunately, it has been proved to be worse than conventional alignment models in alignment accuracy. In this paper, we analyze and explain this issue from the point view of reordering, and propose a supervised attention which is learned with guidance from conventional alignment models. Experiments on two Chinese-to-English translation tasks show that the supervised attention mechanism yields better alignments leading to substantial gains over the standard attention based NMT. Lemao Liu, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
COLING | 3 |
| 2016 | Introducing the Asian Language Treebank (ALT)
Ye Kyaw Thu, Win Pa Pa, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
LREC | 4 |
| 2016 | Agreement on Target-bidirectional Neural Machine TranslationabstractLemao Liu, Masao Utiyama, Andrew Finch, Eiichiro Sumita. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Lemao Liu, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
HLT-NAACL | 3 |
| 2016 | Interlocking Phrases in Phrase-based Statistical Machine TranslationabstractThis paper presents an study of the use of interlocking phrases in phrase-based statistical machine translation. We examine the effect on translation quality when the translation units used in the translation hypotheses are allowed to overlap on the source side, on the target side and on both sides. A large-scale evaluation on 380 language pairs was conducted. Our results show that overall the use of overlapping phrases improved translation quality by 0.3 BLEU points on average. Further analysis revealed that language pairs requiring a larger amount of re-ordering benefited the most from our approach. When the evaluation was restricted to such pairs, the average improvement increased to up to 0.75 BLEU points with over 97% of the pairs improving. Our approach requires only a simple modification to the decoding algorithm and we believe it should be generally applicable to improve the performance of phrase-based decoders. Ye Kyaw Thu, Andrew M. Finch, Eiichiro Sumita |
HLT-NAACL | 2 |
| 2015 | Hierarchical Phrase-based Stream DecodingabstractThis paper proposes a method for hierarchical phrase-based stream decoding.A stream decoder is able to take a continuous stream of tokens as input, and segments this stream into word sequences that are translated and output as a stream of target word sequences.Phrase-based stream decoding techniques have been shown to be effective as a means of simultaneous interpretation.In this paper we transfer the essence of this idea into the framework of hierarchical machine translation.The hierarchical decoding framework organizes the decoding process into a chart; this structure is naturally suited to the process of stream decoding, leading to an efficient stream decoding algorithm that searches a restricted subspace containing only relevant hypotheses.Furthermore, the decoder allows more explicit access to the word re-ordering process that is of critical importance in decoding while interpreting.The decoder was evaluated on TED talk data for English-Spanish and English-Chinese.Our results show that like the phrase-based stream decoder, the hierarchical is capable of approaching the performance of the underlying hierarchical phrase-based machine translation decoder, at useful levels of latency.In addition the hierarchical approach appeared to be robust to the difficulties presented by the more challenging English-Chinese task. Andrew M. Finch, Xiaolin Wang 0002, Masao Utiyama, Eiichiro Sumita |
EMNLP | 1 |
| 2015 | Leave-one-out Word Alignment without Garbage Collector EffectsabstractExpectation-maximization algorithms, such as those implemented in GIZA++ pervade the field of unsupervised word alignment.However, these algorithms have a problem of over-fitting, leading to "garbage collector effects," where rare words tend to be erroneously aligned to untranslated words.This paper proposes a leave-one-out expectationmaximization algorithm for unsupervised word alignment to address this problem.The proposed method excludes information derived from the alignment of a sentence pair from the alignment models used to align it.This prevents erroneous alignments within a sentence pair from supporting themselves.Experimental results on Chinese-English and Japanese-English corpora show that the F 1 , precision and recall of alignment were consistently increased by 5.0% -17.2%, and BLEU scores of end-to-end translation were raised by 0.03 -1.30.The proposed method also outperformed l 0 -normalized GIZA++ and Kneser-Ney smoothed GIZA++. Xiaolin Wang 0002, Masao Utiyama, Andrew M. Finch, Taro Watanabe, Eiichiro Sumita |
EMNLP | 3 |
| 2015 | HMM based myanmar text to speech system
Ye Kyaw Thu, Win Pa Pa, Jinfu Ni, Yoshinori Shiga, Andrew M. Finch, Chiori Hori, Hisashi Kawai, Eiichiro Sumita |
INTERSPEECH | 5 |
| 2015 | Learning bilingual phrase representations with recurrent neural networks
Hideya Mino, Andrew M. Finch, Eiichiro Sumita |
MTSummit | 2 |
| 2015 | A Large-scale Study of Statistical Machine Translation Methods for Khmer Language
Ye Kyaw Thu, Vichet Chea, Andrew M. Finch, Masao Utiyama, Eiichiro Sumita |
PACLIC | 3 |
| 2014 | Refining Word Segmentation Using a Manually Aligned Corpus for Statistical Machine TranslationabstractLanguages that have no explicit word delimiters often have to be segmented for statistical machine translation (SMT).This is commonly performed by automated segmenters trained on manually annotated corpora.However, the word segmentation (WS) schemes of these annotated corpora are handcrafted for general usage, and may not be suitable for SMT.An analysis was performed to test this hypothesis using a manually annotated word alignment (WA) corpus for Chinese-English SMT.An analysis revealed that 74.60% of the sentences in the WA corpus if segmented using an automated segmenter trained on the Penn Chinese Treebank (CTB) will contain conflicts with the gold WA annotations.We formulated an approach based on word splitting with reference to the annotated WA to alleviate these conflicts.Experimental results show that the refined WS reduced word alignment error rate by 6.82% and achieved the highest BLEU improvement (0.63 on average) on the Chinese-English open machine translation (OpenMT) corpora compared to related work. Xiaolin Wang 0002, Masao Utiyama, Andrew M. Finch, Eiichiro Sumita |
EMNLP | 3 |
| 2013 | Inducing Romanization Systems
Keiko Taguchi, Andrew M. Finch, Seiichi Yamamoto, Eiichiro Sumita |
MTSummit | 2 |
| 2013 | A-STAR: Toward translating Asian spoken languages
Sakriani Sakti, Michael Paul, Andrew M. Finch, Shinsuke Sakai, Thang Tat Vu, Noriyuki Kimura, Chiori Hori, Eiichiro Sumita, Satoshi Nakamura 0001, Jun Park, Chai Wutiwiwatchai, Bo Xu 0002, Hammam Riza, Karunesh Arora, Haizhou Li 0001 |
Comput. Speech Lang. | 3 |
| 2013 | A Bayesian Alignment Approach to Transliteration MiningabstractIn this article we present a technique for mining transliteration pairs using a set of simple features derived from a many-to-many bilingual forced-alignment at the grapheme level to classify candidate transliteration word pairs as correct transliterations or not. We use a nonparametric Bayesian method for the alignment process, as this process rewards the reuse of parameters, resulting in compact models that align in a consistent manner and tend not to over-fit. Our approach uses the generative model resulting from aligning the training data to force-align the test data. We rely on the simple assumption that correct transliteration pairs would be well modeled and generated easily, whereas incorrect pairs---being more random in character---would be more costly to model and generate. Our generative model generates by concatenating bilingual grapheme sequence pairs. The many-to-many generation process is essential for handling many languages with non-Roman scripts, and it is hard to train well using a maximum likelihood techniques, as these tend to over-fit the data. Our approach works on the principle that generation using only grapheme sequence pairs that are in the model results in a high probability derivation, whereas if the model is forced to introduce a new parameter in order to explain part of the candidate pair, the derivation probability is substantially reduced and severely reduced if the new parameter corresponds to a sequence pair composed of a large number of graphemes. The features we extract from the alignment of the test data are not only based on the scores from the generative model, but also on the relative proportions of each sequence that are hard to generate. The features are used in conjunction with a support vector machine classifier trained on known positive examples together with synthetic negative examples to determine whether a candidate word pair is a correct transliteration pair. In our experiments, we used all data tracks from the 2010 Named-Entity Workshop (NEWS’10) and use the performance of the best system for each language pair as a reference point. Our results show that the new features we propose are powerfully predictive, enabling our approach to achieve levels of performance on this task that are comparable to the state of the art. Takaaki Fukunishi, Andrew M. Finch, Seiichi Yamamoto, Eiichiro Sumita |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2013 | How to Choose the Best Pivot Language for Automatic Translation of Low-Resource LanguagesabstractRecent research on multilingual statistical machine translation focuses on the usage of pivot languages in order to overcome language resource limitations for certain language pairs. Due to the richness of available language resources, English is, in general, the pivot language of choice. However, factors like language relatedness can also effect the choice of the pivot language for a given language pair, especially for Asian languages, where language resources are currently quite limited. In this article, we provide new insights into what factors make a pivot language effective and investigate the impact of these factors on the overall pivot translation performance for translation between 22 Indo-European and Asian languages. Experimental results using state-of-the-art statistical machine translation techniques revealed that the translation quality of 54.8% of the language pairs improved when a non-English pivot language was chosen. Moreover, 81.0% of system performance variations can be explained by a combination of factors such as language family, vocabulary, sentence length, language perplexity, translation model entropy, reordering, monotonicity, and engine performance. Michael Paul, Andrew M. Finch, Eiichiro Sumita |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2013 | picoTrans: An intelligent icon-driven interface for cross-lingual communicationabstractpicoTrans is a prototype system that introduces a novel icon-based paradigm for cross-lingual communication on mobile devices. Our approach marries a machine translation system with the popular picture book. Users interact with picoTrans by pointing at pictures as if it were a picture book; the system generates natural language from these icons and the user is able to interact with the icon sequence to refine the meaning of the words that are generated. When users are satisfied that the sentence generated represents what they wish to express, they tap a translate button and picoTrans displays the translation. Structuring the process of communication in this way has many advantages. First, tapping icons is a very natural method of user input on mobile devices; typing is cumbersome and speech input errorful. Second, the sequence of icons which is annotated both with pictures and bilingually with words is meaningful to both users, and it opens up a second channel of communication between them that conveys the gist of what is being expressed. We performed a number of evaluations of picoTrans to determine: its coverage of a corpus of in-domain sentences; the input efficiency in terms of the number of key presses required relative to text entry; and users' overall impressions of using the system compared to using a picture book. Our results show that we are able to cover 74% of the expressions in our test corpus using a 2000-icon set; we believe that this icon set size is realistic for a mobile device. We also found that picoTrans requires fewer key presses than typing the input and that the system is able to predict the correct, intended natural language sentence from the icon sequence most of the time, making user interaction with the icon sequence often unnecessary. In the user evaluation, we found that in general users prefer using picoTrans and are able to communicate more rapidly and expressively. Furthermore, users had more confidence that they were able to communicate effectively using picoTrans. Andrew M. Finch, Kumiko Tanaka-Ishii, Keiji Yasuda, Eiichiro Sumita |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2012 | Method to Build a Bilingual Lexicon for Speech-to-Speech Translation Systems
Keiji Yasuda, Andrew M. Finch, Eiichiro Sumita |
CICLing (2) | 2 |
| 2011 | Word Segmentation for Dialect Translation
Michael Paul, Andrew M. Finch, Eiichiro Sumita |
CICLing (2) | 2 |
| 2011 | A Method to Measure the Reading Difficulty of Japanese Words
Keiji Yasuda, Andrew M. Finch, Eiichiro Sumita |
CICLing (2) | 2 |
| 2011 | Unsupervised determination of efficient Korean LVCSR units using a Bayesian Dirichlet process modelabstractKorean is an agglutinative language that does not have explicit word boundaries. It is also a highly inflective language that exhibits severe coarticulation effects. These characteristics pose a challenge in developing large-vocabulary continuous speech recognition (LVCSR) systems. Many existing Korean LVCSR systems attempt to overcome these difficulties by defining a set of "word" units using morphological analysis (rule-based) or statistical methods. These approaches usually require a great deal of linguistic knowledge or at least some explicit information about the statistical distribution of the units. However, exceptions or uncommon words (e.g., foreign proper nouns) still exist that cannot be covered by rules alone. In this paper, we investigate the use of an unsupervised, nonparametric Bayesian approach to automatically determining efficient units for a Korean LVCSR system. Specifically, we utilize a Dirichlet process model trained using Bayesian inference through block Gibbs sampling. Our approach provides a principled way of learning units without explicit linguistic knowledge or any static parameters. Experiments were conducted on a travel domain corpus, which includes many foreign words and proper nouns. In our experiments we compared our method to a set of state-of-the-art baseline systems that relied on either morphological analysis or segmentation heuristics. Our system was able to produce a considerably more compact set of "word" units than the best baseline system (the lexical dictionary was approximately half the size), with a recognition accuracy 5.89% higher in terms of the relative word error rate than the best baseline system. Sakriani Sakti, Andrew M. Finch, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
ICASSP | 2 |
| 2011 | picoTrans: Using Pictures as Input for Machine Translation on Mobile DevicesabstractIn this paper we present a novel user interface that integrates two popular approaches to language translation for travelers allowing multimodal communication between the parties involved: the picture-book, in which the user simply points to multiple picture icons representing what they want to say, and the statistical machine translation (SMT) system that can translate arbitrary word sequences. Our prototype system tightly couples both processes within a translation framework that inherits many of the the positive features of both approaches, while at the same time mitigating their main weaknesses. Our system differs from traditional approaches in that its mode of input is a sequence of pictures, rather than text or speech. Text in the source language is generated automatically, and is used as a detailed representation of the intended meaning. The picture sequence which not only provides a rapid method to communicate basic concepts but also gives a 'second opinion' on the machine transition output that catches machine translation errors and allows the users to retry the translation, avoiding misunderstandings. Andrew M. Finch, Kumiko Tanaka-Ishii, Eiichiro Sumita |
IJCAI | 1 |
| 2011 | picoTrans: an icon-driven user interface for machine translation on mobile devicesabstractIn this paper we present a novel user interface that integrates two popular approaches to language translation for travelers allowing multimodal communication between the parties involved. In our approach we integrate the popular picture-book, in which the user simply points to multiple picture icons representing what they want to say, with a statistical machine translation system that can translate arbitrary word sequences. The simple pointing at pictures paradigm is used as the primary method of user input and the users can use the device as if it were a picture book. The application is then able to generate a complete sentence in the user's native language for what they wish to say from the sequence of picture icons chosen by the user. Once the user is satisfied that the sentence provided by the system adequately represents what they wish to convey, the application can automatically translate the sentence into the language of the other party, who can interpret the intended meaning of the first party by combining evidence from both modes of communication: the picture sequence, and the machine translation. The prototype system we have developed inherits many of the positive features of both approaches, while at the same time mitigating their main weaknesses. The user may combine the pictures in considerably more combinations than is possible with a picture book designed with combinations from only within the same page spread of the book in mind, making the application more expressive than a book. The machine translation system can contribute a detailed and precise translation which is supported by the picture-based mode which not only provides a rapid method to communicate basic concepts but also gives a 'second opinion' on the machine transition output that catches machine translation errors and allows the users to retry the sentence, avoiding misunderstandings. Andrew M. Finch, Kumiko Tanaka-Ishii, Eiichiro Sumita |
IUI | 2 |
| 2009 | Bidirectional Phrase-based Statistical Machine Translation
Andrew M. Finch, Eiichiro Sumita |
EMNLP | 1 |
| 2008 | Phrase-based Machine Transliteration
Andrew M. Finch, Eiichiro Sumita |
IJCNLP | 1 |
| 2007 | Improving statistical machine translation using shallow linguistic knowledge
Young-Sook Hwang, Andrew M. Finch, Yutaka Sasaki |
Comput. Speech Lang. | 2 |
| 2006 | Using Lexical Dependency and Ontological Knowledge to Improve a Detailed Syntactic and Semantic Tagger of English
Andrew M. Finch, Ezra Black, Young-Sook Hwang, Eiichiro Sumita |
ACL | 1 |
| 2004 | How Does Automatic Machine Translation Evaluation Correlate with Human Scoring as the Number of Reference Translations Increases?
Andrew M. Finch, Yasuhiro Akiba, Eiichiro Sumita |
LREC | 1 |
| 2003 | A corpus-centered approach to spoken language translation
Eiichiro Sumita, Yasuhiro Akiba, Takao Doi, Andrew M. Finch, Kenji Imamura, Michael Paul, Mitsuo Shimohata, Taro Watanabe |
EACL | 4 |
| 2002 | Beyond Tag Trigrams: New Local Features for Tagging
Andrew M. Finch, Ezra Black, Ringo Wathelet |
LREC | 1 |
| 2000 | Integrating detailed information into a language modelabstractApplying natural language processing technique to language modeling is a key problem in speech recognition. This paper describes a maximum entropy-based approach to language modeling in which both words together with syntactic and semantic tags in the long history are used as a basis for complex linguistic questions. These questions are integrated with a standard trigram language model or a standard trigram language model combined with long history word triggers and the resulting language model is used to rescore the N-best hypotheses output of the ATRSPREC speech recognition system. The technique removed 24% of the correctable error of the recognition system. Ruiqiang Zhang, Ezra Black, Andrew M. Finch, Yoshinori Sagisaka |
ICASSP | 3 |
| 2000 | A tagger-aided language model with a stack decoder
Ruiqiang Zhang, Ezra Black, Andrew M. Finch, Yoshinori Sagisaka |
INTERSPEECH | 3 |
| 1999 | Applying Extrasentential Context To Maximum Entropy Based Tagging With A Large Semantic And Syntactic Tagset
Ezra Black, Andrew M. Finch, Ruigiang Zhang |
EMNLP | 2 |
| 1999 | Using detailed linguistic structure in language modellingabstractMuch recent research has demonstrated that the correlation between a language model's perplexity and its effect on the word error rate of a speech recognition system is not as strong as was once thought.This represents a major problem for those involved in developing language models.This paper describes the development of new measures of language model quality.These measures retain the ease of computation and task independence that are perplexity's strengths, yet are considerably better correlated with word error rate.This paper also shows that mixture-based language models are improved by applying interpolation weights which are optimised with respect to these new measures, rather than a maximum likelihood criterion. Ruiqiang Zhang, Ezra Black, Andrew M. Finch |
EUROSPEECH | 3 |
| 1998 | An Energy Function and Continuous Edit Process for Graph MatchingabstractThe contributions of this article are twofold. First, we develop a new nonquadratic energy function for graph matching. The starting point is a recently reported mixture model that gauges relational consistency using a series of exponential functions of the Hamming distances between graph neighborhoods. We compute the effective neighborhood potentials associated with the mixture model by identifying the single probability function of zero Kullback divergence. This new energy function is simply a weighted sum of graph Hamming distances. The second contribution is to locate matches by graduated assignment. Rather than solving the mean-field saddle-point equations, which are intractable for our nonquadratic energy function, we apply the soft-assign ansatz to the derivatives of our energy function. Here we introduce a novel departure from the standard graduated assignment formulation of graph matching by allowing the connection strengths of the data graph to update themselves. The aim is to provide a means by which the structure of the data graph can be updated so as to rectify structural errors. The method is evaluated experimentally and is shown to outperform its quadratic counterpart. Andrew M. Finch, Richard C. Wilson 0001, Edwin R. Hancock |
Neural Comput. | 1 |
| 1998 | Symbolic graph matching with the EM algorithm
Andrew M. Finch, Richard C. Wilson 0001, Edwin R. Hancock |
Pattern Recognit. | 1 |
| 1997 | Matching delaunay graphs
Andrew M. Finch, Richard C. Wilson 0001, Edwin R. Hancock |
Pattern Recognit. | 1 |
| 1996 | Relational matching with mean field annealingabstractThis paper describes a new framework for constructing mean-field energy functions for use in relational matching. The starting point is the Bayesian relational consistency model of Wilson and Hancock (1995). Hitherto, the optimisation of the consistency measure has been effected by the deterministic hill climbing process known as discrete relaxation which is prone to local convergence if local maxima are present. By applying ideas from statistical physics to the configurational matching probabilities we determine the effective potentials of an equivalent Boltzmann distribution. Formally, these potentials are weighted sums of the Hamming distances between matched neighbourhoods in the data graph and their counterparts in the model graph. Adopting a simple softening ansatz we derive mean-field equations for minimising the global graph matching potential. This provides an efficient means of locating the global optima of the Bayesian consistency measure. Andrew M. Finch, Richard C. Wilson 0001, Edwin R. Hancock |
ICPR | 1 |
| 1996 | Softening Discrete Relaxation
Andrew M. Finch, Richard C. Wilson 0001, Edwin R. Hancock |
NIPS | 1 |
| 1995 | Matching Delauny Triangulations by Probabilistic Relaxation
Andrew M. Finch, Richard C. Wilson 0001, Edwin R. Hancock |
CAIP | 1 |