EDBT 2026 Demo / reviewers in the wild / expert
David Chiang 0001
dblp:89/233-1
· DBLP profile ↗
75ranked-venue papers
19as first author
23since 2021 · last 2026
0000-0002-0435-4864ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 72 · 18 first-author · 22 since 2021Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Theory of computation · 2Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automatic Speech Recognition for Documenting Endangered Languages: Case Study of Ikema Miyakoan
Chihiro Taguchi, Yukinori Takubo, David Chiang 0001 |
LREC | 3 |
| 2026 | Simulating Hard Attention Using Soft AttentionabstractAbstract We study conditions under which transformers using soft attention can simulate hard attention, that is, effectively focus all attention on a subset of positions. First, we examine several subclasses of languages recognized by hard-attention transformers, which can be defined in variants of linear temporal logic. We demonstrate how soft-attention transformers can compute formulas of these logics using unbounded positional embeddings or temperature scaling. Second, we demonstrate how temperature scaling allows softmax transformers to simulate general hard-attention transformers, using a temperature that depends on the minimum gap between the maximum attention scores and other attention scores. Andy Yang, Lena Strobl, David Chiang 0001, Dana Angluin |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | Languages Still Left Behind: Toward a Better Multilingual Machine Translation BenchmarkabstractMultilingual machine translation (MT) benchmarks play a central role in evaluating the capabilities of modern MT systems.Among them, the FLORES+ benchmark is widely used, offering English-to-many translation data for over 200 languages, curated with strict quality control protocols.However, we study data in four languages (Asante Twi, Japanese, Jinghpaw, and South Azerbaijani) and uncover critical shortcomings in the benchmark's suitability for truly multilingual evaluation.Human assessments reveal that many translations fall below the claimed 90% quality standard, and the annotators report that source sentences are often too domain-specific and culturally biased toward the English-speaking world.We further demonstrate that simple heuristics, such as copying named entities, can yield non-trivial BLEU scores, suggesting vulnerabilities in the evaluation protocol.Notably, we show that MT models trained on high-quality, naturalistic data perform poorly on FLORES+ while achieving significant gains on our domain-relevant evaluation set.Based on these findings, we advocate for multilingual MT benchmarks that use domain-general and culturally neutral source texts rely less on named entities, in order to better reflect real-world translation challenges. 1 * Equal contribution. Chihiro Taguchi, Seng Mai, Keita Kurabe, Yusuke Sakai 0010, Georgina Agyei, Soudabeh Eslami, David Chiang 0001 |
EMNLP | 7 |
| 2025 | Knee-Deep in C-RASP: A Transformer Depth HierarchyabstractIt has been observed that transformers with greater depth (that is, more layers) have more capabilities, but can we establish formally which capabilities are gained? We answer this question with a theoretical proof followed by an empirical study. First, we consider transformers that round to fixed precision except inside attention. We show that this subclass of transformers is expressively equivalent to the programming language $\textsf{C}$-$\textsf{RASP}$ and this equivalence preserves depth. Second, we prove that deeper $\textsf{C}$-$\textsf{RASP}$ programs are more expressive than shallower $\textsf{C}$-$\textsf{RASP}$ programs, implying that deeper transformers are more expressive than shallower transformers (within the subclass mentioned above). The same is also proven for transformers with positional encodings (like RoPE and ALiBi). These results are established by studying a temporal logic with counting operators equivalent to $\textsf{C}$-$\textsf{RASP}$. Finally, we provide empirical evidence that our theory predicts the depth required for transformers without positional encodings to length-generalize on a family of sequential dependency tasks. Andy Yang, Michaël Cadilhac, David Chiang 0001 |
NeurIPS | 3 |
| 2025 | Transformers as TransducersabstractAbstract We study the sequence-to-sequence mapping capacity of transformers by relating them to finite transducers, and find that they can express surprisingly large classes of (total functional) transductions. We do so using variants of RASP, a programming language designed to help people “think like transformers,” as an intermediate representation. We extend the existing Boolean variant B-RASP to sequence-to-sequence transductions and show that it computes exactly the first-order rational transductions (such as string rotation). Then, we introduce two new extensions. B-RASP[pos] enables calculations on positions (such as copying the first half of a string) and contains all first-order regular transductions. S-RASP adds prefix sum, which enables additional arithmetic operations (such as squaring a string) and contains all first-order polyregular transductions. Finally, we show that masked average-hard attention transformers can simulate S-RASP. Lena Strobl, Dana Angluin, David Chiang 0001, Jonathan Rawski, Ashish Sabharwal |
Trans. Assoc. Comput. Linguistics | 3 |
| 2024 | DIALECTBENCH: An NLP Benchmark for Dialects, Varieties, and Closely-Related LanguagesabstractFahim Faisal, Orevaoghene Ahia, Aarohi Srivastava, Kabir Ahuja, David Chiang, Yulia Tsvetkov, Antonios Anastasopoulos. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Fahim Faisal, Orevaoghene Ahia, Aarohi Srivastava, Kabir Ahuja, David Chiang 0001, Yulia Tsvetkov, Antonios Anastasopoulos |
ACL (1) | 5 |
| 2024 | Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn'tabstractWe investigate what linguistic factors affect the performance of Automatic Speech Recognition (ASR) models.We hypothesize that orthographic and phonological complexities both degrade accuracy.To examine this, we finetune the multilingual self-supervised pretrained model Wav2Vec2-XLSR-53 on 25 languages with 15 writing systems, and we compare their ASR accuracy, number of graphemes, unigram grapheme entropy, logographicity (how much word/morpheme-level information is encoded in the writing system), and number of phonemes.The results demonstrate that a high logographicity correlates with low ASR accuracy, while phonological complexity has no strong correlation. Chihiro Taguchi, David Chiang 0001 |
ACL (1) | 2 |
| 2024 | PILA: A Historical-Linguistic Dataset of Proto-Italic and LatinabstractComputational historical linguistics seeks to systematically understand processes of sound change, including during periods at which little to no formal recording of language is attested. At the same time, few computational resources exist which deeply explore phonological and morphological connections between proto-languages and their descendants. This is particularly true for the family of Italic languages. To assist historical linguists in the study of Italic sound change, we introduce the Proto-Italic to Latin (PILA) dataset, which consists of roughly 3,000 pairs of forms from Proto-Italic and Latin. We provide a detailed description of how our dataset was created and organized. Then, we exhibit PILA’s value in two ways. First, we present baseline results for PILA on a pair of traditional computational historical linguistics tasks. Second, we demonstrate PILA’s capability for enhancing other historical-linguistic datasets through a dataset compatibility study. Stephen Bothwell, Brian DuSell, David Chiang 0001, Brian Krostenko |
LREC/COLING | 3 |
| 2024 | Killkan: The Automatic Speech Recognition Dataset for Kichwa with Morphosyntactic InformationabstractThis paper presents Killkan, the first dataset for automatic speech recognition (ASR) in the Kichwa language, an indigenous language of Ecuador. Kichwa is an extremely low-resource endangered language, and there have been no resources before Killkan for Kichwa to be incorporated in applications of natural language processing. The dataset contains approximately 4 hours of audio with transcription, translation into Spanish, and morphosyntactic annotation in the format of Universal Dependencies, all done in ELAN, the annotation software. The audio data was retrieved from a publicly available radio program in Kichwa. This paper also provides corpus-linguistic analyses of the dataset with a special focus on the agglutinative morphology of Kichwa and frequent code-switching with Spanish. The experiments show that the dataset makes it possible to develop the first ASR system for Kichwa with reliable quality despite its small dataset size. This dataset, the ASR model, and the code used to develop them will be publicly available. Thus, our study positively showcases resource building and its applications for low-resource languages and their community. Chihiro Taguchi, Jefferson Saransig, Dayana Velásquez, David Chiang 0001 |
LREC/COLING | 4 |
| 2024 | Stack Attention: Improving the Ability of Transformers to Model Hierarchical PatternsabstractAttention, specifically scaled dot-product attention, has proven effective for natural language, but it does not have a mechanism for handling hierarchical patterns of arbitrary nesting depth, which limits its ability to recognize certain syntactic structures. To address this shortcoming, we propose stack attention: an attention operator that incorporates stacks, inspired by their theoretical connections to context-free languages (CFLs). We show that stack attention is analogous to standard attention, but with a latent model of syntax that requires no syntactic supervision. We propose two variants: one related to deterministic pushdown automata (PDAs) and one based on nondeterministic PDAs, which allows transformers to recognize arbitrary CFLs. We show that transformers with stack attention are very effective at learning CFLs that standard transformers struggle on, achieving strong results on a CFL with theoretically maximal parsing difficulty. We also show that stack attention is more effective at natural language modeling under a constrained parameter budget, and we include results on machine translation. Brian DuSell, David Chiang 0001 |
ICLR | 2 |
| 2024 | Masked Hard-Attention Transformers Recognize Exactly the Star-Free LanguagesabstractThe expressive power of transformers over inputs of unbounded size can be studied through their ability to recognize classes of formal languages. In this paper, we establish exact characterizations of transformers with hard attention (in which all attention is focused on exactly one position) and attention masking (in which each position only attends to positions on one side). With strict masking (each position cannot attend to itself) and without position embeddings, these transformers are expressively equivalent to linear temporal logic (LTL), which defines exactly the star-free languages. A key technique is the use of Boolean RASP as a convenient intermediate language between transformers and LTL. We then take numerous results known for LTL and apply them to transformers, showing how position embeddings, strict masking, and depth all increase expressive power. Andy Yang, David Chiang 0001, Dana Angluin |
NeurIPS | 2 |
| 2024 | What Formal Languages Can Transformers Express? A SurveyabstractAbstract As transformers have gained prominence in natural language processing, some researchers have investigated theoretically what problems they can and cannot solve, by treating problems as formal languages. Exploring such questions can help clarify the power of transformers relative to other models of computation, their fundamental capabilities and limits, and the impact of architectural choices. Work in this subarea has made considerable progress in recent years. Here, we undertake a comprehensive survey of this work, documenting the diverse assumptions that underlie different results and providing a unified framework for harmonizing seemingly contradictory findings. Lena Strobl, William Merrill, Gail Weiss, David Chiang 0001, Dana Angluin |
Trans. Assoc. Comput. Linguistics | 4 |
| 2023 | Convergence and Diversity in the Control HierarchyabstractWeir has defined a hierarchy of language classes whose second member (L 2 ) is generated by tree-adjoining grammars (TAG), linear indexed grammars (LIG), combinatory categorial grammars, and head grammars.The hierarchy is obtained using the mechanism of control, and L 2 is obtained using a contextfree grammar (CFG) whose derivations are controlled by another CFG.We adapt Weir's definition of a controllable CFG to give a definition of controllable pushdown automata (PDAs).This yields three new characterizations of L 2 as the class of languages generated by PDAs controlling PDAs, PDAs controlling CFGs, and CFGs controlling PDAs.We show that these four formalisms are not only weakly equivalent but equivalent in a stricter sense that we call d-weak equivalence.Furthermore, using an even stricter notion of equivalence called d-strong equivalence, we make precise the intuition that a CFG controlling a CFG is a TAG, a PDA controlling a PDA is an embedded PDA, and a PDA controlling a CFG is a LIG.The fourth member of this family, a CFG controlling a PDA, does not correspond to any formalism we know of, so we invent one and call it a Pushdown Adjoining Automaton. Alexandra Butoi, Ryan Cotterell, David Chiang 0001 |
ACL (1) | 3 |
| 2023 | Introducing Rhetorical Parallelism Detection: A New Task with Datasets, Metrics, and BaselinesabstractRhetoric, both spoken and written, involves not only content but also style.One common stylistic tool is parallelism: the juxtaposition of phrases which have the same sequence of linguistic (e.g., phonological, syntactic, semantic) features.Despite the ubiquity of parallelism, the field of natural language processing has seldom investigated it, missing a chance to better understand the nature of the structure, meaning, and intent that humans convey.To address this, we introduce the task of rhetorical parallelism detection.We construct a formal definition of it; we provide one new Latin dataset and one adapted Chinese dataset for it; we establish a family of metrics to evaluate performance on it; and, lastly, we create baseline systems and novel sequence labeling schemes to capture it.On our strictest metric, we attain F 1 scores of 0.40 and 0.43 on our Latin and Chinese datasets, respectively. Stephen Bothwell, Justin DeBenedetto, Theresa Crnkovich, Hildegund Müller, David Chiang 0001 |
EMNLP | 5 |
| 2023 | Efficient Algorithms for Recognizing Weighted Tree-Adjoining LanguagesabstractThe class of tree-adjoining languages can be characterized by various two-level formalisms, consisting of a context-free grammar (CFG) or pushdown automaton (PDA) controlling another CFG or PDA.These four formalisms are equivalent to tree-adjoining grammars (TAG), linear indexed grammars (LIG), pushdownadjoining automata (PAA), and embedded pushdown automata (EPDA).We define semiringweighted versions of the above two-level formalisms, and we design new algorithms for computing their stringsums (the weight of all derivations of a string) and allsums (the weight of all derivations).From these, we also immediately obtain stringsum and allsum algorithms for TAG, LIG, PAA, and EPDA.For LIG, our algorithm is more time-efficient by a factor of O(n|N |) (where n is the string length and |N | is the size of the nonterminal set) and more space-efficient by a factor of O(|Γ|) (where Γ is the size of the stack alphabet) than the algorithm of Vijay-Shanker and Weir (1989).For EPDA, our algorithm is both more spaceefficient and time-efficient than the algorithm of Alonso et al. (2001) by factors of O(|Γ| 2 ) and O(|Γ| 3 ), respectively.Finally, we give the first PAA stringsum and allsum algorithms. Alexandra Butoi, Tim Vieira, Ryan Cotterell, David Chiang 0001 |
EMNLP | 4 |
| 2023 | The Surprising Computational Power of Nondeterministic Stack RNNs
Brian DuSell, David Chiang 0001 |
ICLR | 2 |
| 2023 | Tighter Bounds on the Expressivity of Transformer EncodersabstractCharacterizing neural networks in terms of better-understood formal systems has the potential to yield new insights into the power and limitations of these networks. Doing so for transformers remains an active area of research. Bhattamishra and others have shown that transformer encoders are at least as expressive as a certain kind of counter machine, while Merrill and Sabharwal have shown that fixed-precision transformer encoders recognize only languages in uniform $TC^0$. We connect and strengthen these results by identifying a variant of first-order logic with counting quantifiers that is simultaneously an upper bound for fixed-precision transformer encoders and a lower bound for transformer encoders. This brings us much closer than before to an exact characterization of the languages that transformer encoders recognize. David Chiang 0001, Peter Cholak, Anand Pillay |
ICML | 1 |
| 2023 | Universal Automatic Phonetic Transcription into the International Phonetic Alphabet
Chihiro Taguchi, Yusuke Sakai 0010, Parisa Haghani, David Chiang 0001 |
INTERSPEECH | 4 |
| 2023 | Exact Recursive Probabilistic ProgrammingabstractRecursive calls over recursive data are useful for generating probability distributions, and probabilistic programming allows computations over these distributions to be expressed in a modular and intuitive way. Exact inference is also useful, but unfortunately, existing probabilistic programming languages do not perform exact inference on recursive calls over recursive data, forcing programmers to code many applications manually. We introduce a probabilistic language in which a wide variety of recursion can be expressed naturally, and inference carried out exactly. For instance, probabilistic pushdown automata and their generalizations are easy to express, and polynomial-time parsing algorithms for them are derived automatically. We eliminate recursive data types using program transformations related to defunctionalization and refunctionalization. These transformations are assured correct by a linear type system, and a successful choice of transformations, if there is one, is guaranteed to be found by a greedy algorithm. David Chiang 0001, Colin McDonald, Chung-chieh Shan |
Proc. ACM Program. Lang. | 1 |
| 2022 | Overcoming a Theoretical Limitation of Self-AttentionabstractAlthough transformers are remarkably effective for many tasks, there are some surprisingly easy-looking regular languages that they struggle with.Hahn shows that for languages where acceptance depends on a single input symbol, a transformer's classification decisions become less and less confident (that is, with crossentropy approaching 1 bit per string) as input strings get longer and longer.We examine this limitation using two languages: PAR-ITY, the language of bit strings with an odd number of 1s, and FIRST, the language of bit strings starting with a 1.We demonstrate three ways of overcoming the limitation suggested by Hahn's lemma.First, we settle an open question by constructing a transformer that recognizes PARITY with perfect accuracy, and similarly for FIRST.Second, we use layer normalization to bring the cross-entropy of both models arbitrarily close to zero.Third, when transformers need to focus on a single position, as for FIRST, we find that they can fail to generalize to longer strings; we offer a simple remedy to this problem that also improves length generalization in machine translation. David Chiang 0001, Peter Cholak |
ACL (1) | 1 |
| 2022 | Algorithms for Weighted Pushdown AutomataabstractWeighted pushdown automata (WPDAs) are at the core of many natural language processing tasks, like syntax-based statistical machine translation and transition-based dependency parsing.As most existing dynamic programming algorithms are designed for context-free grammars (CFGs), algorithms for PDAs often resort to a PDA-to-CFG conversion.In this paper, we develop novel algorithms that operate directly on WPDAs.Our algorithms are inspired by Lang's algorithm, but use a more general definition of pushdown automaton and either reduce the space requirements by a factor of |Γ| (the size of the stack alphabet) or reduce the runtime by a factor of more than |𝑄| (the number of states).When run on the same class of PDAs as Lang's algorithm, our algorithm is both more space-efficient by a factor of |Γ| and more time-efficient by a factor of |𝑄| • |Γ|. Alexandra Butoi, Brian DuSell, Tim Vieira, Ryan Cotterell, David Chiang 0001 |
EMNLP | 5 |
| 2022 | Learning Hierarchical Structures with Differentiable Nondeterministic Stacks
Brian DuSell, David Chiang 0001 |
ICLR | 2 |
| 2022 | Measuring Human Perception to Improve Handwritten Document TranscriptionabstractIn this paper, we consider how to incorporate psychophysical measurements of human visual perception into the loss function of a deep neural network being trained for a recognition task, under the assumption that such information can reduce errors. As a case study to assess the viability of this approach, we look at the problem of handwritten document transcription. While good progress has been made towards automatically transcribing modern handwriting, significant challenges remain in transcribing historical documents. Here we describe a general enhancement strategy, underpinned by the new loss formulation, which can be applied to the training regime of any deep learning-based document transcription system. Through experimentation, reliable performance improvement is demonstrated for the standard IAM and RIMES datasets for three different network architectures. Further, we go on to show feasibility for our approach on a new dataset of digitized Latin manuscripts, originally produced by scribes in the Cloister of St. Gall in the the 9th century. Samuel Grieggs, Bingyu Shen 0001, Greta Rauch, David Chiang 0001, Brian L. Price, Walter J. Scheirer |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | Learning Context-free Languages with Nondeterministic Stack RNNsabstractWe present a differentiable stack data structure that simultaneously and tractably encodes an exponential number of stack configurations, based on Lang's algorithm for simulating nondeterministic pushdown automata.We call the combination of this data structure with a recurrent neural network (RNN) controller a Nondeterministic Stack RNN.We compare our model against existing stack RNNs on various formal languages, demonstrating that our model converges more reliably to algorithmic behavior on deterministic tasks, and achieves lower cross-entropy on inherently nondeterministic tasks. Brian DuSell, David Chiang 0001 |
CoNLL | 2 |
| 2020 | Representing Unordered Data Using Complex-Weighted Multiset AutomataabstractUnordered, variable-sized inputs arise in many settings across multiple fields. The ability for set- and multiset-oriented neural networks to handle this type of input has been the focus of much work in recent years. We propose to represent multisets using complex-weighted multiset automata and show how the multiset representations of certain existing neural architectures can be viewed as special cases of ours. Namely, (1) we provide a new theoretical and intuitive justification for the Transformer model’s representation of positions using sinusoidal functions, and (2) we extend the DeepSets model to use complex numbers, enabling it to outperform the existing model on an extension of one of their tasks. Justin DeBenedetto, David Chiang 0001 |
ICML | 2 |
| 2020 | Factor Graph GrammarsabstractWe propose the use of hyperedge replacement graph grammars for factor graphs, or factor graph grammars (FGGs) for short. FGGs generate sets of factor graphs and can describe a more general class of models than plate notation, dynamic graphical models, case-factor diagrams, and sum-product networks can. Moreover, inference can be done on FGGs without enumerating all the generated factor graphs. For finite variable domains (but possibly infinite sets of graphs), a generalization of variable elimination to FGGs allows exact and tractable inference in many situations. For finite sets of graphs (but possibly infinite variable domains), a FGG can be converted to a single factor graph amenable to standard inference techniques. David Chiang 0001, Darcey Riley |
NeurIPS | 1 |
| 2019 | Accelerating Sparse Matrix Operations in Neural Networks on Graphics Processing UnitsabstractGraphics Processing Units (GPUs) are commonly used to train and evaluate neural networks efficiently.While previous work in deep learning has focused on accelerating operations on dense matrices/tensors on GPUs, efforts have concentrated on operations involving sparse data structures.Operations using sparse structures are common in natural language models at the input and output layers, because these models operate on sequences over discrete alphabets.We present two new GPU algorithms: one at the input layer, for multiplying a matrix by a few-hot vector (generalizing the more common operation of multiplication by a one-hot vector) and one at the output layer, for a fused softmax and top-N selection (commonly used in beam search).Our methods achieve speedups over state-of-theart parallel GPU baselines of up to 7× and 50×, respectively.We also illustrate how our methods scale on different GPU architectures. Arturo Argueta, David Chiang 0001 |
ACL (1) | 2 |
| 2019 | Learning Hyperedge Replacement Grammars for Graph GenerationabstractThe discovery and analysis of network patterns are central to the scientific enterprise. In the present work, we developed and evaluated a new approach that learns the building blocks of graphs that can be used to understand and generate new realistic graphs. Our key insight is that a graph's clique tree encodes robust and precise information. We show that a Hyperedge Replacement Grammar (HRG) can be extracted from the clique tree, and we develop a fixed-size graph generation algorithm that can be used to produce new graphs of a specified size. In experiments on large real-world graphs, we show that graphs generated from the HRG approach exhibit a diverse range of properties that are similar to those found in the original networks. In addition to graph properties like degree or eigenvector centrality, what a graph "looks like" ultimately depends on small details in local graph substructures that are difficult to define at a global level. We show that the HRG model can also preserve these local substructures when generating new graphs. Salvador Aguiñaga, David Chiang 0001, Tim Weninger |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Composing Finite State Transducers on GPUsabstractWeighted finite state transducers (FSTs) are frequently used in language processing to handle tasks such as part-of-speech tagging and speech recognition.There has been previous work using multiple CPU cores to accelerate finite state algorithms, but limited attention has been given to parallel graphics processing unit (GPU) implementations.In this paper, we introduce the first (to our knowledge) GPU implementation of the FST composition operation, and we also discuss the optimizations used to achieve the best performance on this architecture.We show that our approach obtains speedups of up to 6× over our serial implementation and 4.5× over OpenFST. Arturo Argueta, David Chiang 0001 |
ACL (1) | 2 |
| 2018 | Part-of-Speech Tagging on an Endangered Language: a Parallel Griko-Italian ResourceabstractMost work on part-of-speech (POS) tagging is focused on high resource languages, or examines low-resource and active learning settings through simulated studies. We evaluate POS tagging techniques on an actual endangered language, Griko. We present a resource that contains 114 narratives in Griko, along with sentence-level translations in Italian, and provides gold annotations for the test set. Based on a previously collected small corpus, we investigate several traditional methods, as well as methods that take advantage of monolingual data or project cross-lingual POS tags. We show that the combination of a semi-supervised method with cross-lingual transfer is more appropriate for this extremely challenging setting, with the best tagger achieving an accuracy of 72.9%. With an applied active learning scheme, which we use to collect sentence-level annotations over the test set, we achieve improvements of more than 21 percentage points. Antonios Anastasopoulos, Marika Lekakou, Josep Quer, Eleni Zimianiti, Justin DeBenedetto, David Chiang 0001 |
COLING | 6 |
| 2018 | Synchronous Hyperedge Replacement Graph Grammars
Corey Pennycuff, Satyaki Sikdar, Catalina Vajiac, David Chiang 0001, Tim Weninger |
ICGT | 4 |
| 2018 | Leveraging Translations for Speech Transcription in Low-resource SettingsabstractRecently proposed data collection frameworks for endangered language documentation aim not only to collect speech in the language of interest, but also to collect translations into a high-resource language that will render the collected resource interpretable. We focus on this scenario and explore whether we can improve transcription quality under these extremely low-resource settings with the assistance of text translations. We present a neural multi-source model and evaluate several variations of it on three low-resource datasets. We find that our multi-source model with shared attention outperforms the baselines, reducing transcription character error rate by up to 12.3%. Antonios Anastasopoulos, David Chiang 0001 |
INTERSPEECH | 2 |
| 2018 | Tied Multitask Learning for Neural Speech TranslationabstractAntonios Anastasopoulos, David Chiang. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Antonios Anastasopoulos, David Chiang 0001 |
NAACL-HLT | 2 |
| 2018 | Combining Character and Word Information in Neural Machine Translation Using a Multi-Level AttentionabstractHuadong Chen, Shujian Huang, David Chiang, Xinyu Dai, Jiajun Chen. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Huadong Chen, Shujian Huang, David Chiang 0001, Xinyu Dai, Jiajun Chen 0001 |
NAACL-HLT | 3 |
| 2018 | Improving Lexical Choice in Neural Machine TranslationabstractWe explore two solutions to the problem of mistranslating rare words in neural machine translation.First, we argue that the standard output layer, which computes the inner product of a vector representing the context with all possible output word embeddings, rewards frequent words disproportionately, and we propose to fix the norms of both vectors to a constant value.Second, we integrate a simple lexical module which is jointly trained with the rest of the model.We evaluate our approaches on eight language pairs with data sizes ranging from 100k to 8M words, and achieve improvements of up to +4.3 BLEU, surpassing phrasebased translation in nearly all settings.1 Toan Q. Nguyen, David Chiang 0001 |
NAACL-HLT | 2 |
| 2018 | Algorithms and Training for Weighted Multiset Automata and Regular Expressions
Justin DeBenedetto, David Chiang 0001 |
CIAA | 2 |
| 2018 | Weighted DAG Automata for Semantic GraphsabstractGraphs have a variety of uses in natural language processing, particularly as representations of linguistic meaning. A deficit in this area of research is a formal framework for creating, combining, and using models involving graphs that parallels the frameworks of finite automata for strings and finite tree automata for trees. A possible starting point for such a framework is the formalism of directed acyclic graph (DAG) automata, defined by Kamimura and Slutzki and extended by Quernheim and Knight. In this article, we study the latter in depth, demonstrating several new results, including a practical recognition algorithm that can be used for inference and learning with models defined on DAG automata. We also propose an extension to graphs with unbounded node degree and show that our results carry over to the extended formalism. David Chiang 0001, Frank Drewes, Daniel Gildea, Adam Lopez, Giorgio Satta |
Comput. Linguistics | 1 |
| 2018 | Incident-Driven Machine Translation and Name Tagging for Low-resource Languages
Ulf Hermjakob, Daniel Marcu, Jonathan May, Sabrina J. Mielke, Nima Pourdamghani, Michael Pust, Kevin Knight, Tomer Levinboim, Kenton Murray, David Chiang 0001, Boliang Zhang, Xiaoman Pan, Di Lu 0003, Heng Ji 0001 |
Mach. Transl. | 12 |
| 2017 | Improved Neural Machine Translation with a Syntax-Aware Encoder and DecoderabstractMost neural machine translation (NMT) models are based on the sequential encoder-decoder framework, which makes no use of syntactic information.In this paper, we improve this model by explicitly incorporating source-side syntactic trees.More specifically, we propose (1) a bidirectional tree encoder which learns both sequential and tree structured representations; (2) a tree-coverage model that lets the attention depend on the source-side syntax.Experiments on Chinese-English translation demonstrate that our proposed models outperform the sequential attentional model as well as a stronger baseline with a bottom-up tree encoder and word coverage.1 Huadong Chen, Shujian Huang, David Chiang 0001, Jiajun Chen 0001 |
ACL (1) | 3 |
| 2017 | Top-Rank Enhanced Listwise Optimization for Statistical Machine TranslationabstractPairwise ranking methods are the basis of many widely used discriminative training approaches for structure prediction problems in natural language processing (NLP).Decomposing the problem of ranking hypotheses into pairwise comparisons enables simple and efficient solutions.However, neglecting the global ordering of the hypothesis list may hinder learning.We propose a listwise learning framework for structure prediction problems such as machine translation.Our framework directly models the entire translation list's ordering to learn parameters which may better fit the given listwise samples.Furthermore, we propose top-rank enhanced loss functions, which are more sensitive to ranking errors at higher positions.Experiments on a large-scale Chinese-English translation task show that both our listwise learning framework and top-rank enhanced listwise losses lead to significant improvements in translation quality. Huadong Chen, Shujian Huang, David Chiang 0001, Xinyu Dai, Jiajun Chen 0001 |
CoNLL | 3 |
| 2017 | Decoding with Finite-State Transducers on GPUsabstractWeighted finite automata and transducers (including hidden Markov models and conditional random fields) are widely used in natural language processing (NLP) to perform tasks such as morphological analysis, part-of-speech tagging, chunking, named entity recognition, speech recognition, and others.Parallelizing finite state algorithms on graphics processing units (GPUs) would benefit many areas of NLP.Although researchers have implemented GPU versions of basic graph algorithms, limited previous work, to our knowledge, has been done on GPU algorithms for weighted finite automata.We introduce a GPU implementation of the Viterbi and forward-backward algorithm, achieving decoding speedups of up to 5.2x over our serial implementation running on different computer architectures and 6093x over OpenFST. Arturo Argueta, David Chiang 0001 |
EACL (1) | 2 |
| 2016 | Growing Graphs from Hyperedge Replacement Graph GrammarsabstractDiscovering the underlying structures present in large real world graphs is a fundamental scientific problem. In this paper we show that a graph's clique tree can be used to extract a hyperedge replacement grammar. If we store an ordering from the extraction process, the extracted graph grammar is guaranteed to generate an isomorphic copy of the original graph. Or, a stochastic application of the graph grammar rules can be used to quickly create random graphs. In experiments on large real world networks, we show that random graphs, generated from extracted graph grammars, exhibit a wide range of properties that are very similar to the original graphs. In addition to graph properties like degree or eigenvector centrality, what a graph ``looks like'' ultimately depends on small details in local graph substructures that are difficult to define at a global level. We show that our generative graph model is able to preserve these local substructures when generating new graphs and performs well on new and difficult tests of model robustness. Salvador Aguiñaga, Rodrigo Palácios, David Chiang 0001, Tim Weninger |
CIKM | 3 |
| 2016 | An Unsupervised Probability Model for Speech-to-Translation Alignment of Low-Resource LanguagesabstractFor many low-resource languages, spoken language resources are more likely to be annotated with translations than with transcriptions. Translated speech data is potentially valuable for documenting endangered languages or for training speech translation systems. A first step towards making use of such data would be to automatically align spoken words with their translations. We present a model that combines Dyer et al.'s reparameterization of IBM Model 2 (fast-align) and k-means clustering using Dynamic Time Warping as a distance metric. The two components are trained jointly using expectation-maximization. In an extremely low-resource scenario, our model performs significantly better than both a neural model and a strong baseline. Antonios Anastasopoulos, David Chiang 0001, Long Duong |
EMNLP | 2 |
| 2016 | An Attentional Model for Speech Translation Without TranscriptionabstractLong Duong, Antonios Anastasopoulos, David Chiang, Steven Bird, Trevor Cohn. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Long Duong, Antonios Anastasopoulos, David Chiang 0001, Steven Bird, Trevor Cohn |
HLT-NAACL | 3 |
| 2015 | Supervised Phrase Table Triangulation with Neural Word Embeddings for Low-Resource LanguagesabstractIn this paper, we develop a supervised learning technique that improves noisy phrase translation scores obtained by phrase table triangulation.In particular, we extract word translation distributions from small amounts of source-target bilingual data (a dictionary or a parallel corpus) with which we learn to assign better scores to translation candidates obtained by triangulation.Our method is able to gain improvement in translation quality on two tasks: (1) On Malagasy-to-French translation via English, we use only 1k dictionary entries to gain +0.5 Bleu over triangulation.(2) On Spanish-to-French via English we use only 4k sentence pairs to gain +0.7 Bleu over triangulation interpolated with a phrase table extracted from the same 4k sentence pairs. Tomer Levinboim, David Chiang 0001 |
EMNLP | 2 |
| 2015 | Auto-Sizing Neural Networks: With Applications to n-gram Language ModelsabstractNeural networks have been shown to improve performance across a range of natural-language tasks.However, designing and training them can be complicated.Frequently, researchers resort to repeated experimentation to pick optimal settings.In this paper, we address the issue of choosing the correct number of units in hidden layers.We introduce a method for automatically adjusting network size by pruning out hidden units through ℓ ∞,1 and ℓ 2,1 regularization.We apply this method to language modeling and demonstrate its ability to correctly choose the number of hidden units while maintaining perplexity.We also include these models in a machine translation decoder and show that these smaller neural models maintain the significant improvements of their unpruned versions. Kenton Murray, David Chiang 0001 |
EMNLP | 2 |
| 2015 | Multi-Task Word Alignment Triangulation for Low-Resource LanguagesabstractWe present a multi-task learning approach that jointly trains three word alignment models over disjoint bitexts of three languages: source, target and pivot.Our approach builds upon model triangulation, following Wang et al., which approximates a source-target model by combining source-pivot and pivot-target models.We develop a MAP-EM algorithm that uses triangulation as a prior, and show how to extend it to a multi-task setting.On a low-resource Czech-English corpus, using French as the pivot, our multi-task learning approach more than doubles the gains in both Fand Bleu scores compared to the interpolation approach of Wang et al.Further experiments reveal that the choice of pivot language does not significantly affect performance. Tomer Levinboim, David Chiang 0001 |
HLT-NAACL | 2 |
| 2015 | Model Invertibility Regularization: Sequence Alignment With or Without Parallel DataabstractWe present Model Invertibility Regularization (MIR), a method that jointly trains two directional sequence alignment models, one in each direction, and takes into account the invertibility of the alignment task.By coupling the two models through their parameters (as opposed to through their inferences, as in Liang et al.'s Alignment by Agreement (ABA), and Ganchev et al.'s Posterior Regularization (PostCAT)), our method seamlessly extends to all IBMstyle word alignment models as well as to alignment without parallel data.Our proposed algorithm is mathematically sound and inherits convergence guarantees from EM.We evaluate MIR on two tasks: (1) On word alignment, applying MIR on fertility based models we attain higher F-scores than ABA and PostCAT.(2) On Japanese-to-English backtransliteration without parallel data, applied to the decipherment model of Ravi and Knight, MIR learns sparser models that close the gap in whole-name error rate by 33% relative to a model trained on parallel data, and further, beats a previous approach by Mylonakis et al. Tomer Levinboim, Ashish Vaswani, David Chiang 0001 |
HLT-NAACL | 3 |
| 2014 | Kneser-Ney Smoothing on Expected CountsabstractWidely used in speech and language processing, Kneser-Ney (KN) smoothing has consistently been shown to be one of the best-performing smoothing methods.However, KN smoothing assumes integer counts, limiting its potential uses-for example, inside Expectation-Maximization.In this paper, we propose a generalization of KN smoothing that operates on fractional counts, or, more precisely, on distributions over counts.We rederive all the steps of KN smoothing to operate on count distributions instead of integral counts, and apply it to two tasks where KN smoothing was not applicable before: one in language model adaptation, and the other in word alignment.In both cases, our method improves performance significantly. David Chiang 0001 |
ACL (1) | 2 |
| 2014 | Improving Word Alignment using Word SimilarityabstractWe show that semantic relationships can be used to improve word alignment, in addition to the lexical and syntactic features that are typically used.In this paper, we present a method based on a neural network to automatically derive word similarity from monolingual data.We present an extension to word alignment models that exploits word similarity.Our experiments, in both large-scale and resourcelimited settings, show improvements in word alignment tasks as well as translation tasks. Theerawat Songyot, David Chiang 0001 |
EMNLP | 2 |
| 2013 | Parsing Graphs with Hyperedge Replacement Grammars
David Chiang 0001, Jacob Andreas, Daniel Bauer 0002, Karl Moritz Hermann, Bevan K. Jones, Kevin Knight |
ACL (1) | 1 |
| 2013 | Decoding with Large-Scale Neural Language Models Improves TranslationabstractWe explore the application of neural language models to machine translation.We develop a new model that combines the neural probabilistic language model of Bengio et al., rectified linear units, and noise-contrastive estimation, and we incorporate it into a machine translation system both by reranking k-best lists and by direct integration into the decoder.Our large-scale, large-vocabulary experiments across four language pairs show that our neural language model improves translation quality by up to 1.1 Bleu. Ashish Vaswani, Yinggong Zhao, Victoria Fossum, David Chiang 0001 |
EMNLP | 4 |
| 2012 | Smaller Alignment Models for Better Translations: Unsupervised Word Alignment with the l0-norm
Ashish Vaswani, Liang Huang 0001, David Chiang 0001 |
ACL (1) | 3 |
| 2012 | Hope and Fear for Discriminative Training of Statistical Translation Models
David Chiang 0001 |
J. Mach. Learn. Res. | 1 |
| 2012 | Soft syntactic constraints for Arabic-English hierarchical phrase-based translation
Yuval Marton, David Chiang 0001, Philip Resnik |
Mach. Transl. | 2 |
| 2011 | Rule Markov Models for Fast Tree-to-String Translation
Ashish Vaswani, Haitao Mi, Liang Huang 0001, David Chiang 0001 |
ACL | 4 |
| 2010 | Learning to Translate with Source and Target Syntax
David Chiang 0001 |
ACL | 1 |
| 2010 | Fast, Greedy Model Minimization for Unsupervised Tagging
Sujith Ravi, Ashish Vaswani, Kevin Knight, David Chiang 0001 |
COLING | 4 |
| 2010 | Bayesian Inference for Finite-State Transducers
David Chiang 0001, Jonathan Graehl, Kevin Knight, Adam Pauls, Sujith Ravi |
HLT-NAACL | 1 |
| 2010 | Unsupervised Syntactic Alignment with Inversion Transduction Grammars
Adam Pauls, Daniel Klein 0001, David Chiang 0001, Kevin Knight |
HLT-NAACL | 3 |
| 2009 | Fast Consensus Decoding over Translation Forests
John DeNero, David Chiang 0001, Kevin Knight |
ACL/IJCNLP | 2 |
| 2009 | 11,001 New Features for Statistical Machine Translation
David Chiang 0001, Kevin Knight, Wei Wang 0006 |
HLT-NAACL | 1 |
| 2009 | Introduction to the Special Issue on Machine Translation of Asian LanguagesabstractNo abstract available. David Chiang 0001, Philipp Koehn |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2008 | Extracting Synchronous Grammar Rules From Word-Level Alignments in Linear Time
Hao Zhang 0010, Daniel Gildea, David Chiang 0001 |
COLING | 3 |
| 2008 | Decomposability of Translation Metrics for Improved Evaluation and Efficient Algorithms
David Chiang 0001, Steve DeNeefe, Yee Seng Chan, Hwee Tou Ng |
EMNLP | 1 |
| 2008 | Online Large-Margin Training of Syntactic and Structural Translation Features
David Chiang 0001, Yuval Marton, Philip Resnik |
EMNLP | 1 |
| 2007 | Word Sense Disambiguation Improves Statistical Machine Translation
Yee Seng Chan, Hwee Tou Ng, David Chiang 0001 |
ACL | 3 |
| 2007 | Forest Rescoring: Faster Decoding with Integrated Language Models
Liang Huang 0001, David Chiang 0001 |
ACL | 2 |
| 2007 | Hierarchical Phrase-Based TranslationabstractWe present a statistical machine translation model that uses hierarchical phrases—phrases that contain subphrases. The model is formally a synchronous context-free grammar but is learned from a parallel text without any syntactic annotations. Thus it can be seen as combining fundamental ideas from both syntax-based translation and phrase-based translation. We describe our system's training and decoding methods in detail, and evaluate it for translation speed and translation accuracy. Using BLEU as a metric of translation accuracy, we find that our system performs significantly better than the Alignment Template System, a state-of-the-art phrase-based system. David Chiang 0001 |
Comput. Linguistics | 1 |
| 2006 | Parsing Arabic Dialects
David Chiang 0001, Mona T. Diab, Nizar Habash, Owen Rambow, Safiullah Shareef |
EACL | 1 |
| 2005 | A Hierarchical Phrase-Based Model for Statistical Machine TranslationabstractWe present a statistical phrase-based translation model that uses hierarchical phrases---phrases that contain subphrases. The model is formally a synchronous context-free grammar but is learned from a bitext without any syntactic information. Thus it can be seen as a shift to the formal machinery of syntax-based translation systems without any linguistic commitment. In our experiments using BLEU as a metric, the hierarchical phrase-based model achieves a relative improvement of 7.5% over Pharaoh, a state-of-the-art phrase-based system. David Chiang 0001 |
ACL | 1 |
| 2002 | Recovering Latent Information in Treebanks
David Chiang 0001, Dan Bikel |
COLING | 1 |
| 2001 | Constraints on Strong Generative PowerabstractWe consider the question "How much strong generative power can be squeezed out of a formal system without increasing its weak generative power?" and propose some theoretical and practical constraints on this problem. We then introduce a formalism which, under these constraints, maximally squeezes strong generative power out of context-free grammar. Finally, we generalize this result to formalisms beyond CFG. David Chiang 0001 |
ACL | 1 |
| 2000 | Statistical Parsing with an Automatically-Extracted Tree Adjoining GrammarabstractWe discuss the advantages of lexicalized tree-adjoining grammar as an alternative to lexicalized PCFG for statistical parsing, describing the induction of a probabilistic LTAG model from the Penn Treebank and evaluating its parsing performance. We find that this induction method is an improvement over the EM-based method of (Hwa, 1998), and that the induced model yields results comparable to lexicalized PCFG. David Chiang 0001 |
ACL | 1 |
| 2000 | Multi-Component TAG and Notions of Formal PowerabstractThis paper presents a restricted version of Set-Local Multi-Component TAGs (Weir, 1988) which retains the strong generative capacity of Tree-Local Multi-Component TAG (i.e. produces the same derived structures) but has a greater derivational generative capacity (i.e. can derive those structures in more ways). This formalism is then applied as a framework for integrating dependency and constituency based linguistic representations. William Schuler, David Chiang 0001, Mark Dras |
ACL | 2 |