James Cross 0003

dblp:90/4769-3 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
10since 2021 · last 2023
0000-0001-8042-1099ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 10 since 2021
YearPublicationVenuePosition
2023 Efficiently Upgrading Multilingual Machine Translation Models to Support More Languages
abstract
With multilingual machine translation (MMT) models continuing to grow in size and number of supported languages, it is natural to reuse and upgrade existing models to save computation as data becomes available in more languages.However, adding new languages requires updating the vocabulary, which complicates the reuse of embeddings.The question of how to reuse existing models while also making architectural changes to provide capacity for both old and new languages has also not been closely studied.In this work, we introduce three techniques that help speed up effective learning of the new languages and alleviate catastrophic forgetting despite vocabulary and architecture mismatches.Our results show that by (1) carefully initializing the network, (2) applying learning rate scaling, and (3) performing data up-sampling, it is possible to exceed the performance of a same-sized baseline model with 30% computation and recover the performance of a larger model trained from scratch with over 50% reduction in computation.Furthermore, our analysis reveals that the introduced techniques help learn the new directions more effectively and alleviate catastrophic forgetting at the same time.We hope our work will guide research into more efficient approaches to growing languages for these MMT models and ultimately maximize the reuse of existing models.* Work done during an internship at Meta AI.
Simeng Sun, Maha Elbayad, Anna Y. Sun, James Cross 0003
EACL4
2022 Alternative Input Signals Ease Transfer in Multilingual Machine Translation
abstract
Simeng Sun, Angela Fan, James Cross, Vishrav Chaudhary, Chau Tran, Philipp Koehn, Francisco Guzmán. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Simeng Sun, Angela Fan, James Cross 0003, Vishrav Chaudhary, Chau Tran, Philipp Koehn, Francisco Guzmán
ACL (1)3
2022 Multilingual Machine Translation with Hyper-Adapters
abstract
Multilingual machine translation suffers from negative interference across languages.A common solution is to relax parameter sharing with language-specific modules like adapters.However, adapters of related languages are unable to transfer information, and their total number of parameters becomes prohibitively expensive as the number of languages grows.In this work, we overcome these drawbacks using hyper-adapters-hyper-networks that generate adapters from language and layer embeddings.While past work had poor results when scaling hyper-networks, we propose a rescaling fix that significantly improves convergence and enables training larger hyper-networks.We find that hyper-adapters are more parameter efficient than regular adapters, reaching the same performance with up to 12 times less parameters.When using the same number of parameters and FLOPS, our approach consistently outperforms regular adapters.Also, hyper-adapters converge faster than alternative approaches and scale better than regular dense networks.Our analysis shows that hyperadapters learn to encode language relatedness, enabling positive transfer across languages.
Christos Baziotis, Mikel Artetxe, James Cross 0003, Shruti Bhosale
EMNLP3
2022 Tricks for Training Sparse Translation Models
abstract
Dheeru Dua, Shruti Bhosale, Vedanuj Goswami, James Cross, Mike Lewis, Angela Fan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Dheeru Dua, Shruti Bhosale, Vedanuj Goswami, James Cross 0003, Mike Lewis, Angela Fan
NAACL-HLT4
2022 Lifting the Curse of Multilinguality by Pre-training Modular Transformers
abstract
Jonas Pfeiffer, Naman Goyal, Xi Lin, Xian Li, James Cross, Sebastian Riedel, Mikel Artetxe. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Jonas Pfeiffer, Naman Goyal 0001, Xi Victoria Lin, Xian Li 0003, James Cross 0003, Sebastian Riedel 0001, Mikel Artetxe
NAACL-HLT5
2021 Improving Zero-Shot Translation by Disentangling Positional Information
abstract
Danni Liu, Jan Niehues, James Cross, Francisco Guzmán, Xian Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jan Niehues, James Cross 0003, Francisco Guzmán, Xian Li 0003
ACL/IJCNLP (1)3
2021 Multilingual Neural Machine Translation with Deep Encoder and Multiple Shallow Decoders
abstract
Xiang Kong, Adithya Renduchintala, James Cross, Yuqing Tang, Jiatao Gu, Xian Li. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Xiang Kong, Adithya Renduchintala, James Cross 0003, Jiatao Gu, Xian Li 0003
EACL3
2021 XLEnt: Mining a Large Cross-lingual Entity Dataset with Lexical-Semantic-Phonetic Word Alignment
abstract
Cross-lingual named-entity lexica are an important resource to multilingual NLP tasks such as machine translation and cross-lingual wikification.While knowledge bases contain a large number of entities in high-resource languages such as English and French, corresponding entities for lower-resource languages are often missing.To address this, we propose Lexical-Semantic-Phonetic Align (LSP-Align), a technique to automatically mine cross-lingual entity lexica from mined web data.We demonstrate LSP-Align outperforms baselines at extracting cross-lingual entity pairs and mine 164 million entity pairs from 120 different languages aligned with English.We release these cross-lingual entity pairs along with the massively multilingual tagged named entity corpus as a resource to the NLP community.
Ahmed El-Kishky, Adithya Renduchintala, James Cross 0003, Francisco Guzmán, Philipp Koehn
EMNLP (1)3
2021 Classification-based Quality Estimation: Small and Efficient Models for Real-world Applications
abstract
Sentence-level Quality Estimation (QE) of machine translation is traditionally formulated as a regression task, and the performance of QE models is typically measured by Pearson correlation with human labels.Recent QE models have achieved previously-unseen levels of correlation with human judgments, but they rely on large multilingual contextualized language models that are computationally expensive and thus infeasible for many real-world applications.In this work, we evaluate several model compression techniques for QE and find that, despite their popularity in other NLP tasks, they lead to poor performance in this regression setting.We observe that a full model parameterization is required to achieve SoTA results in a regression task.However, we argue that the level of expressiveness of a model in a continuous range is unnecessary given the downstream applications of QE, and show that reframing QE as a classification problem and evaluating QE models using classification metrics would better reflect their actual performance in real-world applications.
Ahmed El-Kishky, Vishrav Chaudhary, James Cross 0003, Lucia Specia, Francisco Guzmán
EMNLP (1)4
2021 Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine Translation
Jungo Kasai, Nikolaos Pappas 0002, Hao Peng 0009, James Cross 0003, Noah A. Smith
ICLR4
2020 Monotonic Multihead Attention
Xutai Ma, Juan Pino 0001, James Cross 0003, Liezl Puzon, Jiatao Gu
ICLR3
2020 Non-autoregressive Machine Translation with Disentangled Context Transformer
abstract
State-of-the-art neural machine translation models generate a translation from left to right and every step is conditioned on the previously generated tokens. The sequential nature of this generation process causes fundamental latency in inference since we cannot generate multiple tokens in each sentence in parallel. We propose an attention-masking based model, called Disentangled Context (DisCo) transformer, that simultaneously generates all tokens given different contexts. The DisCo transformer is trained to predict every output token given an arbitrary subset of the other reference tokens. We also develop the parallel easy-first inference algorithm, which iteratively refines every token in parallel and reduces the number of required iterations. Our extensive experiments on 7 translation directions with varying data sizes demonstrate that our model achieves competitive, if not better, performance compared to the state of the art in non-autoregressive machine translation while significantly reducing decoding time on average.
Jungo Kasai, James Cross 0003, Marjan Ghazvininejad, Jiatao Gu
ICML2
2016 Span-Based Constituency Parsing with a Structure-Label System and Provably Optimal Dynamic Oracles
abstract
Parsing accuracy using efficient greedy transition systems has improved dramatically in recent years thanks to neural networks.Despite striking results in dependency parsing, however, neural models have not surpassed stateof-the-art approaches in constituency parsing.To remedy this, we introduce a new shiftreduce system whose stack contains merely sentence spans, represented by a bare minimum of LSTM features.We also design the first provably optimal dynamic oracle for constituency parsing, which runs in amortized O(1) time, compared to O(n 3 ) oracles for standard dependency parsing.Training with this oracle, we achieve the best F 1 scores on both English and French of any parser that does not use reranking or external data.
James Cross 0003, Liang Huang 0001
EMNLP1
2013 Optimal Incremental Parsing via Best-First Dynamic Programming
abstract
We present the first provably optimal polynomial time dynamic programming (DP) algorithm for best-first shift-reduce parsing, which applies the DP idea of Huang and Sagae (2010) to the best-first parser of Sagae and Lavie (2006) in a non-trivial way, reducing the complexity of the latter from exponential to polynomial.We prove the correctness of our algorithm rigorously.Experiments confirm that DP leads to a significant speedup on a probablistic best-first shift-reduce parser, and makes exact search under such a model tractable for the first time.
Kai Zhao 0003, James Cross 0003, Liang Huang 0001
EMNLP2