VLDB 2026 Research / reviewers in the wild / expert
James Cross 0003
dblp:90/4769-3
· DBLP profile ↗
14ranked-venue papers
1as first author
10since 2021 · last 2023
0000-0001-8042-1099ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 1 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Efficiently Upgrading Multilingual Machine Translation Models to Support More LanguagesabstractWith multilingual machine translation (MMT) models continuing to grow in size and number of supported languages, it is natural to reuse and upgrade existing models to save computation as data becomes available in more languages.However, adding new languages requires updating the vocabulary, which complicates the reuse of embeddings.The question of how to reuse existing models while also making architectural changes to provide capacity for both old and new languages has also not been closely studied.In this work, we introduce three techniques that help speed up effective learning of the new languages and alleviate catastrophic forgetting despite vocabulary and architecture mismatches.Our results show that by (1) carefully initializing the network, (2) applying learning rate scaling, and (3) performing data up-sampling, it is possible to exceed the performance of a same-sized baseline model with 30% computation and recover the performance of a larger model trained from scratch with over 50% reduction in computation.Furthermore, our analysis reveals that the introduced techniques help learn the new directions more effectively and alleviate catastrophic forgetting at the same time.We hope our work will guide research into more efficient approaches to growing languages for these MMT models and ultimately maximize the reuse of existing models.* Work done during an internship at Meta AI. Simeng Sun, Maha Elbayad, Anna Y. Sun, James Cross 0003 |
EACL | 4 |
| 2022 | Alternative Input Signals Ease Transfer in Multilingual Machine TranslationabstractSimeng Sun, Angela Fan, James Cross, Vishrav Chaudhary, Chau Tran, Philipp Koehn, Francisco Guzmán. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Simeng Sun, Angela Fan, James Cross 0003, Vishrav Chaudhary, Chau Tran, Philipp Koehn, Francisco Guzmán |
ACL (1) | 3 |
| 2022 | Multilingual Machine Translation with Hyper-AdaptersabstractMultilingual machine translation suffers from negative interference across languages.A common solution is to relax parameter sharing with language-specific modules like adapters.However, adapters of related languages are unable to transfer information, and their total number of parameters becomes prohibitively expensive as the number of languages grows.In this work, we overcome these drawbacks using hyper-adapters-hyper-networks that generate adapters from language and layer embeddings.While past work had poor results when scaling hyper-networks, we propose a rescaling fix that significantly improves convergence and enables training larger hyper-networks.We find that hyper-adapters are more parameter efficient than regular adapters, reaching the same performance with up to 12 times less parameters.When using the same number of parameters and FLOPS, our approach consistently outperforms regular adapters.Also, hyper-adapters converge faster than alternative approaches and scale better than regular dense networks.Our analysis shows that hyperadapters learn to encode language relatedness, enabling positive transfer across languages. Christos Baziotis, Mikel Artetxe, James Cross 0003, Shruti Bhosale |
EMNLP | 3 |
| 2022 | Tricks for Training Sparse Translation ModelsabstractDheeru Dua, Shruti Bhosale, Vedanuj Goswami, James Cross, Mike Lewis, Angela Fan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Dheeru Dua, Shruti Bhosale, Vedanuj Goswami, James Cross 0003, Mike Lewis, Angela Fan |
NAACL-HLT | 4 |
| 2022 | Lifting the Curse of Multilinguality by Pre-training Modular TransformersabstractJonas Pfeiffer, Naman Goyal, Xi Lin, Xian Li, James Cross, Sebastian Riedel, Mikel Artetxe. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jonas Pfeiffer, Naman Goyal 0001, Xi Victoria Lin, Xian Li 0003, James Cross 0003, Sebastian Riedel 0001, Mikel Artetxe |
NAACL-HLT | 5 |
| 2021 | Improving Zero-Shot Translation by Disentangling Positional InformationabstractDanni Liu, Jan Niehues, James Cross, Francisco Guzmán, Xian Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jan Niehues, James Cross 0003, Francisco Guzmán, Xian Li 0003 |
ACL/IJCNLP (1) | 3 |
| 2021 | Multilingual Neural Machine Translation with Deep Encoder and Multiple Shallow DecodersabstractXiang Kong, Adithya Renduchintala, James Cross, Yuqing Tang, Jiatao Gu, Xian Li. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Xiang Kong, Adithya Renduchintala, James Cross 0003, Jiatao Gu, Xian Li 0003 |
EACL | 3 |
| 2021 | XLEnt: Mining a Large Cross-lingual Entity Dataset with Lexical-Semantic-Phonetic Word AlignmentabstractCross-lingual named-entity lexica are an important resource to multilingual NLP tasks such as machine translation and cross-lingual wikification.While knowledge bases contain a large number of entities in high-resource languages such as English and French, corresponding entities for lower-resource languages are often missing.To address this, we propose Lexical-Semantic-Phonetic Align (LSP-Align), a technique to automatically mine cross-lingual entity lexica from mined web data.We demonstrate LSP-Align outperforms baselines at extracting cross-lingual entity pairs and mine 164 million entity pairs from 120 different languages aligned with English.We release these cross-lingual entity pairs along with the massively multilingual tagged named entity corpus as a resource to the NLP community. Ahmed El-Kishky, Adithya Renduchintala, James Cross 0003, Francisco Guzmán, Philipp Koehn |
EMNLP (1) | 3 |
| 2021 | Classification-based Quality Estimation: Small and Efficient Models for Real-world ApplicationsabstractSentence-level Quality Estimation (QE) of machine translation is traditionally formulated as a regression task, and the performance of QE models is typically measured by Pearson correlation with human labels.Recent QE models have achieved previously-unseen levels of correlation with human judgments, but they rely on large multilingual contextualized language models that are computationally expensive and thus infeasible for many real-world applications.In this work, we evaluate several model compression techniques for QE and find that, despite their popularity in other NLP tasks, they lead to poor performance in this regression setting.We observe that a full model parameterization is required to achieve SoTA results in a regression task.However, we argue that the level of expressiveness of a model in a continuous range is unnecessary given the downstream applications of QE, and show that reframing QE as a classification problem and evaluating QE models using classification metrics would better reflect their actual performance in real-world applications. Ahmed El-Kishky, Vishrav Chaudhary, James Cross 0003, Lucia Specia, Francisco Guzmán |
EMNLP (1) | 4 |
| 2021 | Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine Translation
Jungo Kasai, Nikolaos Pappas 0002, Hao Peng 0009, James Cross 0003, Noah A. Smith |
ICLR | 4 |
| 2020 | Monotonic Multihead Attention
Xutai Ma, Juan Pino 0001, James Cross 0003, Liezl Puzon, Jiatao Gu |
ICLR | 3 |
| 2020 | Non-autoregressive Machine Translation with Disentangled Context TransformerabstractState-of-the-art neural machine translation models generate a translation from left to right and every step is conditioned on the previously generated tokens. The sequential nature of this generation process causes fundamental latency in inference since we cannot generate multiple tokens in each sentence in parallel. We propose an attention-masking based model, called Disentangled Context (DisCo) transformer, that simultaneously generates all tokens given different contexts. The DisCo transformer is trained to predict every output token given an arbitrary subset of the other reference tokens. We also develop the parallel easy-first inference algorithm, which iteratively refines every token in parallel and reduces the number of required iterations. Our extensive experiments on 7 translation directions with varying data sizes demonstrate that our model achieves competitive, if not better, performance compared to the state of the art in non-autoregressive machine translation while significantly reducing decoding time on average. Jungo Kasai, James Cross 0003, Marjan Ghazvininejad, Jiatao Gu |
ICML | 2 |
| 2016 | Span-Based Constituency Parsing with a Structure-Label System and Provably Optimal Dynamic OraclesabstractParsing accuracy using efficient greedy transition systems has improved dramatically in recent years thanks to neural networks.Despite striking results in dependency parsing, however, neural models have not surpassed stateof-the-art approaches in constituency parsing.To remedy this, we introduce a new shiftreduce system whose stack contains merely sentence spans, represented by a bare minimum of LSTM features.We also design the first provably optimal dynamic oracle for constituency parsing, which runs in amortized O(1) time, compared to O(n 3 ) oracles for standard dependency parsing.Training with this oracle, we achieve the best F 1 scores on both English and French of any parser that does not use reranking or external data. James Cross 0003, Liang Huang 0001 |
EMNLP | 1 |
| 2013 | Optimal Incremental Parsing via Best-First Dynamic ProgrammingabstractWe present the first provably optimal polynomial time dynamic programming (DP) algorithm for best-first shift-reduce parsing, which applies the DP idea of Huang and Sagae (2010) to the best-first parser of Sagae and Lavie (2006) in a non-trivial way, reducing the complexity of the latter from exponential to polynomial.We prove the correctness of our algorithm rigorously.Experiments confirm that DP leads to a significant speedup on a probablistic best-first shift-reduce parser, and makes exact search under such a model tractable for the first time. Kai Zhao 0003, James Cross 0003, Liang Huang 0001 |
EMNLP | 2 |