Biao Zhang 0006

dblp:83/3266-6 · DBLP profile ↗
← Back
21ranked-venue papers
15as first author
16since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 15 first-author · 16 since 2021
YearPublicationVenuePosition
2025 YouTube-SL-25: A Large-Scale, Open-Domain Multilingual Sign Language Parallel Corpus
abstract
Even for better-studied sign languages like American Sign Language (ASL), data is the bottleneck for machine learning research. The situation is worse yet for the many other sign languages used by Deaf/Hard of Hearing communities around the world. In this paper, we present YouTube-SL-25, a large-scale, open-domain multilingual corpus of sign language videos with seemingly well-aligned captions drawn from YouTube. With >3000 hours of videos across >25 sign languages, YouTube-SL-25 is a) >3x the size of YouTube-ASL, b) the largest parallel sign language dataset to date, and c) the first or largest parallel dataset for many of its component languages. We provide baselines for sign-to-text tasks using a unified multilingual multitask model based on T5 and report scores on benchmarks across 4 sign languages. The results demonstrate that multilingual transfer benefits both higher- and lower-resource sign languages within YouTube-SL-25.
Garrett Tanzer, Biao Zhang 0006
ICLR2
2024 When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
abstract
While large language models (LLMs) often adopt finetuning to unlock their capabilities for downstream applications, our understanding on the inductive biases (especially the scaling properties) of different finetuning methods is still limited. To fill this gap, we conduct systematic experiments studying whether and how different scaling factors, including LLM model size, pretraining data size, new finetuning parameter size and finetuning data size, affect the finetuning performance. We consider two types of finetuning – full-model tuning (FMT) and parameter efficient tuning (PET, including prompt tuning and LoRA), and explore their scaling behaviors in the data-limited regime where the LLM model size substantially outweighs the finetuning data size. Based on two sets of pretrained bilingual LLMs from 1B to 16B and experiments on bilingual machine translation and multilingual summarization benchmarks, we find that 1) LLM finetuning follows a powerbased multiplicative joint scaling law between finetuning data size and each other scaling factor; 2) LLM finetuning benefits more from LLM model scaling than pretraining data scaling, and PET parameter scaling is generally ineffective; and 3) the optimal finetuning method is highly task- and finetuning data-dependent. We hope our findings could shed light on understanding, selecting and developing LLM finetuning methods.
Biao Zhang 0006, Zhongtao Liu, Colin Cherry, Orhan Firat
ICLR1
2024 When Does Monolingual Data Help Multilingual Translation: The Role of Domain and Model Scale
abstract
Christos Baziotis, Biao Zhang, Alexandra Birch, Barry Haddow. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Christos Baziotis, Biao Zhang 0006, Alexandra Birch, Barry Haddow
NAACL-HLT2
2024 Scaling Sign Language Translation
abstract
Sign language translation (SLT) addresses the problem of translating information from a sign language in video to a spoken language in text. Existing studies, while showing progress, are often limited to narrow domains and/or few sign languages and struggle with open-domain tasks. In this paper, we push forward the frontier of SLT by scaling pretraining data, model size, and number of translation directions. We perform large-scale SLT pretraining on different data including 1) noisy multilingual Youtube SLT data, 2) parallel text corpora, and 3) SLT data augmented by translating video captions to other languages with off-the-shelf machine translation models. We unify different pretraining tasks with task-specific prompts under the encoder-decoder architecture, and initialize the SLT model with pretrained (m/By)T5 models across model sizes. SLT pretraining results on How2Sign and FLEURS-ASL\#0 (ASL to 42 spoken languages) demonstrate the significance of data/model scaling and cross-lingual cross-modal transfer, as well as the feasibility of zero-shot SLT. We finetune the pretrained SLT models on 5 downstream open-domain SLT benchmarks covering 5 sign languages. Experiments show substantial quality improvements over the vanilla baselines, surpassing the previous state-of-the-art (SOTA) by wide margins.
Biao Zhang 0006, Garrett Tanzer, Orhan Firat
NeurIPS1
2023 Self-training Reduces Flicker in Retranslation-based Simultaneous Translation
abstract
In simultaneous translation, the retranslation approach has the advantage of requiring no modifications to the inference engine.However, in order to reduce the undesirable flicker in the output, previous work has resorted to increasing the latency through masking, and introducing specialised inference, thus losing the simplicity of the approach.In this work, we show that self-training improves the flickerlatency tradeoff, while maintaining similar translation quality to the original.Our analysis indicates that self-training reduces flicker by controlling monotonicity.Furthermore, selftraining can be combined with biased beam search to further improve the flicker-latency tradeoff.
Sukanta Sen, Rico Sennrich, Biao Zhang 0006, Barry Haddow
EACL3
2023 Efficient CTC Regularization via Coarse Labels for End-to-End Speech Translation
abstract
For end-to-end speech translation, regularizing the encoder with the Connectionist Temporal Classification (CTC) objective using the source transcript or target translation as labels can greatly improve quality metrics.However, CTC demands an extra prediction layer over the vocabulary space, bringing in nonnegligible model parameters and computational overheads, although this layer is typically not used for inference.In this paper, we re-examine the need for genuine vocabulary labels for CTC for regularization and explore strategies to reduce the CTC label space, targeting improved efficiency without quality degradation.We propose coarse labeling for CTC (CoLaCTC), which merges vocabulary labels via simple heuristic rules, such as using truncation, division or modulo (MOD) operations.Despite its simplicity, our experiments on 4 source and 8 target languages show that CoLaCTC with MOD particularly can compress the label space aggressively to 256 and even further, gaining training efficiency (1.18× ∼ 1.77× speedup depending on the original vocabulary size) yet still delivering comparable or better performance than the CTC baseline.We also show that CoLaCTC successfully generalizes to CTC regularization regardless of using transcript or translation for labeling.
Biao Zhang 0006, Barry Haddow, Rico Sennrich
EACL1
2023 SLTUNET: A Simple Unified Model for Sign Language Translation
Biao Zhang 0006, Mathias Müller 0002, Rico Sennrich
ICLR1
2023 Prompting Large Language Model for Machine Translation: A Case Study
abstract
Research on prompting has shown excellent performance with little or even no supervised training across many tasks. However, prompting for machine translation is still under-explored in the literature. We fill this gap by offering a systematic study on prompting strategies for translation, examining various factors for prompt template and demonstration example selection. We further explore the use of monolingual data and the feasibility of cross-lingual, cross-domain, and sentence-to-document transfer learning in prompting. Extensive experiments with GLM-130B (Zeng et al., 2022) as the testbed show that 1) the number and the quality of prompt examples matter, where using suboptimal examples degenerates translation; 2) several features of prompt examples, such as semantic similarity, show significant Spearman correlation with their prompting performance; yet, none of the correlations are strong enough; 3) using pseudo parallel prompt examples constructed from monolingual data via zero-shot prompting could improve translation; and 4) improved performance is achievable by transferring knowledge from prompt examples selected in other settings. We finally provide an analysis on the model outputs and discuss several problems that prompting still suffers from.
Biao Zhang 0006, Barry Haddow, Alexandra Birch
ICML1
2023 MADLAD-400: A Multilingual And Document-Level Large Audited Dataset
abstract
We introduce MADLAD-400, a manually audited, general domain 3T token monolingual dataset based on CommonCrawl, spanning 419 languages. We discuss the limitations revealed by self-auditing MADLAD-400, and the role data auditing had in the dataset creation process. We then train and release a 10.7B-parameter multilingual machine translation model on 250 billion tokens covering over 450 languages using publicly available data, and find that it is competitive with models that are significantly larger, and report the results on different domains. In addition, we train a 8B-parameter language model, and assess the results on few-shot translation. We make the baseline models available to the research community.
Sneha Reddy Kudugunta, Isaac Caswell, Biao Zhang 0006, Xavier Garcia, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, Orhan Firat
NeurIPS3
2022 Multilingual Document-Level Translation Enables Zero-Shot Transfer From Sentences to Documents
abstract
Biao Zhang, Ankur Bapna, Melvin Johnson, Ali Dabirmoghaddam, Naveen Arivazhagan, Orhan Firat. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Biao Zhang 0006, Ankur Bapna, Melvin Johnson, Ali Dabirmoghaddam, Naveen Arivazhagan, Orhan Firat
ACL (1)1
2022 Data Scaling Laws in NMT: The Effect of Noise and Architecture
abstract
In this work, we study the effect of varying the architecture and training data quality on the data scaling properties of Neural Machine Translation (NMT). First, we establish that the test loss of encoder-decoder transformer models scales as a power law in the number of training samples, with a dependence on the model size. Then, we systematically vary aspects of the training setup to understand how they impact the data scaling laws. In particular, we change the following (1) Architecture and task setup: We compare to a transformer-LSTM hybrid, and a decoder-only transformer with a language modeling loss (2) Noise level in the training distribution: We experiment with filtering, and adding iid synthetic noise. In all the above cases, we find that the data scaling exponents are minimally impacted, suggesting that marginally worse architectures or training data can be compensated for by adding more data. Lastly, we find that using back-translated data instead of parallel data, can significantly degrade the scaling exponent.
Yamini Bansal, Behrooz Ghorbani, Ankush Garg, Biao Zhang 0006, Colin Cherry, Behnam Neyshabur, Orhan Firat
ICML4
2022 Examining Scaling and Transfer of Language Model Architectures for Machine Translation
abstract
Natural language understanding and generation models follow one of the two dominant architectural paradigms: language models (LMs) that process concatenated sequences in a single stack of layers, and encoder-decoder models (EncDec) that utilize separate layer stacks for input and output processing. In machine translation, EncDec has long been the favoured approach, but with few studies investigating the performance of LMs. In this work, we thoroughly examine the role of several architectural design choices on the performance of LMs on bilingual, (massively) multilingual and zero-shot translation tasks, under systematic variations of data conditions and model sizes. Our results show that: (i) Different LMs have different scaling properties, where architectural differences often have a significant impact on model performance at small scales, but the performance gap narrows as the number of parameters increases, (ii) Several design choices, including causal masking and language-modeling objectives for the source sequence, have detrimental effects on translation quality, and (iii) When paired with full-visible masking for source sequences, LMs could perform on par with EncDec on supervised bilingual and multilingual translation tasks, and improve greatly on zero-shot directions by facilitating the reduction of off-target translations.
Biao Zhang 0006, Behrooz Ghorbani, Ankur Bapna, Yong Cheng 0003, Xavier Garcia, Jonathan Shen, Orhan Firat
ICML1
2022 Revisiting End-to-End Speech-to-Text Translation From Scratch
abstract
End-to-end (E2E) speech-to-text translation (ST) often depends on pretraining its encoder and/or decoder using source transcripts via speech recognition or text translation tasks, without which translation performance drops substantially. However, transcripts are not always available, and how significant such pretraining is for E2E ST has rarely been studied in the literature. In this paper, we revisit this question and explore the extent to which the quality of E2E ST trained on speech-translation pairs alone can be improved. We reexamine several techniques proven beneficial to ST previously, and offer a set of best practices that biases a Transformer-based E2E ST system toward training from scratch. Besides, we propose parameterized distance penalty to facilitate the modeling of locality in the self-attention model for speech. On four benchmarks covering 23 languages, our experiments show that, without using any transcripts or pretraining, the proposed system reaches and even outperforms previous studies adopting pretraining, although the gap remains in (extremely) low-resource settings. Finally, we discuss neural acoustic feature modeling, where a neural model is designed to extract acoustic features from raw speech signals directly, with the goal to simplify inductive biases and add freedom to the model in describing speech. For the first time, we demonstrate its feasibility and show encouraging results on ST tasks.
Biao Zhang 0006, Barry Haddow, Rico Sennrich
ICML1
2021 Beyond Sentence-Level End-to-End Speech Translation: Context Helps
abstract
Biao Zhang, Ivan Titov, Barry Haddow, Rico Sennrich. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Biao Zhang 0006, Ivan Titov 0001, Barry Haddow, Rico Sennrich
ACL/IJCNLP (1)1
2021 Sparse Attention with Linear Units
abstract
Recently, it has been argued that encoderdecoder models can be made more interpretable by replacing the softmax function in the attention with its sparse variants.In this work, we introduce a novel, simple method for achieving sparsity in attention: we replace the softmax activation with a ReLU, and show that sparsity naturally emerges from such a formulation.Training stability is achieved with layer normalization with either a specialized initialization or an additional gating function.Our model, which we call Rectified Linear Attention (ReLA), is easy to implement and more efficient than previously proposed sparse attention mechanisms.We apply ReLA to the Transformer and conduct experiments on five machine translation tasks.ReLA achieves translation performance comparable to several strong baselines, with training and decoding speed similar to that of the vanilla attention.Our analysis shows that ReLA delivers high sparsity rate and head diversity, and the induced cross attention achieves better accuracy with respect to source-target word alignment than recent sparsified softmax-based models.Intriguingly, ReLA heads also learn to attend to nothing (i.e.'switch off') for some queries, which is not possible with sparsified softmax alternatives.1
Biao Zhang 0006, Ivan Titov 0001, Rico Sennrich
EMNLP (1)1
2021 Share or Not? Learning to Schedule Language-Specific Capacity for Multilingual Translation
Biao Zhang 0006, Ankur Bapna, Rico Sennrich, Orhan Firat
ICLR1
2020 Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation
abstract
Massively multilingual models for neural machine translation (NMT) are theoretically attractive, but often underperform bilingual models and deliver poor zero-shot translations.In this paper, we explore ways to improve them.We argue that multilingual NMT requires stronger modeling capacity to support language pairs with varying typological characteristics, and overcome this bottleneck via language-specific components and deepening NMT architectures.We identify the off-target translation issue (i.e.translating into a wrong target language) as the major source of the inferior zero-shot performance, and propose random online backtranslation to enforce the translation of unseen training language pairs.Experiments on OPUS-100 (a novel multilingual dataset with 100 languages) show that our approach substantially narrows the performance gap with bilingual models in both oneto-many and many-to-many settings, and improves zero-shot performance by ∼10 BLEU, approaching conventional pivot-based methods. 1
Biao Zhang 0006, Philip Williams, Ivan Titov 0001, Rico Sennrich
ACL1
2019 Revisiting Low-Resource Neural Machine Translation: A Case Study
abstract
It has been shown that the performance of neural machine translation (NMT) drops starkly in low-resource conditions, underperforming phrase-based statistical machine translation (PBSMT) and requiring large amounts of auxiliary data to achieve competitive results.In this paper, we re-assess the validity of these results, arguing that they are the result of lack of system adaptation to low-resource settings.We discuss some pitfalls to be aware of when training low-resource NMT systems, and recent techniques that have shown to be especially helpful in low-resource settings, resulting in a set of best practices for low-resource NMT.In our experiments on German-English with different amounts of IWSLT14 training data, we show that, without the use of any auxiliary monolingual or multilingual data, an optimized NMT system can outperform PBSMT with far less data than previously claimed.We also apply these techniques to a low-resource Korean-English dataset, surpassing previously reported results by 4 BLEU.
Rico Sennrich, Biao Zhang 0006
ACL (1)2
2019 A Lightweight Recurrent Network for Sequence Modeling
abstract
Recurrent networks have achieved great success on various sequential tasks with the assistance of complex recurrent units, but suffer from severe computational inefficiency due to weak parallelization.One direction to alleviate this issue is to shift heavy computations outside the recurrence.In this paper, we propose a lightweight recurrent network, or LRN.LRN uses input and forget gates to handle long-range dependencies as well as gradient vanishing and explosion, with all parameterrelated calculations factored outside the recurrence.The recurrence in LRN only manipulates the weight assigned to each token, tightly connecting LRN with self-attention networks.We apply LRN as a drop-in replacement of existing recurrent units in several neural sequential models.Extensive experiments on six NLP tasks show that LRN yields the best running efficiency with little or no loss in model performance.1
Biao Zhang 0006, Rico Sennrich
ACL (1)1
2019 Improving Deep Transformer with Depth-Scaled Initialization and Merged Attention
abstract
Biao Zhang, Ivan Titov, Rico Sennrich. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Biao Zhang 0006, Ivan Titov 0001, Rico Sennrich
EMNLP/IJCNLP (1)1
2019 Root Mean Square Layer Normalization
abstract
Layer normalization (LayerNorm) has been successfully applied to various deep neural networks to help stabilize training and boost model convergence because of its capability in handling re-centering and re-scaling of both inputs and weight matrix. However, the computational overhead introduced by LayerNorm makes these improvements expensive and significantly slows the underlying network, e.g. RNN in particular. In this paper, we hypothesize that re-centering invariance in LayerNorm is dispensable and propose root mean square layer normalization, or RMSNorm. RMSNorm regularizes the summed inputs to a neuron in one layer according to root mean square (RMS), giving the model re-scaling invariance property and implicit learning rate adaptation ability. RMSNorm is computationally simpler and thus more efficient than LayerNorm. We also present partial RMSNorm, or pRMSNorm where the RMS is estimated from p% of the summed inputs without breaking the above properties. Extensive experiments on several tasks using diverse network architectures show that RMSNorm achieves comparable performance against LayerNorm but reduces the running time by 7%~64% on different models. Source code is available at https://github.com/bzhangGo/rmsnorm.
Biao Zhang 0006, Rico Sennrich
NeurIPS1