Shanbo Cheng

dblp:185/5589 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-6115-9483ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
YearPublicationVenuePosition
2026 Improving Long-Context Translation via Self-Supervised Dual Learning
abstract
Large language models (LLMs) with long context windows offer the potential to translate entire documents in a single pass, yet they frequently suffer from catastrophic information distortion, undermining the strict faithfulness required for translation.This challenge is compounded by the scarcity of documentlevel parallel data, which makes both supervised fine-tuning and reliable evaluation prohibitively expensive.We propose LongDu, a self-supervised post-training framework that improves long-document translation reliability via round-trip consistency.Given monolingual documents, LongDu samples multiple candidate translations, back-translates each candidate, and optimizes the model to prefer translations that best reconstruct the source.To make this signal robust for long-form generation, we design a reward that filters trivial failure modes (e.g., copying and local language drift) before applying a reconstruction and fluency score, enabling stable reinforcement learning without human annotations.We additionally introduce Long-CIRT, an automatic evaluation protocol that quantifies information distortion by measuring how much a LLM's performance degrades after a translation cycle.Across multiple base models, LongDu substantially improves information retention and translation quality, with gains that generalize beyond the training length range and to unseen target languages.
Shanbo Cheng, Shuaijie She, Jiajun Chen 0001, Shujian Huang
ACL (1)1
2025 From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
abstract
Recent advances in Automatic Speech Recognition (ASR) have been largely fueled by massive speech corpora.However, extending coverage to diverse languages with limited resources remains a formidable challenge.This paper introduces Speech Back-Translation, a scalable pipeline that improves multilingual ASR models by converting large-scale text corpora into synthetic speech via off-the-shelf text-tospeech (TTS) models.We demonstrate that just tens of hours of real transcribed speech can effectively train TTS models to generate synthetic speech at hundreds of times the original volume while maintaining high quality.To evaluate synthetic speech quality, we develop an intelligibility-based assessment framework and establish clear thresholds for when synthetic data benefits ASR training.Using Speech Back-Translation, we generate more than 500,000 hours of synthetic speech in ten languages and continue pre-training Whisperlarge-v3, achieving average transcription error reductions of over 30%.These results highlight the scalability and effectiveness of Speech Back-Translation for enhancing multilingual ASR systems.
Tianduo Wang, Shanbo Cheng
EMNLP4
2025 EnAnchored-X2X: English-Anchored Optimization for Many-to-Many Translation
abstract
Large language models (LLMs) have demonstrated strong machine translation capabilities for English-centric language pairs but underperform in direct non-English (x2x) translation.This work addresses this limitation through a synthetic data generation framework that leverages models' established English-to-x (en2x) capabilities.By extending English parallel corpora into omnidirectional datasets and developing an English-referenced quality evaluation proxy, we enable effective collection of high-quality x2x training data.Combined with preference-based optimization, our method achieves significant improvement across 72 x2x directions for widely used LLMs, while generalizing to enhance en2x performance.The results demonstrate that strategic exploitation of English-centric strengths can bootstrap comprehensive multilingual translation capabilities in LLMs.We release codes, datasets, and model checkpoints at
Sen Yang 0015, Jiajun Chen 0001, Shujian Huang, Shanbo Cheng
EMNLP6
2024 Retaining Key Information under High Compression Ratios: Query-Guided Compressor for LLMs
abstract
Zhiwei Cao, Qian Cao, Yu Lu, Ningxin Peng, Luyang Huang, Shanbo Cheng, Jinsong Su. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Ningxin Peng, Luyang Huang, Shanbo Cheng, Jinsong Su
ACL (1)6
2024 G-DIG: Towards Gradient-based DIverse and hiGh-quality Instruction Data Selection for Machine Translation
abstract
Large Language Models (LLMs) have demonstrated remarkable abilities in general scenarios.Instruction finetuning empowers them to align with humans in various tasks.Nevertheless, the Diversity and Quality of the instruction data remain two main challenges for instruction finetuning.With regard to this, in this paper, we propose a novel gradient-based method to automatically select high-quality and diverse instruction finetuning data for machine translation.Our key innovation centers around analyzing how individual training examples influence the model during training.Specifically, we select training examples that exert beneficial influences on the model as high-quality ones by means of Influence Function plus a small high-quality seed dataset.Moreover, to enhance the diversity of the training data we maximize the variety of influences they have on the model by clustering on their gradients and resampling.Extensive experiments on WMT22 and FLORES translation tasks demonstrate the superiority of our methods, and in-depth analysis further validates their effectiveness and generalization.1
Xingyuan Pan, Luyang Huang, Liyan Kang, Shanbo Cheng
ACL (1)6
2024 MT-PATCHER: Selective and Extendable Knowledge Distillation from Large Language Models for Machine Translation
abstract
Jiahuan Li, Shanbo Cheng, Shujian Huang, Jiajun Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jiahuan Li, Shanbo Cheng, Shujian Huang, Jiajun Chen 0001
NAACL-HLT2
2024 Eliciting the Translation Ability of Large Language Models via Multilingual Finetuning with Translation Instructions
abstract
Abstract Large-scale pretrained language models (LLMs), such as ChatGPT and GPT4, have shown strong abilities in multilingual translation, without being explicitly trained on parallel corpora. It is intriguing how the LLMs obtain their ability to carry out translation instructions for different languages. In this paper, we present a detailed analysis by finetuning a multilingual pretrained language model, XGLM-7.5B, to perform multilingual translation following given instructions. Firstly, we show that multilingual LLMs have stronger translation abilities than previously demonstrated. For a certain language, the translation performance depends on its similarity to English and the amount of data used in the pretraining phase. Secondly, we find that LLMs’ ability to carry out translation instructions relies on the understanding of translation instructions and the alignment among different languages. With multilingual finetuning with translation instructions, LLMs could learn to perform the translation task well even for those language pairs unseen during the instruction tuning phase.
Jiahuan Li, Hao Zhou 0012, Shujian Huang, Shanbo Cheng, Jiajun Chen 0001
Trans. Assoc. Comput. Linguistics4
2023 Visual Information Matters for ASR Error Correction
abstract
Aiming to improve the Automatic Speech Recognition (ASR) outputs with a post-processing step, ASR error correction (EC) techniques have been widely developed due to their efficiency in using parallel text data. Previous works mainly focus on using text or/ and speech data, which hinders the performance gain when not only text and speech information, but other modalities, such as visual information are critical for EC. The challenges are mainly two folds: one is that previous work fails to emphasize visual information, thus rare exploration has been studied. The other is that the community lacks a high-quality benchmark where visual information matters for the EC models. Therefore, this paper provides 1) simple yet effective methods, namely gated fusion and image captions as prompts to incorporate visual information to help EC; 2) large-scale benchmark datasets, namely Visual-ASR-EC, where each item in the training data consists of visual, speech, and text information, and the test data are carefully selected by human annotators to ensure that even humans could make mistakes when visual information is missing. Experimental results show that using captions as prompts could effectively use the visual information and surpass state-of-the-art methods by upto 1.2% in Word Error Rate(WER), which also indicates that visual information is critical in our proposed Visual-ASR-EC dataset.
Vanya Bannihatti Kumar, Shanbo Cheng, Ningxin Peng
ICASSP2
2023 Only 5% Attention Is All You Need: Efficient Long-range Document-level Neural Machine Translation
abstract
Zihan Liu, Zewei Sun, Shanbo Cheng, Shujian Huang, Mingxuan Wang. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Zewei Sun, Shanbo Cheng, Shujian Huang, Mingxuan Wang
IJCNLP (1)3
2022 Unified Multimodal Punctuation Restoration Framework for Mixed-Modality Corpus
abstract
The punctuation restoration task aims to correctly punctuate the output transcriptions of automatic speech recognition systems. Previous punctuation models, either using text only or demanding the corresponding audio, tend to be constrained by real scenes, where unpunctuated sentences are a mixture of those with and without audio. This paper proposes a unified multimodal punctuation restoration framework, named UniPunc, to punctuate the mixed sentences with a single model. UniPunc jointly represents audio and non-audio samples in a shared latent space, based on which the model learns a hybrid representation and punctuates both kinds of samples. We validate the effectiveness of the UniPunc on real-world datasets, which outperforms various strong baselines (e.g. BERT, MuSe) by at least 0.8 overall F1 scores, making a new state-of-the-art. Extensive experiments show that UniPunc’s design is a pervasive solution: by grafting onto previous models, UniPunc enables them to punctuate on the mixed corpus. Our code is available at github.com/Yaoming95/UniPunc
Yaoming Zhu, Shanbo Cheng, Mingxuan Wang
ICASSP3
2022 switch-GLAT: Multilingual Parallel Machine Translation Via Code-Switch Decoder
Zhenqiao Song, Hao Zhou 0012, Lihua Qian, Jingjing Xu 0001, Shanbo Cheng, Mingxuan Wang, Lei Li 0005
ICLR5
2021 Learning Kernel-Smoothed Machine Translation with Retrieved Examples
abstract
How to effectively adapt neural machine translation (NMT) models according to emerging cases without retraining?Despite the great success of neural machine translation, updating the deployed models online remains a challenge.Existing non-parametric approaches that retrieve similar examples from a database to guide the translation process are promising but are prone to overfit the retrieved examples.However, non-parametric methods are prone to overfit the retrieved examples.In this work, we propose to learn Kernel-Smoothed Translation with Example Retrieval (KSTER), an effective approach to adapt neural machine translation models online.Experiments on domain adaptation and multi-domain machine translation datasets show that even without expensive retraining, KSTER is able to achieve improvement of 1.1 to 1.5 BLEU scores over the best existing online adaptation methods.The code and trained models are released at https://github.com/jiangqn/KSTER.
Qingnan Jiang, Mingxuan Wang, Shanbo Cheng, Shujian Huang, Lei Li 0005
EMNLP (1)4
2020 Acquiring Knowledge from Pre-Trained Model to Neural Machine Translation
abstract
Pre-training and fine-tuning have achieved great success in natural language process field. The standard paradigm of exploiting them includes two steps: first, pre-training a model, e.g. BERT, with a large scale unlabeled monolingual data. Then, fine-tuning the pre-trained model with labeled data from downstream tasks. However, in neural machine translation (NMT), we address the problem that the training objective of the bilingual task is far different from the monolingual pre-trained model. This gap leads that only using fine-tuning in NMT can not fully utilize prior language knowledge. In this paper, we propose an Apt framework for acquiring knowledge from pre-trained model to NMT. The proposed approach includes two modules: 1). a dynamic fusion mechanism to fuse task-specific features adapted from general knowledge into NMT network, 2). a knowledge distillation paradigm to learn language knowledge continuously during the NMT training process. The proposed approach could integrate suitable knowledge from pre-trained models to improve the NMT. Experimental results on WMT English to German, German to English and Chinese to English machine translation tasks show that our model outperforms strong baselines and the fine-tuning counterparts.
Rongxiang Weng, Heng Yu 0006, Shujian Huang, Shanbo Cheng, Weihua Luo
AAAI4
2020 Language-aware Interlingua for Multilingual Neural Machine Translation
abstract
Multilingual neural machine translation (NMT) has led to impressive accuracy improvements in low-resource scenarios by sharing common linguistic information across languages.However, the traditional multilingual model fails to capture the diversity and specificity of different languages, resulting in inferior performance compared with individual models that are sufficiently trained.In this paper, we incorporate a language-aware interlingua into the Encoder-Decoder architecture.The interlingual network enables the model to learn a language-independent representation from the semantic spaces of different languages, while still allowing for language-specific specialization of a particular language-pair.Experiments show that our proposed method achieves remarkable improvements over state-of-the-art multilingual NMT baselines and produces comparable performance with strong individual models.
Changfeng Zhu, Heng Yu 0006, Shanbo Cheng, Weihua Luo
ACL3
2016 PRIMT: A Pick-Revise Framework for Interactive Machine Translation
abstract
Shanbo Cheng, Shujian Huang, Huadong Chen, Xin-Yu Dai, Jiajun Chen. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Shanbo Cheng, Shujian Huang, Huadong Chen, Xinyu Dai, Jiajun Chen 0001
HLT-NAACL1