Shaolin Zhu

dblp:263/1683 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference
abstract
Mixture-of-Experts Multimodal Large Language Models (MoE MLLMs) suffer from a significant efficiency bottleneck during Expert Parallelism (EP) inference due to the straggler effect.This issue is worsened in the multimodal context, as existing token-count-based load balancing methods fail to address two unique challenges: (1) Information Heterogeneity, where numerous redundant visual tokens are treated equally to semantically critical ones, and (2) Modality Dynamics, where varying visual to text ratios across tasks lead to resource misallocation.To address these challenges, we propose MACS (Modality-Aware Capacity Scaling), a training-free inference framework.Specifically, MACS introduces an Entropy-Weighted Load mechanism to quantify the semantic value of visual tokens, addressing information heterogeneity.Additionally, the Dynamic Modality-Adaptive Capacity mechanism allocates expert resources based on the real-time modal composition of the input.Extensive experiments demonstrate that MACS significantly outperforms existing methods on various multimodal benchmarks, providing a novel and robust solution for the efficient deployment of MoE MLLMs in EP inference.
Shaolin Zhu
ACL (1)3
2026 AdaDPI: Document-level Translation Adaptive Agent via Dynamic Parametric Internalization
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities in machine translation.However, maintaining discourse coherence and terminological consistency remains a persistent challenge in documentlevel translation (DocMT).Existing solutions, such as memory-based agents, predominantly rely on explicit context concatenation.This paradigm treats historical context as a static external resource, which often leads to context dilution, high inference latency, and superficial knowledge integration.To address these limitations, we propose AdaDPI, an adaptive agentic framework that shifts the DocMT paradigm from static retrieval to dynamic parametric internalization.Specifically, we design a linguistic uncertainty monitor (LUM) to actively detect critical discourse discontinuities by the model's epistemic uncertainty.Upon detection, a context-to-parameter integrator (CPI) compiles retrieved external constraints directly into the model's intrinsic state via an online parameter adaptation mechanism.Through the online parameter adaptation on a lightweight adapter, AdaDPI internalizes document-specific norms into the model's intrinsic representations, enabling a progressive evolution of the translation strategy as the discourse unfolds.Extensive experiments on the discourse-rich GuoFeng and IWSLT2017 datasets demonstrate that AdaDPI significantly outperforms the SoTA baselines by more than 5 points on the consistency metric.
Hong Ren, Liting Deng, Shaolin Zhu, Deyi Xiong
ACL (1)3
2026 AMART: A multi-agent reflective framework for detecting and correcting faithfulness errors in translation
Shaolin Zhu, Deyi Xiong
Inf. Process. Manag.2
2025 LRM-LLaVA: Overcoming the Modality Gap of Multilingual Large Language-Vision Model for Low-Resource Languages
abstract
Multilingual large language-vision models (LVLMs), which understand and generate both text and images across multiple languages, have achieved remarkable performance on English-centric multimodal generation tasks. However, their performance on non-English tasks has been underwhelming. One major challenge with multilingual LVLMs is the modality gap between visual inputs and multilingual textual inputs/outputs due to the lack of high-quality multilingual training data. In this paper, we propose LRM-LLaVA, a multilingual large language-vision model designed for low-resource languages to overcome the modality gap. It is composed of four components: a visual encoder, a multilingual large language model, a vision-text representation projector, and a cross-modal regularizer. Both the projector and regularizer aim at reducing the modality gap and improving multilingual performance. To train LRM-LLaVA, we employ a two-stage training strategy including pre-training and instruction fine-tuning. Meanwhile, we construct a multilingual visual question answering dataset based on English open-source datasets and adopt multiple task instructions. To evaluate the performance of LVLMs across various languages, we construct four multilingual benchmarks for 10 languages, based on English open-source benchmarks. Experimental results show that LRM-LLaVA achieves competitive performance compared to other multilingual LVLMs of similar parameters.
Junchen Li, Bojian Jiang, Shaolin Zhu, Qingxuan Sun
AAAI4
2025 MLAS-LoRA: Language-Aware Parameters Detection and LoRA-Based Knowledge Transfer for Multilingual Machine Translation
abstract
Large language models (LLMs) have achieved remarkable progress in multilingual machine translation (MT), demonstrating strong performance even with limited parallel data.However, effectively fine-tuning LLMs for MT is challenging due to parameter interference, which arises from the conflicting demands of different language pairs and the risk of overwriting pre-trained knowledge.To address this issue, we propose MLAS-LoRA, a novel multiple language-aware LoRA knowledge transfer framework.MLAS-LoRA efficiently adapts LLMs to MT by selectively transferring knowledge from a large teacher to a small student model.Our approach first evaluates the awareness of neurons and extracts linguistic knowledge in the teacher model to both the general MT task and specific language pairs.We then propose a multiple language-specific LoRA architecture to inject the extracted knowledge into the student model.During fine-tuning, only the parameters of the relevant languagegeneral and language-specific LoRA modules are updated.Experimental results on diverse multilingual language pairs demonstrate that MLAS-LoRA significantly outperforms strong baselines by +1.7 BLEU on average, including standard fine-tuning and other parameterefficient methods.
Tianyu Dong, Bo Li 0131, Shaolin Zhu, Deyi Xiong
ACL (1)4
2025 MIT-10M: A Large Scale Parallel Corpus of Multilingual Image Translation
abstract
Image Translation (IT) holds immense potential across diverse domains, enabling the translation of textual content within images into various languages. However, existing datasets often suffer from limitations in scale, diversity, and quality, hindering the development and evaluation of IT models. To address this issue, we introduce MIT-10M, a large-scale parallel corpus of multilingual image translation with over 10M image-text pairs derived from real-world data, which has undergone extensive data cleaning and multilingual translation validation. It contains 0.8M images in three sizes, 28 categories, tasks with three levels of difficulty and 14 languages image-text pairs, which is a considerable improvement on existing datasets. We conduct extensive experiments to evaluate and train models on MIT-10M. The experimental results clearly indicate that our dataset has higher adaptability when it comes to evaluating the performance of the models in tackling challenging and complex image translation tasks in the real world. Moreover, the performance of the model fine-tuned with MIT-10M has tripled compared to the baseline model, further confirming its superiority.
Bo Li 0131, Shaolin Zhu, Lijie Wen 0001
COLING2
2025 CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency
abstract
Zhanming Shen, Hao Chen, Yulei Tang, Shaolin Zhu, Wentao Ye, Xiaomeng Hu, Haobo Wang, Gang Chen, Junbo Zhao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zhanming Shen, Hao Chen 0081, Yulei Tang, Shaolin Zhu, Wentao Ye, Xiaomeng Hu, Haobo Wang 0001, Gang Chen 0001, Junbo Zhao 0002
EMNLP4
2025 DCIS: Efficient Length Extrapolation of LLMs via Divide-and-Conquer Scaling Factor Search
abstract
Large language models (LLMs) based on the Transformer architecture usually have their context length limited due to the high training cost.Recent advancements extend the context window by adjusting the scaling factors of RoPE and fine-tuning.However, suboptimal initialization of these factors results in increased fine-tuning costs and reduced performance at target length.To address these challenges, we propose a novel RoPE-based fine-tuning framework that diverges from conventional scaling factors search.Specifically, we present a Divide-and-Conquer Incremental Search (DCIS) algorithm that strategically determines the better scaling factors.Further finetuning with the identified scaling factors effectively extends the context window of LLMs.Empirical results demonstrate that our methodology not only mitigates performance decay at extended target lengths but also allows the model to fine-tune on short contexts and generalize to long contexts, thereby reducing the cost of fine-tuning.The scaling factors obtained through DCIS can even perform effectively without fine-tuning.Further analysis of the search space reveals that DCIS achieves twice the search efficiency compared to other methods.We also examine the impact of the non-strictly increasing scaling factors utilized in DCIS and evaluate the general capabilities of LLMs across various context lengths.
Shaoyang Xu, Jianxiang Peng, Shaolin Zhu, Deyi Xiong
EMNLP4
2025 Multi-source knowledge fusion for multilingual loanword identification
Chenggang Mi 0001, Shaolin Zhu
Expert Syst. Appl.2
2025 Overcoming language barriers via machine translation with sparse Mixture-of-Experts fusion of large language models
Shaolin Zhu, Leiyu Pan, Dong Jian, Deyi Xiong
Inf. Process. Manag.1
2025 CoLE: A collaborative legal expert prompting framework for large language models in law
Bo Li 0131, Shuang Fan, Shaolin Zhu, Lijie Wen 0001
Knowl. Based Syst.3
2024 LANDeRMT: Dectecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine Translation
abstract
Recent advancements in large language models (LLMs) have shown promising results in multilingual translation even with limited bilingual supervision.The major challenges are catastrophic forgetting and parameter interference 1 for finetuning LLMs when provided parallel training data.To address these challenges, we propose LANDeRMT, a Language-Aware Neuron Detecting and Routing framework that selectively finetunes LLMs to Machine Translation with diverse translation training data.In LANDeRMT, we evaluate the awareness of neurons to MT tasks and categorize them into language-general and languagespecific neurons.This categorization enables selective parameter updates during finetuning, mitigating parameter interference and catastrophic forgetting issues.For the detected neurons, we further propose a conditional awareness-based routing mechanism to dynamically adjust language-general and languagespecific capacity within LLMs, guided by translation signals.Experimental results demonstrate that the proposed LANDeRMT is very effective in learning translation knowledge, significantly improving translation quality over various strong baselines for multiple language pairs.
Shaolin Zhu, Leiyu Pan, Bo Li 0131, Deyi Xiong
ACL (1)1
2024 Towards Robust In-Context Learning for Machine Translation with Large Language Models
abstract
Using large language models (LLMs) for machine translation via in-context learning (ICL) has become an interesting research direction of machine translation (MT) in recent years. Its main idea is to retrieve a few translation pairs as demonstrations from an additional datastore (parallel corpus) to guide translation without updating the LLMs. However, the underlying noise of retrieved demonstrations usually dramatically deteriorate the performance of LLMs. In this paper, we propose a robust method to enable LLMs to achieve robust translation with ICL. The method incorporates a multi-view approach, considering both sentence- and word-level information, to select demonstrations that effectively avoid noise. At the sentence level, a margin-based score is designed to avoid semantic noise. At the word level, word embeddings are utilized to evaluate the related tokens and change the weight of words in demonstrations. By considering both sentence- and word-level similarity, the proposed method provides fine-grained demonstrations that effectively prompt the translation of LLMs. Experimental results demonstrate the effectiveness of our method, particularly in domain adaptation.
Shaolin Zhu, Menglong Cui, Deyi Xiong
LREC/COLING1
2024 Mining parallel sentences from internet with multi-view knowledge distillation for low-resource language pairs
Shaolin Zhu, Shiwei Gu, Shangjie Li, Deyi Xiong
Knowl. Inf. Syst.1
2023 PEIT: Bridging the Modality Gap with Pre-trained Models for End-to-End Image Translation
abstract
Image translation is a task that translates an image containing text in the source language to the target language.One major challenge with image translation is the modality gap between visual text inputs and textual inputs/outputs of machine translation (MT).In this paper, we propose PEIT, an end-to-end image translation framework that bridges the modality gap with pre-trained models.It is composed of four essential components: a visual encoder, a shared encoder-decoder backbone network, a vision-text representation aligner equipped with the shared encoder and a cross-modal regularizer stacked over the shared decoder.Both the aligner and regularizer aim at reducing the modality gap.To train PEIT, we employ a twostage pre-training strategy with an auxiliary MT task: (1) pre-training the MT model on the MT training data to initialize the shared encoder-decoder backbone network; and (2) pre-training PEIT with the aligner and regularizer on a synthesized dataset with rendered images containing text from the MT training data.In order to facilitate the evaluation of PEIT and promote research on image translation, we create a large-scale image translation corpus ECOIT containing 480K imagetranslation pairs via crowd-sourcing and manual post-editing from real-world images in the e-commerce domain.Experiments on the curated ECOIT benchmark dataset demonstrate that PEIT substantially outperforms both cascaded image translation systems (OCR+MT) and previous strong end-to-end image translation model, with fewer parameters and faster decoding speed.Codes are available at https: //github.com/lishangjie1/PEIT.
Shaolin Zhu, Shangjie Li, Yikun Lei, Deyi Xiong
ACL (1)1
2023 MMNMT: Modularizing Multilingual Neural Machine Translation with Flexibly Assembled MoE and Dense Blocks
abstract
Mixture-of-Experts (MoE) based sparse architectures can significantly increase model capacity with sublinear computational overhead, which are hence widely used in massively multilingual neural machine translation (MNMT).However, they are prone to overfitting on lowresource language translation.In this paper, we propose a modularized MNMT framework that is able to flexibly assemble dense and MoEbased sparse modules to achieve the best of both worlds.The training strategy of the modularized MNMT framework consists of three stages: (1) Pre-training basic MNMT models with different training objectives or model structures, (2) Initializing modules of the framework with pre-trained couterparts (e.g., encoder, decoder and embedding layers) from the basic models and (3) Fine-tuning the modularized MNMT framework to fit modules from different models together.We pre-train three basic MNMT models from scratch: a dense model, an MoE-based sparse model and a new MoE model, termed as MoE-LGR that explores multiple Language-Group-specifc Routers to incorporate language group knowledge into MNMT.The strengths of these pre-trained models are either on low-resource language translation, highresource language translation or zero-shot translation.Our modularized MNMT framework attempts to incorporate these advantages into a single model with reasonable initialization and fine-tuning.Experiments on widely-used benchmark datasets demonstrate that the proposed modularized MNMT framwork substantially outperforms both MoE and dense models on high-and low-resource language translation as well as zero-shot translation.Our framework facilitates the combination of different methods with their own strengths and recycling off-the-shelf models for multilingual neural machine translation.Codes are available at https://github.com/lishangjie1/MMNMT.
Shangjie Li, Xiangpeng Wei, Shaolin Zhu, Baosong Yang, Deyi Xiong
EMNLP3
2023 Ternary Data, Triangle Decoding, Three Tasks, a Multitask Learning Speech Translation Model
Boxing Chen, Shaolin Zhu, Luo Si
ICANN (3)3
2023 Unsupervised Parallel Sentences of Machine Translation for Asian Language Pairs
abstract
Parallel sentence pairs play a very important role in many natural language processing tasks, especially cross-lingual tasks such as machine translation. So far, many Asian language pairs lack bilingual parallel sentences. As collecting bilingual parallel data is very time-consuming and difficult, it is very important for many low-resource Asian language pairs. While existing methods have shown encouraging results, they rely on bilingual data seriously or have some drawbacks in an unsupervised situation. To address these issues, we propose a new unsupervised similarity calculation and dynamic selection metric to obtain parallel sentence pairs in an unsupervised situation. First, our method maps bilingual word embedding by postdoc adversarial training, which rotates the source space to match the target without parallel data. Then, we introduce a new cross-domain similarity adaption to obtain parallel sentence pairs. Experimental results on real-world datasets show that our model can obtain better accuracy and recall on mining parallel sentence pairs. We also show that the extracted bilingual sentence corpora can significantly improve the performance of neural machine translation.
Shaolin Zhu, Chenggang Mi 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.1