Shuming Ma

dblp:190/7739 · DBLP profile ↗
← Back
53ranked-venue papers
6as first author
31since 2021 · last 2025
0000-0003-1091-1206ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 6 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Bitnet.cpp: Efficient Edge Inference for Ternary LLMs
abstract
Jinheng Wang, Hansong Zhou, Ting Song, Shijie Cao, Yan Xia, Ting Cao, Jianyu Wei, Shuming Ma, Hongyu Wang, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jinheng Wang, Hansong Zhou, Shijie Cao, Yan Xia 0005, Jianyu Wei, Shuming Ma, Furu Wei
ACL (1)8
2025 Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
abstract
Recent studies have shown that making a model spend more time thinking through longer Chain of Thoughts (CoTs) enables it to gain significant improvements in complex reasoning tasks. While current researches continue to explore the benefits of increasing test-time compute by extending the CoT lengths of Large Language Models (LLMs), we are concerned about a potential issue hidden behind the current pursuit of test-time scaling: Would excessively scaling the CoT length actually bring adverse effects to a model's reasoning performance? Our explorations on mathematical reasoning tasks reveal an unexpected finding that scaling with longer CoTs can indeed impair the reasoning performance of LLMs in certain domains. Moreover, we discover that there exists an optimal scaled length distribution that differs across different domains. Based on these insights, we propose a Thinking-Optimal Scaling strategy. Our method first uses a small set of seed data with varying response length distributions to teach the model to adopt different reasoning efforts for deep thinking. Then, the model selects its shortest correct response under different reasoning efforts on additional problems for self-improvement. Our self-improved models built upon Qwen2.5-32B-Instruct outperform other distillation-based 32B o1-like models across various math benchmarks, and achieve performance on par with the teacher model QwQ-32B-Preview that produces the seed data.
Wenkai Yang, Shuming Ma, Yankai Lin 0001, Furu Wei
NeurIPS2
2025 BitNet: 1-bit Pre-training for Large Language Models
abstract
The increasing size of large language models (LLMs) has posed challenges for deployment and raised concerns about environmental impact due to high energy consumption. Previous research typically applies quantization after pre-training. While these methods avoid the need for model retraining, they often cause notable accuracy loss at extremely low bit-widths. In this work, we explore the feasibility and scalability of 1-bit pre-training. We introduce BitNet b1 and BitNet b1.58, the scalable and stable 1-bit Transformer architecture designed for LLMs. Specifically, we introduce BitLinear as a drop-in replacement of the nn.Linear layer in order to train 1-bit weights from scratch. Experimental results show that BitNet b1 achieves competitive performance, compared to state-of-the-art 8-bit quantization methods and FP16 Transformer baselines. With the ternary weight, BitNet b1.58 matches the half-precision Transformer LLM with the same model size and training tokens in terms of both perplexity and end-task performance, while being significantly more cost-effective in terms of latency, memory, throughput, and energy consumption. More profoundly, BitNet defines a new scaling law and recipe for training new generations of LLMs that are both high-performance and cost-effective. It enables a new computation paradigm and opens the door for designing specific hardware optimized for 1-bit LLMs.
Hongyu Wang 0009, Shuming Ma, Lingxiao Ma, Lei Wang 0222, Wenhui Wang 0003, Li Dong 0004, Shaohan Huang, Huaijie Wang, Jilong Xue, Yi Wu 0013, Furu Wei
J. Mach. Learn. Res.2
2024 Grounding Multimodal Large Language Models to the World
abstract
We introduce Kosmos-2, a Multimodal Large Language Model (MLLM), enabling new capabilities of perceiving object descriptions (e.g., bounding boxes) and grounding text to the visual world. Specifically, we represent text spans (i.e., referring expressions and noun phrases) as links in Markdown, i.e., [text span](bounding boxes), where object descriptions are sequences of location tokens. To train the model, we construct a large-scale dataset about grounded image-text pairs (GrIT) together with multimodal corpora. In addition to the existing capabilities of MLLMs (e.g., perceiving general modalities, following instructions, and performing in-context learning), Kosmos-2 integrates the grounding capability to downstream applications, while maintaining the conventional capabilities of MLLMs (e.g., perceiving general modalities, following instructions, and performing in-context learning). Kosmos-2 is evaluated on a wide range of tasks, including (i) multimodal grounding, such as referring expression comprehension and phrase grounding, (ii) multimodal referring, such as referring expression generation, (iii) perception-language tasks, and (iv) language understanding and generation. This study sheds a light on the big convergence of language, multimodal perception, and world modeling, which is a key step toward artificial general intelligence. Code can be found in [https://aka.ms/kosmos-2](https://aka.ms/kosmos-2).
Zhiliang Peng, Wenhui Wang 0003, Li Dong 0004, Yaru Hao, Shaohan Huang, Shuming Ma, Qixiang Ye, Furu Wei
ICLR6
2024 KOSMOS-E : Learning to Follow Instruction for Robotic Grasping
abstract
Tuning on instruction-following data has been shown to enhance the capabilities and controllability of language models, but the idea is less explored in the robotic field. In this work, we introduce KOSMOS-E, a Multimodal Large Language Model (MLLM) that leverages instruction-following robotic grasping data to enhance capabilities for precise and intricate robotic grasping maneuvers. To achieve this, we craft a large-scale instruction-following robotic grasping dataset, termed INSTRUCT-GRASP, primarily comprising two aspects: (i) grasp a single object following varying levels of granularity descriptions, e.g., different angles and aspects, and (ii) grasp a specific object within a multi-object environment following specific attributes, e.g., color and shape. Extensive experiments show the effectiveness of KOSMOS-E on robotic grasping tasks across a variety of environments.
Shaohan Huang, Li Dong 0004, Wenhui Wang 0003, Shuming Ma, Furu Wei
IROS6
2024 You Only Cache Once: Decoder-Decoder Architectures for Language Models
abstract
We introduce a decoder-decoder architecture, YOCO, for large language models, which only caches key-value pairs once. It consists of two components, i.e., a cross-decoder stacked upon a self-decoder. The self-decoder efficiently encodes global key-value (KV) caches that are reused by the cross-decoder via cross-attention. The overall model behaves like a decoder-only Transformer, although YOCO only caches once. The design substantially reduces GPU memory demands, yet retains global attention capability. Additionally, the computation flow enables prefilling to early exit without changing the final output, thereby significantly speeding up the prefill stage. Experimental results demonstrate that YOCO achieves favorable performance compared to Transformer in various settings of scaling up model size and number of training tokens. We also extend YOCO to 1M context length with near-perfect needle retrieval accuracy. The profiling results show that YOCO improves inference memory, prefill latency, and throughput by orders of magnitude across context lengths and model sizes.
Yutao Sun, Li Dong 0004, Shaohan Huang, Wenhui Wang 0003, Shuming Ma, Quanlu Zhang, Jianyong Wang 0001, Furu Wei
NeurIPS6
2024 Multi-Head Mixture-of-Experts
abstract
Sparse Mixtures of Experts (SMoE) scales model capacity without significant increases in computational costs. However, it exhibits the low expert activation issue, i.e., only a small subset of experts are activated for optimization, leading to suboptimal performance and limiting its effectiveness in learning a larger number of experts in complex tasks. In this paper, we propose Multi-Head Mixture-of-Experts (MH-MoE). MH-MoE split each input token into multiple sub-tokens, then these sub-tokens are assigned to and processed by a diverse set of experts in parallel, and seamlessly reintegrated into the original token form. The above operations enables MH-MoE to significantly enhance expert activation while collectively attend to information from various representation spaces within different experts to deepen context understanding. Besides, it's worth noting that our MH-MoE is straightforward to implement and decouples from other SMoE frameworks, making it easy to integrate with these frameworks for enhanced performance. Extensive experimental results across different parameter scales (300M to 7B) and three pre-training tasks—English-focused language modeling, multi-lingual language modeling and masked multi-modality modeling—along with multiple downstream validation tasks, demonstrate the effectiveness of MH-MoE.
Shaohan Huang, Wenhui Wang 0003, Shuming Ma, Li Dong 0004, Furu Wei
NeurIPS4
2024 DeepNet: Scaling Transformers to 1,000 Layers
abstract
In this paper, we propose a simple yet effective method to stabilize extremely deep Transformers. Specifically, we introduce a new normalization function (DeepNorm) to modify the residual connection in Transformer, accompanying with theoretically derived initialization. In-depth theoretical analysis shows that model updates can be bounded in a stable way. The proposed method combines the best of two worlds, i.e., good performance of Post-LN and stable training of Pre-LN, makingDeepNorma preferred alternative. We successfully scale Transformers up to 1,000 layers (i.e., 2,500 attention and feed-forward network sublayers) without difficulty, which is one order of magnitude deeper than previous deep Transformers. Extensive experiments demonstrate thatDeepNethas superior performance across various benchmarks, including machine translation, language modeling (i.e., BERT, GPT) and vision pre-training (i.e., BEiT). Remarkably, on a multilingual benchmark with 7,482 translation directions, our 200-layer model with 3.2B parameters significantly outperforms the 48-layer state-of-the-art model with 12B parameters by 5 BLEU points, which indicates a promising scaling direction. Our code is available athttps://aka.ms/torchscale.
Hongyu Wang 0009, Shuming Ma, Li Dong 0004, Shaohan Huang, Dongdong Zhang 0001, Furu Wei
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Discourse-Centric Evaluation of Document-level Machine Translation with a New Densely Annotated Parallel Corpus of Novels
abstract
Yuchen Eleanor Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang, Mrinmaya Sachan, Ryan Cotterell. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yuchen Eleanor Jiang, Tianyu Liu 0004, Shuming Ma, Dongdong Zhang 0001, Mrinmaya Sachan, Ryan Cotterell
ACL (1)3
2023 A Length-Extrapolatable Transformer
abstract
Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, Furu Wei. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yutao Sun, Li Dong 0004, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Furu Wei
ACL (1)4
2023 GanLM: Encoder-Decoder Pre-training with an Auxiliary Discriminator
abstract
Jian Yang, Shuming Ma, Li Dong, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang, Liqun Yang, Furu Wei, Zhoujun Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jian Yang 0030, Shuming Ma, Li Dong 0004, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang 0001, Liqun Yang, Furu Wei, Zhoujun Li 0001
ACL (1)2
2023 HanoiT: Enhancing Context-aware Translation via Selective Context
Jian Yang 0030, Yuwei Yin, Shuming Ma, Liqun Yang, Hongcheng Guo, Haoyang Huang, Dongdong Zhang 0001, Yutao Zeng, Zhoujun Li 0001, Furu Wei
DASFAA (3)3
2023 Are More Layers Beneficial to Graph Transformers?
Haiteng Zhao, Shuming Ma, Dongdong Zhang 0001, Zhi-Hong Deng 0001, Furu Wei
ICLR2
2023 Magneto: A Foundation Transformer
abstract
A big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name ”Transformers”, the above areas use different implementations for better performance, e.g., Post-LayerNorm for BERT, and Pre-LayerNorm for GPT and vision Transformers. We call for the development of Foundation Transformer for true general-purpose modeling, which serves as a go-to architecture for various tasks and modalities with guaranteed training stability. In this work, we introduce a Transformer variant, named Magneto, to fulfill the goal. Specifically, we propose Sub-LayerNorm for good expressivity, and the initialization strategy theoretically derived from DeepNet for stable scaling up. Extensive experiments demonstrate its superior performance and better stability than the de facto Transformer variants designed for various applications, including language modeling (i.e., BERT, and GPT), machine translation, vision pretraining (i.e., BEiT), speech recognition, and multimodal pretraining (i.e., BEiT-3).
Hongyu Wang 0009, Shuming Ma, Shaohan Huang, Li Dong 0004, Wenhui Wang 0003, Zhiliang Peng, Payal Bajaj, Saksham Singhal, Alon Benhaim, Barun Patra, Zhun Liu, Vishrav Chaudhary, Furu Wei
ICML2
2023 On the Pareto Front of Multilingual Neural Machine Translation
abstract
In this work, we study how the performance of a given direction changes with its sampling ratio in Multilingual Neural Machine Translation (MNMT). By training over 200 multilingual models with various model sizes, data sizes, and language directions, we find it interesting that the performance of certain translation direction does not always improve with the increase of its weight in the multi-task optimization objective. Accordingly, scalarization method leads to a multitask trade-off front that deviates from the traditional Pareto front when there exists data imbalance in the training corpus, which poses a great challenge to improve the overall performance of all directions. Based on our observations, we propose the Double Power Law to predict the unique performance trade-off front in MNMT, which is robust across various languages, data adequacy, and the number of tasks. Finally, we formulate the sample ratio selection problem in MNMT as an optimization problem based on the Double Power Law. Extensive experiments show that it achieves better performance than temperature searching and gradient manipulation methods with only 1/5 to 1/2 of the total training budget. We release the code at https://github.com/pkunlp-icler/ParetoMNMT for reproduction.
Liang Chen 0024, Shuming Ma, Dongdong Zhang 0001, Furu Wei, Baobao Chang
NeurIPS2
2023 Language Is Not All You Need: Aligning Perception with Language Models
abstract
A big convergence of language, multimodal perception, action, and world modeling is a key step toward artificial general intelligence. In this work, we introduce KOSMOS-1, a Multimodal Large Language Model (MLLM) that can perceive general modalities, learn in context (i.e., few-shot), and follow instructions (i.e., zero-shot). Specifically, we train KOSMOS-1 from scratch on web-scale multi-modal corpora, including arbitrarily interleaved text and images, image-caption pairs, and text data. We evaluate various settings, including zero-shot, few-shot, and multimodal chain-of-thought prompting, on a wide range of tasks without any gradient updates or finetuning. Experimental results show that KOSMOS-1 achieves impressive performance on (i) language understanding, generation, and even OCR-free NLP (directly fed with document images), (ii) perception-language tasks, including multimodal dialogue, image captioning, visual question answering, and (iii) vision tasks, such as image recognition with descriptions (specifying classification via text instructions). We also show that MLLMs can benefit from cross-modal transfer, i.e., transfer knowledge from language to multimodal, and from multimodal to language. In addition, we introduce a dataset of Raven IQ test, which diagnoses the nonverbal reasoning capability of MLLMs.
Shaohan Huang, Li Dong 0004, Wenhui Wang 0003, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui 0001, Owais Khan Mohammed, Barun Patra, Kriti Aggarwal, Zewen Chi, Johan Bjorck, Vishrav Chaudhary, Subhojit Som, Furu Wei
NeurIPS6
2023 GTrans: Grouping and Fusing Transformer Layers for Neural Machine Translation
abstract
Transformer structure, stacked by a sequence of encoder and decoder network layers, achieves significant development in neural machine translation. However, vanilla Transformer mainly exploits the top-layer representation, assuming the lower layers provide trivial or redundant information and thus ignoring the bottom-layer feature that is potentially valuable. In this work, we propose theGroup-Transformer model (GTrans) that flexibly divides multi-layer representations of both encoder and decoder into different groups and then fuses these group features to generate target words. To corroborate the effectiveness of the proposed method, extensive experiments and analytic experiments are conducted on three bilingual translation benchmarks and three multilingual translation tasks, including the IWLST-14, IWLST-17, LDC, WMT-14, WMT-21 and OPUS-100 benchmark. Experimental and analytical results demonstrate that our model outperforms its Transformer counterparts by a consistent gain. Furthermore, it can be successfully scaled up to 60 encoder layers and 36 decoder layers.
Jian Yang 0030, Yuwei Yin, Liqun Yang, Shuming Ma, Haoyang Huang, Dongdong Zhang 0001, Furu Wei, Zhoujun Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Towards Making the Most of Cross-Lingual Transfer for Zero-Shot Neural Machine Translation
abstract
Guanhua Chen, Shuming Ma, Yun Chen, Dongdong Zhang, Jia Pan, Wenping Wang, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Guanhua Chen 0001, Shuming Ma, Yun Chen 0007, Dongdong Zhang 0001, Jia Pan 0001, Wenping Wang 0001, Furu Wei
ACL (1)2
2022 XLM-E: Cross-lingual Language Model Pre-training via ELECTRA
abstract
Zewen Chi, Shaohan Huang, Li Dong, Shuming Ma, Bo Zheng, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Zewen Chi, Shaohan Huang, Li Dong 0004, Shuming Ma, Bo Zheng 0010, Saksham Singhal, Payal Bajaj, Xianling Mao, Heyan Huang, Furu Wei
ACL (1)4
2022 StableMoE: Stable Routing Strategy for Mixture of Experts
abstract
The Mixture-of-Experts (MoE) technique can scale up the model size of Transformers with an affordable computational overhead.We point out that existing learning-to-route MoE methods suffer from the routing fluctuation issue, i.e., the target expert of the same input may change along with training, but only one expert will be activated for the input during inference.The routing fluctuation tends to harm sample efficiency because the same input updates different experts but only one is finally used.In this paper, we propose STABLEMOE with two training stages to address the routing fluctuation problem.In the first training stage, we learn a balanced and cohesive routing strategy and distill it into a lightweight router decoupled from the backbone model.In the second training stage, we utilize the distilled router to determine the token-to-expert assignment and freeze it for a stable routing strategy.We validate our method on language modeling and multilingual machine translation.The results show that STABLEMOE outperforms existing MoE methods in terms of both convergence speed and performance.
Damai Dai, Li Dong 0004, Shuming Ma, Bo Zheng 0010, Zhifang Sui, Baobao Chang, Furu Wei
ACL (1)3
2022 PAEG: Phrase-level Adversarial Example Generation for Neural Machine Translation
abstract
While end-to-end neural machine translation (NMT) has achieved impressive progress, noisy input usually leads models to become fragile and unstable. Generating adversarial examples as the augmented data has been proved to be useful to alleviate this problem. Existing methods for adversarial example generation (AEG) are word-level or character-level, which ignore the ubiquitous phrase structure. In this paper, we propose a Phrase-level Adversarial Example Generation (PAEG) framework to enhance the robustness of the translation model. Our method further improves the gradient-based word-level AEG method by adopting a phrase-level substitution strategy. We verify our method on three benchmarks, including LDC Chinese-English, IWSLT14 German-English, and WMT14 English-German tasks. Experimental results demonstrate that our approach significantly improves translation performance and robustness to noise compared to previous strong baselines.
Juncheng Wan, Jian Yang 0030, Shuming Ma, Dongdong Zhang 0001, Weinan Zhang 0001, Yong Yu 0001, Zhoujun Li 0001
COLING3
2022 Zero-shot Cross-lingual Transfer of Prompt-based Tuning with a Unified Multilingual Prompt
abstract
Prompt-based tuning has been proven effective for pretrained language models (PLMs).While most of the existing work focuses on the monolingual prompts, we study the multilingual prompts for multilingual PLMs, especially in the zero-shot cross-lingual setting.To alleviate the effort of designing different prompts for multiple languages, we propose a novel model that uses a unified prompt for all languages, called UniPrompt.Different from the discrete prompts and soft prompts, the unified prompt is model-based and languageagnostic.Specifically, the unified prompt is initialized by a multilingual PLM to produce language-independent representation, after which is fused with the text input.During inference, the prompts can be pre-computed so that no extra computation cost is needed.To collocate with the unified prompt, we propose a new initialization method for the target label word to further improve the model's transferability across languages.Extensive experiments show that our proposed methods can significantly outperform the strong baselines across different languages.We release data and code to facilitate future research 1 .
Lianzhe Huang, Shuming Ma, Dongdong Zhang 0001, Furu Wei, Houfeng Wang
EMNLP2
2022 A Unified Strategy for Multilingual Grammatical Error Correction with Pre-trained Cross-Lingual Language Model
abstract
Synthetic data construction of Grammatical Error Correction (GEC) for non-English languages relies heavily on human-designed and language-specific rules, which produce limited error-corrected patterns. In this paper, we propose a generic and language-independent strategy for multilingual GEC, which can train a GEC system effectively for a new non-English language with only two easy-to-access resources: 1) a pre-trained cross-lingual language model (PXLM) and 2) parallel translation data between English and the language. Our approach creates diverse parallel GEC data without any language-specific operations by taking the non-autoregressive translation generated by PXLM and the gold translation as error-corrected sentence pairs. Then, we reuse PXLM to initialize the GEC model and pre-train it with the synthetic data generated by itself, which yields further improvement. We evaluate our approach on three public benchmarks of GEC in different languages. It achieves the state-of-the-art results on the NLPCC 2018 Task 2 dataset (Chinese) and obtains competitive performance on Falko-Merlin (German) and RULEC-GEC (Russian). Further analysis demonstrates that our data construction method is complementary to rule-based approaches.
Xin Sun 0013, Tao Ge 0001, Shuming Ma, Jingjing Li 0007, Furu Wei, Houfeng Wang
IJCAI3
2022 High-resource Language-specific Training for Multilingual Neural Machine Translation
abstract
Multilingual neural machine translation (MNMT) trained in multiple language pairs has attracted considerable attention due to fewer model parameters and lower training costs by sharing knowledge among multiple languages. Nonetheless, multilingual training is plagued by language interference degeneration in shared parameters because of the negative interference among different translation directions, especially on high-resource languages. In this paper, we propose the multilingual translation model with the high-resource language-specific training (HLT-MT) to alleviate the negative interference, which adopts the two-stage training with the language-specific selection mechanism. Specifically, we first train the multilingual model only with the high-resource pairs and select the language-specific modules at the top of the decoder to enhance the translation quality of high-resource directions. Next, the model is further trained on all available corpora to transfer knowledge from high-resource languages (HRLs) to low-resource languages (LRLs). Experimental results show that HLT-MT outperforms various strong baselines on WMT-10 and OPUS-100 benchmarks. Furthermore, the analytic experiments validate the effectiveness of our method in mitigating the negative interference in multilingual training.
Jian Yang 0030, Yuwei Yin, Shuming Ma, Dongdong Zhang 0001, Zhoujun Li 0001, Furu Wei
IJCAI3
2022 UM4: Unified Multilingual Multiple Teacher-Student Model for Zero-Resource Neural Machine Translation
abstract
Most translation tasks among languages belong to the zero-resource translation problem where parallel corpora are unavailable. Multilingual neural machine translation (MNMT) enables one-pass translation using shared semantic space for all languages compared to the two-pass pivot translation but often underperforms the pivot-based method. In this paper, we propose a novel method, named as Unified Multilingual Multiple teacher-student Model for NMT (UM4). Our method unifies source-teacher, target-teacher, and pivot-teacher models to guide the student model for the zero-resource translation. The source teacher and target teacher force the student to learn the direct source-target translation by the distilled knowledge on both source and target sides. The monolingual corpus is further leveraged by the pivot-teacher model to enhance the student model. Experimental results demonstrate that our model of 72 directions significantly outperforms previous methods on the WMT benchmark.
Jian Yang 0030, Yuwei Yin, Shuming Ma, Dongdong Zhang 0001, Shuangzhi Wu, Hongcheng Guo, Zhoujun Li 0001, Furu Wei
IJCAI3
2022 BlonDe: An Automatic Evaluation Metric for Document-level Machine Translation
abstract
Yuchen Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang, Jian Yang, Haoyang Huang, Rico Sennrich, Ryan Cotterell, Mrinmaya Sachan, Ming Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Tianyu Liu 0004, Shuming Ma, Dongdong Zhang 0001, Jian Yang 0030, Haoyang Huang, Rico Sennrich, Ryan Cotterell, Mrinmaya Sachan, Ming Zhou 0001
NAACL-HLT3
2022 On the Representation Collapse of Sparse Mixture of Experts
abstract
Sparse mixture of experts provides larger model capacity while requiring a constant computational overhead. It employs the routing mechanism to distribute input tokens to the best-matched experts according to their hidden representations. However, learning such a routing mechanism encourages token clustering around expert centroids, implying a trend toward representation collapse. In this work, we propose to estimate the routing scores between tokens and experts on a low-dimensional hypersphere. We conduct extensive experiments on cross-lingual language model pre-training and fine-tuning on downstream tasks. Experimental results across seven multilingual benchmarks show that our method achieves consistent gains. We also present a comprehensive analysis on the representation and routing behaviors of our models. Our method alleviates the representation collapse issue and achieves more consistent routing than the baseline mixture-of-experts methods.
Zewen Chi, Li Dong 0004, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xianling Mao, Heyan Huang, Furu Wei
NeurIPS5
2021 Zero-Shot Cross-Lingual Transfer of Neural Machine Translation with Multilingual Pretrained Encoders
abstract
Previous work mainly focuses on improving cross-lingual transfer for NLU tasks with a multilingual pretrained encoder (MPE), or improving the performance on supervised machine translation with BERT.However, it is under-explored that whether the MPE can help to facilitate the cross-lingual transferability of NMT model.In this paper, we focus on a zero-shot cross-lingual transfer task in NMT.In this task, the NMT model is trained with parallel dataset of only one language pair and an off-the-shelf MPE, then it is directly tested on zero-shot language pairs.We propose SixT, a simple yet effective model for this task.SixT leverages the MPE with a two-stage training schedule and gets further improvement with a position disentangled encoder and a capacity-enhanced decoder.Using this method, SixT significantly outperforms mBART, a pretrained multilingual encoderdecoder model explicitly designed for NMT, with an average improvement of 7.1 BLEU on zero-shot any-to-English test sets across 14 source languages.Furthermore, with much less training computation cost and training data, our model achieves better performance on 15 any-to-English test sets than CRISS and m2m-100, two strong multilingual NMT baselines.
Guanhua Chen 0001, Shuming Ma, Yun Chen 0007, Li Dong 0004, Dongdong Zhang 0001, Jia Pan 0001, Wenping Wang 0001, Furu Wei
EMNLP (1)2
2021 mT6: Multilingual Pretrained Text-to-Text Transformer with Translation Pairs
abstract
Zewen Chi, Li Dong, Shuming Ma, Shaohan Huang, Saksham Singhal, Xian-Ling Mao, Heyan Huang, Xia Song, Furu Wei. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Zewen Chi, Li Dong 0004, Shuming Ma, Shaohan Huang, Saksham Singhal, Xianling Mao, Heyan Huang, Furu Wei
EMNLP (1)3
2021 Smart-Start Decoding for Neural Machine Translation
abstract
Jian Yang, Shuming Ma, Dongdong Zhang, Juncheng Wan, Zhoujun Li, Ming Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Jian Yang 0030, Shuming Ma, Dongdong Zhang 0001, Juncheng Wan, Zhoujun Li 0001, Ming Zhou 0001
NAACL-HLT2
2021 Learning to Select Relevant Knowledge for Neural Machine Translation
Jian Yang 0030, Juncheng Wan, Shuming Ma, Haoyang Huang, Dongdong Zhang 0001, Yong Yu 0001, Zhoujun Li 0001, Furu Wei
NLPCC (1)3
2020 Alternating Language Modeling for Cross-Lingual Pre-Training
abstract
Language model pre-training has achieved success in many natural language processing tasks. Existing methods for cross-lingual pre-training adopt Translation Language Model to predict masked words with the concatenation of the source sentence and its target equivalent. In this work, we introduce a novel cross-lingual pre-training method, called Alternating Language Modeling (ALM). It code-switches sentences of different languages rather than simple concatenation, hoping to capture the rich cross-lingual context of words and phrases. More specifically, we randomly substitute source phrases with target translations to create code-switched sentences. Then, we use these code-switched data to train ALM model to learn to predict words of different languages. We evaluate our pre-training ALM on the downstream tasks of machine translation and cross-lingual classification. Experiments show that ALM can outperform the previous pre-training methods on three benchmarks.1
Jian Yang 0030, Shuming Ma, Dongdong Zhang 0001, Shuangzhi Wu, Zhoujun Li 0001, Ming Zhou 0001
AAAI2
2020 A Simple and Effective Unified Encoder for Document-Level Machine Translation
abstract
Most of the existing models for documentlevel machine translation adopt dual-encoder structures.The representation of the source sentences and the document-level contexts 1 are modeled with two separate encoders.Although these models can make use of the document-level contexts, they do not fully model the interaction between the contexts and the source sentences, and can not directly adapt to the recent pre-training models (e.g., BERT) which encodes multiple sentences with a single encoder.In this work, we propose a simple and effective unified encoder that can outperform the baseline models of dualencoder models in terms of BLEU and ME-TEOR scores.Moreover, the pre-training models can further boost the performance of our proposed model.
Shuming Ma, Dongdong Zhang 0001, Ming Zhou 0001
ACL1
2020 Improving Neural Machine Translation with Soft Template Prediction
abstract
Although neural machine translation (NMT) has achieved significant progress in recent years, most previous NMT models only depend on the source text to generate translation.Inspired by the success of template-based and syntax-based approaches in other fields, we propose to use extracted templates from tree structures as soft target templates to guide the translation procedure.In order to learn the syntactic structure of the target sentences, we adopt the constituency-based parse tree to generate candidate templates.We incorporate the template information into the encoder-decoder framework to jointly utilize the templates and source text.Experiments show that our model significantly outperforms the baseline models on four benchmarks and demonstrate the effectiveness of soft target templates.
Jian Yang 0030, Shuming Ma, Dongdong Zhang 0001, Zhoujun Li 0001, Ming Zhou 0001
ACL2
2020 Multimodal Matching Transformer for Live Commenting
abstract
Automatic live commenting aims to provide real-time comments on videos for viewers. It encourages users engagement on online video sites, and is also a good benchmark for video-to-text generation. Recent work on this task adopts encoder-decoder models to generate comments. However, these methods do not model the interaction between videos and comments explicitly, so they tend to generate popular comments that are often irrelevant to the videos. In this work, we aim to improve the relevance between live comments and videos by modeling the cross-modal interactions among different modalities. To this end, we propose a multimodal matching transformer to capture the relationships among comments, vision, and audio. The proposed model is based on the transformer framework and can iteratively learn the attention-aware representations for each modality. We evaluate the model on a publicly available live commenting dataset. Experiments show that the multimodal matching transformer model outperforms the state-of-the-art methods.
Chaoqun Duan, Lei Cui 0001, Shuming Ma, Furu Wei, Conghui Zhu, Tiejun Zhao
ECAI3
2020 Training Simplification and Model Simplification for Deep Learning : A Minimal Effort Back Propagation Method
abstract
We propose a simple yet effective technique to simplify the training and the resulting model of neural networks. In back propagation, only a small subset of the full gradient is computed to update the model parameters. The gradient vectors are sparsified in such a way that only the top-k elements (in terms of magnitude) are kept. As a result, only k rows or columns (depending on the layout) of the weight matrix are modified, leading to a linear reduction in the computational cost. Based on the sparsified gradients, we further simplify the model by eliminating the rows or columns that are seldom updated, which will reduce the computational cost both in the training and decoding, and potentially accelerate decoding in real-world applications. Surprisingly, experimental results demonstrate that most of the time we only need to update fewer than 5 percent of the weights at each back propagation pass. More interestingly, the accuracy of the resulting models is actually improved rather than degraded, and a detailed analysis is given. The model simplification results show that we could adaptively simplify the model which could often be reduced by around 9x, without any loss on accuracy or even with improved accuracy.
Xu Sun 0001, Xuancheng Ren, Shuming Ma, Bingzhen Wei, Wei Li 0101, Jingjing Xu 0001, Houfeng Wang, Yi Zhang 0050
IEEE Trans. Knowl. Data Eng.3
2019 Hierarchical Encoder with Auxiliary Supervision for Neural Table-to-Text Generation: Learning Better Representation for Tables
abstract
Generating natural language descriptions for the structured tables which consist of multiple attribute-value tuples is a convenient way to help people to understand the tables. Most neural table-to-text models are based on the encoder-decoder framework. However, it is hard for a vanilla encoder to learn the accurate semantic representation of a complex table. The challenges are two-fold: firstly, the table-to-text datasets often contain large number of attributes across different domains, thus it is hard for the encoder to incorporate these heterogeneous resources. Secondly, the single encoder also has difficulties in modeling the complex attribute-value structure of the tables. To this end, we first propose a two-level hierarchical encoder with coarse-to-fine attention to handle the attribute-value structure of the tables. Furthermore, to capture the accurate semantic representations of the tables, we propose 3 joint tasks apart from the prime encoder-decoder learning, namely auxiliary sequence labeling task, text autoencoder and multi-labeling classification, as the auxiliary supervisions for the table encoder. We test our models on the widely used dataset WIKIBIO which contains Wikipedia infoboxes and related descriptions. The dataset contains complex tables as well as large number of attributes across different domains. We achieve the state-of-the-art performance on both automatic and human evaluation metrics.
Tianyu Liu 0001, Fuli Luo, Qiaolin Xia, Shuming Ma, Baobao Chang, Zhifang Sui
AAAI4
2019 LiveBot: Generating Live Video Comments Based on Visual and Textual Contexts
abstract
We introduce the task of automatic live commenting. Live commenting, which is also called “video barrage”, is an emerging feature on online video sites that allows real-time comments from viewers to fly across the screen like bullets or roll at the right side of the screen. The live comments are a mixture of opinions for the video and the chit chats with other comments. Automatic live commenting requires AI agents to comprehend the videos and interact with human viewers who also make the comments, so it is a good testbed of an AI agent’s ability to deal with both dynamic vision and language. In this work, we construct a large-scale live comment dataset with 2,361 videos and 895,929 live comments. Then, we introduce two neural models to generate live comments based on the visual and textual contexts, which achieve better performance than previous neural baselines such as the sequence-to-sequence model. Finally, we provide a retrieval-based evaluation protocol for automatic live commenting where the model is asked to sort a set of candidate comments based on the log-likelihood score, and evaluated on metrics such as mean-reciprocal-rank. Putting it all together, we demonstrate the first “LiveBot”. The datasets and the codes can be found at https://github.com/lancopku/livebot.
Shuming Ma, Lei Cui 0001, Damai Dai, Furu Wei, Xu Sun 0001
AAAI1
2019 Key Fact as Pivot: A Two-Stage Model for Low Resource Table-to-Text Generation
abstract
Table -to-text generation aims to translate the structured data into the unstructured text.Most existing methods adopt the encoder-decoder framework to learn the transformation, which requires large-scale training samples.However, the lack of large parallel data is a major practical problem for many domains.In this work, we consider the scenario of low resource table-to-text generation, where only limited parallel data is available.We propose a novel model to separate the generation into two stages: key fact prediction and surface realization.It first predicts the key facts from the tables, and then generates the text with the key facts.The training of key fact prediction needs much fewer annotated data, while surface realization can be trained with pseudo parallel corpus.We evaluate our model on a biography generation dataset.Our model can achieve 27.34 BLEU score with only 1, 000 parallel data, while the baseline model only obtain the performance of 9.71 BLEU score. 1
Shuming Ma, Tianyu Liu 0001, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001
ACL (1)1
2019 A Deep Reinforced Sequence-to-Set Model for Multi-Label Classification
abstract
Multi-label classification (MLC) aims to predict a set of labels for a given instance.Based on a pre-defined label order, the sequence-tosequence (Seq2Seq) model trained via maximum likelihood estimation method has been successfully applied to the MLC task and shows powerful ability to capture high-order correlations between labels.However, the output labels are essentially an unordered set rather than an ordered sequence.This inconsistency tends to result in some intractable problems, e.g., sensitivity to the label order.To remedy this, we propose a simple but effective sequence-to-set model.The proposed model is trained via reinforcement learning, where reward feedback is designed to be independent of the label order.In this way, we can reduce the dependence of the model on the label order, as well as capture high-order correlations between labels.Extensive experiments show that our approach can substantially outperform competitive baselines, as well as effectively reduce the sensitivity to the label order. 1
Fuli Luo, Shuming Ma, Junyang Lin, Xu Sun 0001
ACL (1)3
2019 Predicting Popular News Comments Based on Multi-Target Text Matching Model
Deli Chen, Shuming Ma, Qi Su 0001
NLPCC (1)2
2019 Towards easier and faster sequence labeling for natural language processing: A search-based probabilistic online learning framework (SAPO)
Xu Sun 0001, Shuming Ma, Yi Zhang 0050, Xuancheng Ren
Inf. Sci.2
2018 Deconvolution-Based Global Decoding for Neural Machine Translation
abstract
A great proportion of sequence-to-sequence (Seq2Seq) models for Neural Machine Translation (NMT) adopt Recurrent Neural Network (RNN) to generate translation word by word following a sequential order. As the studies of linguistics have proved that language is not linear word sequence but sequence of complex structure, translation at each step should be conditioned on the whole target-side context. To tackle the problem, we propose a new NMT model that decodes the sequence with the guidance of its structural prediction of the context of the target sequence. Our model generates translation based on the structural prediction of the target-side context so that the translation can be freed from the bind of sequential order. Experimental results demonstrate that our model is more competitive compared with the state-of-the-art methods, and the analysis reflects that our model is also robust to translating sentences of different lengths and it also reduces repetition with the instruction from the target-side context for decoding.
Junyang Lin, Xu Sun 0001, Xuancheng Ren, Shuming Ma, Jinsong Su, Qi Su 0001
COLING4
2018 A Neural Question Answering Model Based on Semi-Structured Tables
abstract
Most question answering (QA) systems are based on raw text and structured knowledge graph. However, raw text corpora are hard for QA system to understand, and structured knowledge graph needs intensive manual work, while it is relatively easy to obtain semi-structured tables from many sources directly, or build them automatically. In this paper, we build an end-to-end system to answer multiple choice questions with semi-structured tables as its knowledge. Our system answers queries by two steps. First, it finds the most similar tables. Then the system measures the relevance between each question and candidate table cells, and choose the most related cell as the source of answer. The system is evaluated with TabMCQ dataset, and gets a huge improvement compared to the state of the art.
Xiaodong Zhang 0022, Shuming Ma, Xu Sun 0001, Houfeng Wang, Mengxiang Wang
COLING3
2018 SGM: Sequence Generation Model for Multi-label Classification
abstract
Multi-label classification is an important yet challenging task in natural language processing. It is more complex than single-label classification in that the labels tend to be correlated. Existing methods tend to ignore the correlations between labels. Besides, different parts of the text can contribute differently for predicting different labels, which is not considered by existing models. In this paper, we propose to view the multi-label classification task as a sequence generation problem, and apply a sequence generation model with a novel decoder structure to solve it. Extensive experimental results show that our proposed methods outperform previous work by a substantial margin. Further analysis of experimental results demonstrates that the proposed methods not only capture the correlations between labels, but also select the most informative words automatically when predicting different labels.
Xu Sun 0001, Wei Li 0101, Shuming Ma, Wei Wu 0044, Houfeng Wang
COLING4
2018 Does Higher Order LSTM Have Better Accuracy for Segmenting and Labeling Sequence Data?
abstract
Existing neural models usually predict the tag of the current token independent of the neighboring tags. The popular LSTM-CRF model considers the tag dependencies between every two consecutive tags. However, it is hard for existing neural models to take longer distance dependencies between tags into consideration. The scalability is mainly limited by the complex model structures and the cost of dynamic programming during training. In our work, we first design a new model called “high order LSTM” to predict multiple tags for the current token which contains not only the current tag but also the previous several tags. We call the number of tags in one prediction as “order”. Then we propose a new method called Multi-Order BiLSTM (MO-BiLSTM) which combines low order and high order LSTMs together. MO-BiLSTM keeps the scalability to high order models with a pruning technique. We evaluate MO-BiLSTM on all-phrase chunking and NER datasets. Experiment results show that MO-BiLSTM achieves the state-of-the-art result in chunking and highly competitive results in two NER datasets.
Yi Zhang 0050, Xu Sun 0001, Shuming Ma, Yang Yang 0125, Xuancheng Ren
COLING3
2018 Semantic-Unit-Based Dilated Convolution for Multi-Label Text Classification
abstract
We propose a novel model for multi-label text classification, which is based on sequenceto-sequence learning.The model generates higher-level semantic unit representations with multi-level dilated convolution as well as a corresponding hybrid attention mechanism that extracts both the information at the word-level and the level of the semantic unit.Our designed dilated convolution effectively reduces dimension and supports an exponential expansion of receptive fields without loss of local information, and the attention-overattention mechanism is able to capture more summary relevant information from the source context.Results of our experiments show that the proposed model has significant advantages over the baseline models on the dataset RCV1-V2 and Ren-CECps, and our analysis demonstrates that our model is competitive to the deterministic hierarchical models and it is more robust to classifying low-frequency labels 1 .
Junyang Lin, Qi Su 0001, Shuming Ma, Xu Sun 0001
EMNLP4
2018 Phrase-level Self-Attention Networks for Universal Sentence Encoding
abstract
Universal sentence encoding is a hot topic in recent NLP research.Attention mechanism has been an integral part in many sentence encoding models, allowing the models to capture context dependencies regardless of the distance between elements in the sequence.Fully attention-based models have recently attracted enormous interest due to their highly parallelizable computation and significantly less training time.However, the memory consumption of their models grows quadratically with sentence length, and the syntactic information is neglected.To this end, we propose Phrase-level Self-Attention Networks (PSAN) that perform self-attention across words inside a phrase to capture context dependencies at the phrase level, and use the gated memory updating mechanism to refine each word's representation hierarchically with longer-term context dependencies captured in a larger phrase.As a result, the memory consumption can be reduced because the self-attention is performed at the phrase level instead of the sentence level.At the same time, syntactic information can be easily integrated in the model.Experiment results show that PSAN can achieve the state-ofthe-art transfer performance across a plethora of NLP tasks including sentence classification, natural language inference and sentence textual similarity.
Wei Wu 0044, Houfeng Wang, Tianyu Liu 0001, Shuming Ma
EMNLP4
2018 A Hierarchical End-to-End Model for Jointly Improving Text Summarization and Sentiment Classification
abstract
Text summarization and sentiment classification both aim to capture the main ideas of the text but at different levels. Text summarization is to describe the text within a few sentences, while sentiment classification can be regarded as a special type of summarization which ``summarizes'' the text into a even more abstract fashion, i.e., a sentiment class. Based on this idea, we propose a hierarchical end-to-end model for joint learning of text summarization and sentiment classification, where the sentiment classification label is treated as the further ``summarization'' of the text summarization output. Hence, the sentiment classification layer is put upon the text summarization layer, and a hierarchical structure is derived. Experimental results on Amazon online reviews datasets show that our model achieves better performance than the strong baseline systems on both abstractive summarization and sentiment classification.
Shuming Ma, Xu Sun 0001, Junyang Lin, Xuancheng Ren
IJCAI1
2018 Query and Output: Generating Words by Querying Distributed Word Representations for Paraphrase Generation
abstract
Shuming Ma, Xu Sun, Wei Li, Sujian Li, Wenjie Li, Xuancheng Ren. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Shuming Ma, Xu Sun 0001, Wei Li 0101, Sujian Li, Wenjie Li 0002, Xuancheng Ren
NAACL-HLT1
2018 Accelerating Graph-Based Dependency Parsing with Lock-Free Parallel Perceptron
Shuming Ma, Xu Sun 0001, Yi Zhang 0050, Bingzhen Wei
NLPCC (1)1
2017 meProp: Sparsified Back Propagation for Accelerated Deep Learning with Reduced Overfitting
abstract
We propose a simple yet effective technique for neural network learning. The forward propagation is computed as usual. In back propagation, only a small subset of the full gradient is computed to update the model parameters. The gradient vectors are sparsified in such a way that only the top-$k$ elements (in terms of magnitude) are kept. As a result, only $k$ rows or columns (depending on the layout) of the weight matrix are modified, leading to a linear reduction ($k$ divided by the vector dimension) in the computational cost. Surprisingly, experimental results demonstrate that we can update only 1–4\% of the weights at each back propagation pass. This does not result in a larger number of training iterations. More interestingly, the accuracy of the resulting models is actually improved rather than degraded, and a detailed analysis is given.
Xu Sun 0001, Xuancheng Ren, Shuming Ma, Houfeng Wang
ICML3
2017 Transfer Deep Learning for Low-Resource Chinese Word Segmentation with a Novel Neural Network
Jingjing Xu 0001, Shuming Ma, Yi Zhang 0050, Bingzhen Wei, Xiaoyan Cai, Xu Sun 0001
NLPCC2