Ruotian Ma

dblp:246/3164 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
abstract
Yue Wang, Ruotian Ma, Xingyu Chen, Zhengliang Shi, Morunliu Yang, Wanshun Chen, Huang Liu, Jiadi Yao, Xin He, Qu Yang, Qingxuan Jiang, Fanghua Ye, Juntao Li, Zhaopeng Tu, Xiaolong Li, Liefeng Bo, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yue Wang 0039, Ruotian Ma, Zhengliang Shi, Morunliu Yang, Wanshun Chen, Huang Liu, Jiadi Yao, Qu Yang, Qingxuan Jiang, Fanghua Ye 0004, Juntao Li 0005, Zhaopeng Tu, Liefeng Bo, Min Zhang 0005
ACL (1)2
2025 S²R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
abstract
Recent studies have demonstrated the effectiveness of LLM test-time scaling. However, existing approaches to incentivize LLMs’ deep thinking abilities generally require large-scale data or significant training efforts. Meanwhile, it remains unclear how to improve the thinking abilities of less powerful base models. In this work, we introduce S^2R, an efficient framework that enhances LLM reasoning by teaching models to self-verify and self-correct during inference. Specifically, we first initialize LLMs with iterative self-verification and self-correction behaviors through supervised fine-tuning on carefully curated data. The self-verification and self-correction skills are then further strengthened by outcome-level and process-level reinforcement learning with minimized resource requirements. Our results demonstrate that, with only 3.1k behavior initialization samples, Qwen2.5-math-7B achieves an accuracy improvement from 51.0% to 81.6%, outperforming models trained on an equivalent amount of long-CoT distilled data. We also discuss the effect of different RL strategies on enhancing LLMs’ deep reasoning. Extensive experiments and analysis based on three base models across both in-domain and out-of-domain benchmarks validate the effectiveness of S^2R.
Ruotian Ma, Peisong Wang 0002, Xingyan Liu, Bang Zhang, Jia Li 0009
ACL (1)1
2025 SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
abstract
Evaluating the step-by-step reliability of large language model (LLM) reasoning, such as Chain-of-Thought, remains challenging due to the difficulty and cost of obtaining high-quality step-level supervision. In this paper, we introduce Self-Play Critic (SPC), a novel approach where a critic model evolves its ability to assess reasoning steps through adversarial self-play games, eliminating the need for manual step-level annotation. SPC involves fine-tuning two copies of a base model to play two roles, namely a "sneaky generator" that deliberately produces erroneous steps designed to be difficult to detect, and a "critic" that analyzes the correctness of reasoning steps. These two models engage in an adversarial game in which the generator aims to fool the critic, while the critic model seeks to identify the generator's errors. Using reinforcement learning based on the game outcomes, the models iteratively improve; the winner of each confrontation receives a positive reward and the loser receives a negative reward, driving continuous self-evolution. Experiments on three reasoning process benchmarks (ProcessBench, PRM800K, DeltaBench) demonstrate that our SPC progressively enhances its error detection capabilities (e.g., accuracy increases from 70.8% to 77.7% on ProcessBench) and surpasses strong baselines, including distilled R1 model. Furthermore, SPC can guide the test-time search of diverse LLMs and significantly improve their mathematical reasoning performance on MATH500 and AIME2024, surpassing those guided by state-of-the-art process reward models.
Bang Zhang, Ruotian Ma, Peisong Wang 0002, Xiaodan Liang, Zhaopeng Tu, Kwan-Yee Kenneth Wong
NeurIPS3
2025 Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
abstract
Mixture-of-Experts (MoE) architectures within Large Reasoning Models (LRMs) have achieved impressive reasoning capabilities by selectively activating experts to facilitate structured cognitive processes. Despite notable advances, existing reasoning models often suffer from cognitive inefficiencies like overthinking and underthinking. To address these limitations, we introduce a novel inference-time steering methodology called Reinforcing Cognitive Experts (RICE), designed to improve reasoning depth and efficiency without additional training or complex heuristics. Leveraging normalized Pointwise Mutual Information (nPMI), we systematically identify specialized experts, termed cognitive experts that orchestrate meta-level reasoning operations characterized by tokens like <think>. Empirical evaluations with leading MoE-based LRMs (DeepSeek-R1 and Qwen3-235B) on rigorous quantitative and scientific reasoning benchmarks (AIME and GPQA Diamond) demonstrate noticeable and consistent improvements in reasoning accuracy, cognitive efficiency, and cross-domain generalization. Crucially, our lightweight approach substantially outperforms prevalent reasoning-steering techniques, such as prompt design and decoding constraints, while preserving the model's general instruction-following skills. These results highlight reinforcing cognitive experts as a promising, practical, and interpretable direction to enhance cognitive efficiency within advanced reasoning models.
Yue Wang 0039, Zhiwei He 0002, Qiuzhi Liu, Yunzhi Yao, Wenxuan Wang 0001, Ruotian Ma, Haitao Mi, Ningyu Zhang 0001, Zhaopeng Tu, Dong Yu 0001
NeurIPS10
2024 Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs
abstract
Xin Zhou, Ping Nie, Yiwen Guo, Haojie Wei, Zhanqiu Zhang, Pasquale Minervini, Ruotian Ma, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Xin Zhou 0012, Ping Nie, Yiwen Guo, Haojie Wei, Zhanqiu Zhang, Pasquale Minervini, Ruotian Ma, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP7
2023 Learning "O" Helps for Learning More: Handling the Unlabeled Entity Problem for Class-incremental NER
abstract
Ruotian Ma, Xuanting Chen, Zhang Lin, Xin Zhou, Junzhe Wang, Tao Gui, Qi Zhang, Xiang Gao, Yun Wen Chen. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ruotian Ma, Xuanting Chen, Xin Zhou 0012, Junzhe Wang 0001, Tao Gui, Qi Zhang 0001, Xiang Gao 0017, Yun Wen Chen
ACL (1)1
2023 Towards Building More Robust NER datasets: An Empirical Study on NER Dataset Bias from a Dataset Difficulty View
abstract
Recently, many studies have illustrated the robustness problem of Named Entity Recognition (NER) systems: the NER models often rely on superficial entity patterns for predictions, without considering evidence from the context.Consequently, even state-of-theart NER models generalize poorly to outof-domain scenarios when out-of-distribution (OOD) entity patterns are introduced.Previous research attributes the robustness problem to the existence of NER dataset bias, where simpler and regular entity patterns induce shortcut learning.In this work, we bring new insights into this problem by comprehensively investigating the NER dataset bias from a dataset difficulty view.We quantify the entity-context difficulty distribution in existing datasets and explain their relationship with model robustness.Based on our findings, we explore three potential ways to de-bias the NER datasets by altering entity-context distribution, and we validate the feasibility with intensive experiments.Finally, we show that the de-biased datasets can transfer to different models and even benefit existing model-based robustness-improving methods, indicating that building more robust datasets is fundamental for building more robust NER systems. * Equal contribution. † Corresponding authors.Ms. Hall expects Cathay 's profit to grow around 13 % annually this year.
Ruotian Ma, Xin Zhou 0012, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP1
2022 Making Parameter-efficient Tuning More Efficient: A Unified Framework for Classification Tasks
abstract
Large pre-trained language models (PLMs) have demonstrated superior performance in industrial applications. Recent studies have explored parameter-efficient PLM tuning, which only updates a small amount of task-specific parameters while achieving both high efficiency and comparable performance against standard fine-tuning. However, all these methods ignore the inefficiency problem caused by the task-specific output layers, which is inflexible for us to re-use PLMs and introduces non-negligible parameters. In this work, we focus on the text classification task and propose plugin-tuning, a framework that further improves the efficiency of existing parameter-efficient methods with a unified classifier. Specifically, we re-formulate both token and sentence classification tasks into a unified language modeling task, and map label spaces of different tasks into the same vocabulary space. In this way, we can directly re-use the language modeling heads of PLMs, avoiding introducing extra parameters for different tasks. We conduct experiments on six classification benchmarks. The experimental results show that plugin-tuning can achieve comparable performance against fine-tuned PLMs, while further saving around 50% parameters on top of other parameter-efficient methods.
Xin Zhou 0012, Ruotian Ma, Yicheng Zou, Xuanting Chen, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Rui Xie 0005, Wei Wu 0014
COLING2
2022 Cross-Linguistic Syntactic Difference in Multilingual BERT: How Good is It and How Does It Affect Transfer?
abstract
Multilingual BERT (mBERT) has demonstrated considerable cross-lingual syntactic ability, whereby it enables effective zero-shot cross-lingual transfer of syntactic knowledge.The transfer is more successful between some languages, but it is not well understood what leads to this variation and whether it fairly reflects difference between languages.In this work, we investigate the distributions of grammatical relations induced from mBERT in the context of 24 typologically different languages.We demonstrate that the distance between the distributions of different languages is highly consistent with the syntactic difference in terms of linguistic formalisms.Such difference learnt via self-supervision plays a crucial role in the zero-shot transfer performance and can be predicted by variation in morphosyntactic properties between languages.These results suggest that mBERT properly encodes languages in a way consistent with linguistic diversity and provide insights into the mechanism of crosslingual transfer.
Ningyu Xu, Tao Gui, Ruotian Ma, Qi Zhang 0001, Jingting Ye, Menghan Zhang, Xuanjing Huang 0001
EMNLP3
2022 TextFusion: Privacy-Preserving Pre-trained Model Inference via Token Fusion
abstract
Xin Zhou, Jinzhu Lu, Tao Gui, Ruotian Ma, Zichu Fei, Yuran Wang, Yong Ding, Yibo Cheung, Qi Zhang, Xuanjing Huang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Xin Zhou 0012, Jinzhu Lu, Tao Gui, Ruotian Ma, Zichu Fei, Yibo Cheung, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP4
2022 Searching for Optimal Subword Tokenization in Cross-domain NER
abstract
Input distribution shift is one of the vital problems in unsupervised domain adaptation (UDA). The most popular UDA approaches focus on domain-invariant representation learning, trying to align the features from different domains into a similar feature distribution. However, these approaches ignore the direct alignment of input word distributions between domains, which is a vital factor in word-level classification tasks such as cross-domain NER. In this work, we shed new light on cross-domain NER by introducing a subword-level solution, X-Piece, for input word-level distribution shift in NER. Specifically, we re-tokenize the input words of the source domain to approach the target subword distribution, which is formulated and solved as an optimal transport problem. As this approach focuses on the input level, it can also be combined with previous DIRL methods for further improvement. Experimental results show the effectiveness of the proposed method based on BERT-tagger on four benchmark NER datasets. Also, the proposed method is proved to benefit DIRL methods such as DANN.
Ruotian Ma, Yiding Tan, Xin Zhou 0012, Xuanting Chen, Di Liang, Wei Wu 0014, Tao Gui
IJCAI1
2022 Template-free Prompt Tuning for Few-shot NER
abstract
Ruotian Ma, Xin Zhou, Tao Gui, Yiding Tan, Linyang Li, Qi Zhang, Xuanjing Huang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Ruotian Ma, Xin Zhou 0012, Tao Gui, Yiding Tan, Linyang Li, Qi Zhang 0001, Xuanjing Huang 0001
NAACL-HLT1
2021 SENT: Sentence-level Distant Relation Extraction via Negative Training
abstract
Ruotian Ma, Tao Gui, Linyang Li, Qi Zhang, Xuanjing Huang, Yaqian Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ruotian Ma, Tao Gui, Linyang Li, Qi Zhang 0001, Xuanjing Huang 0001, Yaqian Zhou 0001
ACL/IJCNLP (1)1
2021 Backdoor Attacks on Pre-trained Models by Layerwise Weight Poisoning
abstract
Pre-Trained Models have been widely applied and recently proved vulnerable under backdoor attacks: the released pre-trained weights can be maliciously poisoned with certain triggers.When the triggers are activated, even the fine-tuned model will predict pre-defined labels, causing a security threat.These backdoors generated by the poisoning methods can be erased by changing hyper-parameters during fine-tuning or detected by finding the triggers.In this paper, we propose a stronger weight-poisoning attack method that introduces a layerwise weight poisoning strategy to plant deeper backdoors; we also introduce a combinatorial trigger that cannot be easily detected.The experiments on text classification tasks show that previous defense methods cannot resist our weight-poisoning method, which indicates that our method can be widely applied and may provide hints for future model robustness studies.
Linyang Li, Demin Song, Jiehang Zeng, Ruotian Ma, Xipeng Qiu
EMNLP (1)5
2020 Simplify the Usage of Lexicon in Chinese NER
abstract
Recently, many works have tried to augment the performance of Chinese named entity recognition (NER) using word lexicons.As a representative, Lattice-LSTM (Zhang and Yang, 2018) has achieved new benchmark results on several public Chinese NER datasets.However, Lattice-LSTM has a complex model architecture.This limits its application in many industrial areas where real-time NER responses are needed.In this work, we propose a simple but effective method for incorporating the word lexicon into the character representations.This method avoids designing a complicated sequence modeling architecture, and for any neural NER model, it requires only subtle adjustment of the character representation layer to introduce the lexicon information.Experimental studies on four benchmark Chinese NER datasets show that our method achieves an inference speed up to 6.15 times faster than those of state-ofthe-art methods, along with a better performance.The experimental results also show that the proposed method can be easily incorporated with pre-trained models like BERT. 1 * Equal contribution.
Ruotian Ma, Minlong Peng, Qi Zhang 0001, Zhongyu Wei, Xuanjing Huang 0001
ACL1
2020 BERT-ATTACK: Adversarial Attack Against BERT Using BERT
abstract
Adversarial attacks for discrete data (such as texts) have been proved significantly more challenging than continuous data (such as images) since it is difficult to generate adversarial samples with gradient-based methods.Current successful attack methods for texts usually adopt heuristic replacement strategies on the character or word level, which remains challenging to find the optimal solution in the massive space of possible combinations of replacements while preserving semantic consistency and language fluency.In this paper, we propose BERT-Attack, a high-quality and effective method to generate adversarial samples using pre-trained masked language models exemplified by BERT.We turn BERT against its fine-tuned models and other deep neural models in downstream tasks so that we can successfully mislead the target models to predict incorrectly.Our method outperforms state-of-theart attack strategies in both success rate and perturb percentage, while the generated adversarial samples are fluent and semantically preserved.Also, the cost of calculation is low, thus possible for large-scale generations.The code is available at https://github.com/ LinyangLee/BERT-Attack.
Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 0001, Xipeng Qiu
EMNLP (1)2
2019 CNN-Based Chinese NER with Lexicon Rethinking
abstract
Character-level Chinese named entity recognition (NER) that applies long short-term memory (LSTM) to incorporate lexicons has achieved great success. However, this method fails to fully exploit GPU parallelism and candidate lexicons can conflict. In this work, we propose a faster alternative to Chinese NER: a convolutional neural network (CNN)-based method that incorporates lexicons using a rethinking mechanism. The proposed method can model all the characters and potential words that match the sentence in parallel. In addition, the rethinking mechanism can address the word conflict by feeding back the high-level features to refine the networks. Experimental results on four datasets show that the proposed method can achieve better performance than both word-level and character-level baseline methods. In addition, the proposed method performs up to 3.21 times faster than state-of-the-art methods, while realizing better performance.
Tao Gui, Ruotian Ma, Qi Zhang 0001, Lujun Zhao, Yu-Gang Jiang 0001, Xuanjing Huang 0001
IJCAI2