Rui Men

dblp:170/0093 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-4429-3461ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Deep learning architectures and training · 35% Language models and text generation · 29% Representation and self-supervised learning · 14%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%
Computer networks
1 paper
Datacenter networks · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 21 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining
1.122022
OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework · ICML 2022
M6: Multi-Modality-to-Multi-Modality Multitask Mega-transformer for Unified Pretraining · KDD 2021
Natural language and speech › Language models and text generation › large language model reasoning
multi-turn reasoning
1.012026
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation · ACL (1) 2026
Natural language and speech › Language models and text generation › evaluation of language models
reasoning evaluation
1.012026
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation · ACL (1) 2026
Machine learning › Deep learning architectures and training
attention mechanism
0.912025
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free · NeurIPS 2025
Machine learning › Trustworthy machine learning › language model interpretability
attention sink
0.912025
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free · NeurIPS 2025
Machine learning › Deep learning architectures and training › attention mechanism › attention module
gated attention
0.912025
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free · NeurIPS 2025
Machine learning › Efficient and distributed learning › distributed training › large model training
mixture-of-experts training
0.912025
Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models · ACL (1) 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free · NeurIPS 2025
Distributed systems › distributed machine learning
distributed training
0.912025
Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization · HPCA 2025
Machine learning › Deep learning architectures and training › sequence modeling
sequence-to-sequence learning
0.612022
OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework · ICML 2022
Computer vision › Vision and language
vision-language model
0.612022
OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework · ICML 2022
Natural language and speech › Language models and text generation › large language model
chinese language model
0.512021
M6: Multi-Modality-to-Multi-Modality Multitask Mega-transformer for Unified Pretraining · KDD 2021
Machine learning › Representation and self-supervised learning › multimodal representation learning
cross-modal representation learning
0.512021
M6: Multi-Modality-to-Multi-Modality Multitask Mega-transformer for Unified Pretraining · KDD 2021
Natural language and speech › Language models and text generation
multilingual language models
0.512021
M6: Multi-Modality-to-Multi-Modality Multitask Mega-transformer for Unified Pretraining · KDD 2021
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
non-autoregressive generation
0.512021
UFC-BERT: Unifying Multi-Modal Controls for Conditional Image Synthesis · NeurIPS 2021
Information retrieval
cross-modal retrieval
0.512021
Learning Relation Alignment for Calibrated Cross-modal Retrieval · ACL/IJCNLP (1) 2021
Visual content generation and editing › image generation
conditional image synthesis
0.512021
UFC-BERT: Unifying Multi-Modal Controls for Conditional Image Synthesis · NeurIPS 2021
Visual content generation and editing
multimodal control
0.512021
UFC-BERT: Unifying Multi-Modal Controls for Conditional Image Synthesis · NeurIPS 2021
Natural language and speech › Language models and text generation
large language model training
0.312025
Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models · ACL (1) 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.312025
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free · NeurIPS 2025
Distributed systems
fault tolerance
0.312025
Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization · HPCA 2025

Methods — techniques the papers use, named apart from their topics

traffic planning · 1.7collective communication analysis · 1.7transformer · 1.5non-autoregressive generation · 1.0benchmarking · 1.0sigmoid gating · 0.9scaled dot-product attention · 0.9load-balancing loss · 0.9sequence-to-sequence · 0.6instruction-based learning · 0.6multimodal pretraining · 0.5cross-modal alignment · 0.5
YearPublicationVenuePosition
2026 MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
abstract
Xiaoyuan Li, Keqin Bao, Yubo Ma, Moxin Li, Wenjie Wang, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiaoyuan Li 0001, Keqin Bao, Yubo Ma, Moxin Li, Wenjie Wang 0007, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu
ACL (1)6
2026 A Lightweight Continuous Identity Authentication-Based Security Offloading Scheme in Vehicular Edge Computing
abstract
Vehicular Edge Computing (VEC) is pivotal for latency-sensitive vehicular applications but confronts three critical challenges: traditional one-time identity authentication cannot adapt to high mobility, leaving privacy and data security vulnerabilities, offloading system security levels lack quantifiability, and security-performance optimization objectives are inherently conflicting. To address these limitations, we propose a lightweight continuous identity authentication-based secure offloading scheme for VEC. First, a three-entity collaborative architecture is designed, which leverages chameleon hash function (CHF) to reduce vehicle-side signature overhead, and Bloom filter (BF) to enable real-time verification during vehicle-roadside unit (RSU) handovers. Second, a dual-dimensional security framework that quantifies authentication and data signature levels is established, enabling on-demand security adjustment for diverse tasks. Third, to balance task latency minimization and security maximization, Unlike prior works that optimize security and offloading separately, this framework unifies both identity authentication and data transmission security into a holistic latency-oriented offloading design, filling critical research gaps in high-mobility VEC scenarios. To tackle this intricate problem, we decompose it into four sub-problems. These sub-problems are solved using Lagrangian duality, the Newton-Raphson method, and the branch-and-bound algorithm to obtain stable and highquality feasible solutions efficiently. Extensive simulations against four baseline schemes demonstrate that the proposed approach achieves fast convergence and priority performance.
Rui Men, Axida Shan, Celimuge Wu, Jie Li 0002
IEEE Internet Things J.1
2025 Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
abstract
Zihan Qiu, Zeyu Huang, Bo Zheng, Kaiyue Wen, Zekun Wang, Rui Men, Ivan Titov, Dayiheng Liu, Jingren Zhou, Junyang Lin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zihan Qiu, Bo Zheng 0007, Kaiyue Wen, Rui Men, Ivan Titov 0001, Dayiheng Liu, Jingren Zhou 0001, Junyang Lin
ACL (1)6
2025 Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization
abstract
The emergence of Large Language Models (LLMs) has necessitated the adoption of distributed training techniques, involving the deployment of thousands of GPUs to train a single model. Unfortunately, the efficiency of large-scale distributed training systems is often suboptimal due to the increased likelihood of hardware errors in high-end GPU products and the heightened risk of network traffic collisions. Specifically, GPUs involved in the same job require periodic synchronization to exchange necessary data, such as gradients, parameters, or activations. As a result, any local hardware failure can disrupt training tasks, and the inability to swiftly identify faulty components leads to a significant waste of GPU resources. Moreover, prolonged communication due to traffic collisions can substantially increase GPU waiting times. To address these challenges, we propose a communication-driven solution, namely the C 4. The key insights of C 4 are twofold. First, the load in distributed training exhibits homogeneous characteristics and is divided into iterations through periodic synchronization, therefore hardware anomalies would incur certain syndrome in collective communication. By leveraging this feature, $\mathbf{C} 4$ can rapidly identify the faulty components, swiftly isolate the anomaly, and restart the task, thereby avoiding resource wastage caused by delays in anomaly detection. Second, the predictable communication model of collective communication, involving a limited number of long-lived flows, allows C 4 to efficiently execute traffic planning, substantially reducing bandwidth competition among these flows. The $\mathbf{C 4}$ has been extensively deployed across real-world production systems in a hyperscale cloud provider, yielding a significant improvement in system efficiency, from 30% to $\mathbf{4 5 \%}$. This enhancement is attributed to a $\mathbf{3 0 \%}$ reduction in error-induced overhead and a 15% reduction in communication costs.
Jianbo Dong, Yikai Zhu, Hairong Jiao, Ennan Zhai, Wencong Xiao, Man Yuan, Siran Yang, Jiamang Wang, Rui Men, Dennis Cai, Binzhang Fu
HPCA20
2025 Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
abstract
Gating mechanisms have been widely utilized, from early models like LSTMs and Highway Networks to recent state space models, linear attention, and also softmax attention. Yet, existing literature rarely examines the specific effects of gating. In this work, we conduct comprehensive experiments to systematically investigate gating-augmented softmax attention variants. Specifically, we perform a comprehensive comparison over 30 variants of 15B Mixture-of-Experts (MoE) models and 1.7B dense models trained on a 3.5 trillion token dataset. Our central finding is that a simple modification—applying a head-specific sigmoid gate after the Scaled Dot-Product Attention (SDPA)—consistently improves performance. This modification also enhances training stability, tolerates larger learning rates, and improves scaling properties. By comparing various gating positions and computational variants, we attribute this effectiveness to two key factors: (1) introducing non-linearity upon the low-rank mapping in the softmax attention, and (2) applying query-dependent sparse gating scores to modulate the SDPA output. Notably, we find this sparse gating mechanism mitigates `massive activation`, `attention sink` and enhances long-context extrapolation performance. We also release related codes (https://github.com/qiuzh20/gated_attention}) and models (https://huggingface.co/QwQZh/gated_attention) to facilitate future research. Furthermore, the most effective SDPA output gating is used in the Qwen3-Next models (https://huggingface.co/collections/Qwen/qwen3-next).
Zihan Qiu, Bo Zheng 0007, Kaiyue Wen, Rui Men, Suozhi Huang, Dayiheng Liu, Jingren Zhou 0001, Junyang Lin
NeurIPS7
2025 Semantic communication based on bi-level routing attention in IoT environment
Fan Xiumei, Kok-Lim Alvin Yau, Zhixin Xie, Rui Men
J. Supercomput.5
2024 Mobility-aware parallel offloading and resource allocation scheme for vehicular edge computing
Rui Men, Xiumei Fan, Kok-Lim Alvin Yau, Axida Shan
Ad Hoc Networks1
2022 OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
abstract
In this work, we pursue a unified paradigm for multimodal pretraining to break the shackles of complex task/modality-specific customization. We propose OFA, a Task-Agnostic and Modality-Agnostic framework that supports Task Comprehensiveness. OFA unifies a diverse set of cross-modal and unimodal tasks, including image generation, visual grounding, image captioning, image classification, language modeling, etc., in a simple sequence-to-sequence learning framework. OFA follows the instruction-based learning in both pretraining and finetuning stages, requiring no extra task-specific layers for downstream tasks. In comparison with the recent state-of-the-art vision & language models that rely on extremely large cross-modal datasets, OFA is pretrained on only 20M publicly available image-text pairs. Despite its simplicity and relatively small-scale training data, OFA achieves new SOTAs in a series of cross-modal tasks while attaining highly competitive performances on uni-modal tasks. Our further analysis indicates that OFA can also effectively transfer to unseen tasks and unseen domains. Our code and models are publicly available at https://github.com/OFA-Sys/OFA.
Peng Wang 0028, An Yang, Rui Men, Junyang Lin, Shuai Bai, Chang Zhou 0005, Jingren Zhou 0001, Hongxia Yang
ICML3
2021 Learning Relation Alignment for Calibrated Cross-modal Retrieval
abstract
Shuhuai Ren, Junyang Lin, Guangxiang Zhao, Rui Men, An Yang, Jingren Zhou, Xu Sun, Hongxia Yang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Shuhuai Ren, Junyang Lin, Guangxiang Zhao, Rui Men, An Yang, Jingren Zhou 0001, Xu Sun 0001, Hongxia Yang
ACL/IJCNLP (1)4
2021 M6: Multi-Modality-to-Multi-Modality Multitask Mega-transformer for Unified Pretraining
abstract
Multimodal pretraining has demonstrated success in the downstream tasks of cross-modal representation learning. However, it is limited to the English data, and there is still a lack of large-scale dataset for multimodal pretraining in Chinese. In this work, we propose the largest dataset for pretraining in Chinese, which consists of over 1.9TB images and 292GB texts. The dataset has large coverage over domains, including encyclopedia, question answering, forum discussion, etc. Besides, we propose a method called M6, referring to Multi-Modality-to-Multi-Modality Multitask Mega-transformer, for unified pretraining on the data of single modality and multiple modalities. The model is pretrained with our proposed tasks, including text-to-text transfer, image-to-text transfer, as well as multi-modality-to-text transfer. The tasks endow the model with strong capability of understanding and generation. We scale the model to 10 billion parameters, and build the largest pretrained model in Chinese. Experimental results show that our proposed M6 outperforms the baseline in a number of downstream tasks concerning both single modality and multiple modalities, and the 10B-parameter pretrained model demonstrates strong potential in the setting of zero-shot learning.
Junyang Lin, Rui Men, An Yang, Chang Zhou 0005, Yichang Zhang, Peng Wang 0028, Jingren Zhou 0001, Jie Tang 0001, Hongxia Yang
KDD2
2021 UFC-BERT: Unifying Multi-Modal Controls for Conditional Image Synthesis
abstract
Conditional image synthesis aims to create an image according to some multi-modal guidance in the forms of textual descriptions, reference images, and image blocks to preserve, as well as their combinations. In this paper, instead of investigating these control signals separately, we propose a new two-stage architecture, UFC-BERT, to unify any number of multi-modal controls. In UFC-BERT, both the diverse control signals and the synthesized image are uniformly represented as a sequence of discrete tokens to be processed by Transformer. Different from existing two-stage autoregressive approaches such as DALL-E and VQGAN, UFC-BERT adopts non-autoregressive generation (NAR) at the second stage to enhance the holistic consistency of the synthesized image, to support preserving specified image blocks, and to improve the synthesis speed. Further, we design a progressive algorithm that iteratively improves the non-autoregressively generated image, with the help of two estimators developed for evaluating the compliance with the controls and evaluating the fidelity of the synthesized image, respectively. Extensive experiments on a newly collected large-scale clothing dataset M2C-Fashion and a facial dataset Multi-Modal CelebA-HQ verify that UFC-BERT can synthesize high-fidelity images that comply with flexible multi-modal controls.
Chang Zhou 0005, Rui Men, Ming Ding 0004, Jie Tang 0001, Jingren Zhou 0001, Hongxia Yang
NeurIPS4
2019 Deep-AutoCoder: Learning to Complete Code Precisely with Induced Code Tokens
abstract
Code completion is an essential part of modern IDEs. It assists the developers to speed up the process of coding and reducing typos. In this paper, we exploit the deep learning technique called LSTM to learn language models over large code corpus and make predictions of code elements. Unlike natural language, the innumerable identifiers lead to the vocabulary explosion and more difficult to predict. Therefore, we propose a new approach, the Induced Token based LSTM, to deal with the massive identifiers, thus decrease the vocabulary size. In order to induce the code tokens, we present two approaches, one is a constraint character-level LSTM and the other one is encoding identifiers with various preceding context before feeding them into a token-level LSTM. Based on the two approaches, a tool named Deep-AutoCoder is developed and evaluated in two classic completion scenarios, that is, method invocation completion and random completion. The experiment results indicate that Deep-AutoCoder outperforms the state-of-the-arts on method invocation completion and random code completion. Additionally, the empirical results of Deep-AutoCoder indicate that reducing the size of vocabulary can effectively improve the precision of code completion.
Xing Hu 0008, Rui Men, Ge Li 0001, Zhi Jin 0001
COMPSAC (1)2