VLDB 2026 Research / reviewers in the wild / expert
Rui Men
dblp:170/0093
· DBLP profile ↗
12ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-4429-3461ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Deep learning architectures and training · 35% Language models and text generation · 29% Representation and self-supervised learning · 14% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 100% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% | |
| Computer networks
1 paper |
Datacenter networks · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining |
1.1 | 2 | 2022 | OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework · ICML 2022 M6: Multi-Modality-to-Multi-Modality Multitask Mega-transformer for Unified Pretraining · KDD 2021 |
Natural language and speech › Language models and text generation › large language model reasoning
multi-turn reasoning |
1.0 | 1 | 2026 | MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation · ACL (1) 2026 |
Natural language and speech › Language models and text generation › evaluation of language models
reasoning evaluation |
1.0 | 1 | 2026 | MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation · ACL (1) 2026 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.9 | 1 | 2025 | Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › language model interpretability
attention sink |
0.9 | 1 | 2025 | Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › attention mechanism › attention module
gated attention |
0.9 | 1 | 2025 | Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › distributed training › large model training
mixture-of-experts training |
0.9 | 1 | 2025 | Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models · ACL (1) 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.9 | 1 | 2025 | Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free · NeurIPS 2025 |
Distributed systems › distributed machine learning
distributed training |
0.9 | 1 | 2025 | Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization · HPCA 2025 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence-to-sequence learning |
0.6 | 1 | 2022 | OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework · ICML 2022 |
Computer vision › Vision and language
vision-language model |
0.6 | 1 | 2022 | OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework · ICML 2022 |
Natural language and speech › Language models and text generation › large language model
chinese language model |
0.5 | 1 | 2021 | M6: Multi-Modality-to-Multi-Modality Multitask Mega-transformer for Unified Pretraining · KDD 2021 |
Machine learning › Representation and self-supervised learning › multimodal representation learning
cross-modal representation learning |
0.5 | 1 | 2021 | M6: Multi-Modality-to-Multi-Modality Multitask Mega-transformer for Unified Pretraining · KDD 2021 |
Natural language and speech › Language models and text generation
multilingual language models |
0.5 | 1 | 2021 | M6: Multi-Modality-to-Multi-Modality Multitask Mega-transformer for Unified Pretraining · KDD 2021 |
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
non-autoregressive generation |
0.5 | 1 | 2021 | UFC-BERT: Unifying Multi-Modal Controls for Conditional Image Synthesis · NeurIPS 2021 |
Information retrieval
cross-modal retrieval |
0.5 | 1 | 2021 | Learning Relation Alignment for Calibrated Cross-modal Retrieval · ACL/IJCNLP (1) 2021 |
Visual content generation and editing › image generation
conditional image synthesis |
0.5 | 1 | 2021 | UFC-BERT: Unifying Multi-Modal Controls for Conditional Image Synthesis · NeurIPS 2021 |
Visual content generation and editing
multimodal control |
0.5 | 1 | 2021 | UFC-BERT: Unifying Multi-Modal Controls for Conditional Image Synthesis · NeurIPS 2021 |
Natural language and speech › Language models and text generation
large language model training |
0.3 | 1 | 2025 | Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models · ACL (1) 2025 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.3 | 1 | 2025 | Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free · NeurIPS 2025 |
Distributed systems
fault tolerance |
0.3 | 1 | 2025 | Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization · HPCA 2025 |
Methods — techniques the papers use, named apart from their topics
traffic planning · 1.7collective communication analysis · 1.7transformer · 1.5non-autoregressive generation · 1.0benchmarking · 1.0sigmoid gating · 0.9scaled dot-product attention · 0.9load-balancing loss · 0.9sequence-to-sequence · 0.6instruction-based learning · 0.6multimodal pretraining · 0.5cross-modal alignment · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning EvaluationabstractXiaoyuan Li, Keqin Bao, Yubo Ma, Moxin Li, Wenjie Wang, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaoyuan Li 0001, Keqin Bao, Yubo Ma, Moxin Li, Wenjie Wang 0007, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu |
ACL (1) | 6 |
| 2026 | A Lightweight Continuous Identity Authentication-Based Security Offloading Scheme in Vehicular Edge ComputingabstractVehicular Edge Computing (VEC) is pivotal for latency-sensitive vehicular applications but confronts three critical challenges: traditional one-time identity authentication cannot adapt to high mobility, leaving privacy and data security vulnerabilities, offloading system security levels lack quantifiability, and security-performance optimization objectives are inherently conflicting. To address these limitations, we propose a lightweight continuous identity authentication-based secure offloading scheme for VEC. First, a three-entity collaborative architecture is designed, which leverages chameleon hash function (CHF) to reduce vehicle-side signature overhead, and Bloom filter (BF) to enable real-time verification during vehicle-roadside unit (RSU) handovers. Second, a dual-dimensional security framework that quantifies authentication and data signature levels is established, enabling on-demand security adjustment for diverse tasks. Third, to balance task latency minimization and security maximization, Unlike prior works that optimize security and offloading separately, this framework unifies both identity authentication and data transmission security into a holistic latency-oriented offloading design, filling critical research gaps in high-mobility VEC scenarios. To tackle this intricate problem, we decompose it into four sub-problems. These sub-problems are solved using Lagrangian duality, the Newton-Raphson method, and the branch-and-bound algorithm to obtain stable and highquality feasible solutions efficiently. Extensive simulations against four baseline schemes demonstrate that the proposed approach achieves fast convergence and priority performance. Rui Men, Axida Shan, Celimuge Wu, Jie Li 0002 |
IEEE Internet Things J. | 1 |
| 2025 | Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert ModelsabstractZihan Qiu, Zeyu Huang, Bo Zheng, Kaiyue Wen, Zekun Wang, Rui Men, Ivan Titov, Dayiheng Liu, Jingren Zhou, Junyang Lin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zihan Qiu, Bo Zheng 0007, Kaiyue Wen, Rui Men, Ivan Titov 0001, Dayiheng Liu, Jingren Zhou 0001, Junyang Lin |
ACL (1) | 6 |
| 2025 | Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication OptimizationabstractThe emergence of Large Language Models (LLMs) has necessitated the adoption of distributed training techniques, involving the deployment of thousands of GPUs to train a single model. Unfortunately, the efficiency of large-scale distributed training systems is often suboptimal due to the increased likelihood of hardware errors in high-end GPU products and the heightened risk of network traffic collisions. Specifically, GPUs involved in the same job require periodic synchronization to exchange necessary data, such as gradients, parameters, or activations. As a result, any local hardware failure can disrupt training tasks, and the inability to swiftly identify faulty components leads to a significant waste of GPU resources. Moreover, prolonged communication due to traffic collisions can substantially increase GPU waiting times. To address these challenges, we propose a communication-driven solution, namely the C 4. The key insights of C 4 are twofold. First, the load in distributed training exhibits homogeneous characteristics and is divided into iterations through periodic synchronization, therefore hardware anomalies would incur certain syndrome in collective communication. By leveraging this feature, $\mathbf{C} 4$ can rapidly identify the faulty components, swiftly isolate the anomaly, and restart the task, thereby avoiding resource wastage caused by delays in anomaly detection. Second, the predictable communication model of collective communication, involving a limited number of long-lived flows, allows C 4 to efficiently execute traffic planning, substantially reducing bandwidth competition among these flows. The $\mathbf{C 4}$ has been extensively deployed across real-world production systems in a hyperscale cloud provider, yielding a significant improvement in system efficiency, from 30% to $\mathbf{4 5 \%}$. This enhancement is attributed to a $\mathbf{3 0 \%}$ reduction in error-induced overhead and a 15% reduction in communication costs. Jianbo Dong, Yikai Zhu, Hairong Jiao, Ennan Zhai, Wencong Xiao, Man Yuan, Siran Yang, Jiamang Wang, Rui Men, Dennis Cai, Binzhang Fu |
HPCA | 20 |
| 2025 | Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeabstractGating mechanisms have been widely utilized, from early models like LSTMs and Highway Networks to recent state space models, linear attention, and also softmax attention.
Yet, existing literature rarely examines the specific effects of gating.
In this work, we conduct comprehensive experiments to systematically investigate gating-augmented softmax attention variants.
Specifically, we perform a comprehensive comparison over 30 variants of 15B Mixture-of-Experts (MoE) models and 1.7B dense models trained on a 3.5 trillion token dataset.
Our central finding is that a simple modification—applying a head-specific sigmoid gate after the Scaled Dot-Product Attention (SDPA)—consistently improves performance.
This modification also enhances training stability, tolerates larger learning rates, and improves scaling properties.
By comparing various gating positions and computational variants, we attribute this effectiveness to two key factors: (1) introducing non-linearity upon the low-rank mapping in the softmax attention, and (2) applying query-dependent sparse gating scores to modulate the SDPA output.
Notably, we find this sparse gating mechanism mitigates `massive activation`, `attention sink` and enhances long-context extrapolation performance.
We also release related codes (https://github.com/qiuzh20/gated_attention}) and models (https://huggingface.co/QwQZh/gated_attention) to facilitate future research.
Furthermore, the most effective SDPA output gating is used in the Qwen3-Next models (https://huggingface.co/collections/Qwen/qwen3-next). Zihan Qiu, Bo Zheng 0007, Kaiyue Wen, Rui Men, Suozhi Huang, Dayiheng Liu, Jingren Zhou 0001, Junyang Lin |
NeurIPS | 7 |
| 2025 | Semantic communication based on bi-level routing attention in IoT environment
Fan Xiumei, Kok-Lim Alvin Yau, Zhixin Xie, Rui Men |
J. Supercomput. | 5 |
| 2024 | Mobility-aware parallel offloading and resource allocation scheme for vehicular edge computing
Rui Men, Xiumei Fan, Kok-Lim Alvin Yau, Axida Shan |
Ad Hoc Networks | 1 |
| 2022 | OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkabstractIn this work, we pursue a unified paradigm for multimodal pretraining to break the shackles of complex task/modality-specific customization. We propose OFA, a Task-Agnostic and Modality-Agnostic framework that supports Task Comprehensiveness. OFA unifies a diverse set of cross-modal and unimodal tasks, including image generation, visual grounding, image captioning, image classification, language modeling, etc., in a simple sequence-to-sequence learning framework. OFA follows the instruction-based learning in both pretraining and finetuning stages, requiring no extra task-specific layers for downstream tasks. In comparison with the recent state-of-the-art vision & language models that rely on extremely large cross-modal datasets, OFA is pretrained on only 20M publicly available image-text pairs. Despite its simplicity and relatively small-scale training data, OFA achieves new SOTAs in a series of cross-modal tasks while attaining highly competitive performances on uni-modal tasks. Our further analysis indicates that OFA can also effectively transfer to unseen tasks and unseen domains. Our code and models are publicly available at https://github.com/OFA-Sys/OFA. Peng Wang 0028, An Yang, Rui Men, Junyang Lin, Shuai Bai, Chang Zhou 0005, Jingren Zhou 0001, Hongxia Yang |
ICML | 3 |
| 2021 | Learning Relation Alignment for Calibrated Cross-modal RetrievalabstractShuhuai Ren, Junyang Lin, Guangxiang Zhao, Rui Men, An Yang, Jingren Zhou, Xu Sun, Hongxia Yang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Shuhuai Ren, Junyang Lin, Guangxiang Zhao, Rui Men, An Yang, Jingren Zhou 0001, Xu Sun 0001, Hongxia Yang |
ACL/IJCNLP (1) | 4 |
| 2021 | M6: Multi-Modality-to-Multi-Modality Multitask Mega-transformer for Unified PretrainingabstractMultimodal pretraining has demonstrated success in the downstream tasks of cross-modal representation learning. However, it is limited to the English data, and there is still a lack of large-scale dataset for multimodal pretraining in Chinese. In this work, we propose the largest dataset for pretraining in Chinese, which consists of over 1.9TB images and 292GB texts. The dataset has large coverage over domains, including encyclopedia, question answering, forum discussion, etc. Besides, we propose a method called M6, referring to Multi-Modality-to-Multi-Modality Multitask Mega-transformer, for unified pretraining on the data of single modality and multiple modalities. The model is pretrained with our proposed tasks, including text-to-text transfer, image-to-text transfer, as well as multi-modality-to-text transfer. The tasks endow the model with strong capability of understanding and generation. We scale the model to 10 billion parameters, and build the largest pretrained model in Chinese. Experimental results show that our proposed M6 outperforms the baseline in a number of downstream tasks concerning both single modality and multiple modalities, and the 10B-parameter pretrained model demonstrates strong potential in the setting of zero-shot learning. Junyang Lin, Rui Men, An Yang, Chang Zhou 0005, Yichang Zhang, Peng Wang 0028, Jingren Zhou 0001, Jie Tang 0001, Hongxia Yang |
KDD | 2 |
| 2021 | UFC-BERT: Unifying Multi-Modal Controls for Conditional Image SynthesisabstractConditional image synthesis aims to create an image according to some multi-modal guidance in the forms of textual descriptions, reference images, and image blocks to preserve, as well as their combinations. In this paper, instead of investigating these control signals separately, we propose a new two-stage architecture, UFC-BERT, to unify any number of multi-modal controls. In UFC-BERT, both the diverse control signals and the synthesized image are uniformly represented as a sequence of discrete tokens to be processed by Transformer. Different from existing two-stage autoregressive approaches such as DALL-E and VQGAN, UFC-BERT adopts non-autoregressive generation (NAR) at the second stage to enhance the holistic consistency of the synthesized image, to support preserving specified image blocks, and to improve the synthesis speed. Further, we design a progressive algorithm that iteratively improves the non-autoregressively generated image, with the help of two estimators developed for evaluating the compliance with the controls and evaluating the fidelity of the synthesized image, respectively. Extensive experiments on a newly collected large-scale clothing dataset M2C-Fashion and a facial dataset Multi-Modal CelebA-HQ verify that UFC-BERT can synthesize high-fidelity images that comply with flexible multi-modal controls. Chang Zhou 0005, Rui Men, Ming Ding 0004, Jie Tang 0001, Jingren Zhou 0001, Hongxia Yang |
NeurIPS | 4 |
| 2019 | Deep-AutoCoder: Learning to Complete Code Precisely with Induced Code TokensabstractCode completion is an essential part of modern IDEs. It assists the developers to speed up the process of coding and reducing typos. In this paper, we exploit the deep learning technique called LSTM to learn language models over large code corpus and make predictions of code elements. Unlike natural language, the innumerable identifiers lead to the vocabulary explosion and more difficult to predict. Therefore, we propose a new approach, the Induced Token based LSTM, to deal with the massive identifiers, thus decrease the vocabulary size. In order to induce the code tokens, we present two approaches, one is a constraint character-level LSTM and the other one is encoding identifiers with various preceding context before feeding them into a token-level LSTM. Based on the two approaches, a tool named Deep-AutoCoder is developed and evaluated in two classic completion scenarios, that is, method invocation completion and random completion. The experiment results indicate that Deep-AutoCoder outperforms the state-of-the-arts on method invocation completion and random code completion. Additionally, the empirical results of Deep-AutoCoder indicate that reducing the size of vocabulary can effectively improve the precision of code completion. Xing Hu 0008, Rui Men, Ge Li 0001, Zhi Jin 0001 |
COMPSAC (1) | 2 |