Zhenyu Ming

dblp:279/0003 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 33% Language models and text generation · 33% Planning, search and constraint satisfaction · 33%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model
large language model training and inference
0.912025
BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference · AAAI 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.912025
Improving Monte Carlo Tree Search for Symbolic Regression · NeurIPS 2025
Program synthesis and code generation › inductive program synthesis
symbolic regression
0.912025
Improving Monte Carlo Tree Search for Symbolic Regression · NeurIPS 2025
Distributed systems › distributed machine learning
distributed training
0.312025
BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference · AAAI 2025
Mathematical optimization
combinatorial optimization
0.312025
Improving Monte Carlo Tree Search for Symbolic Regression · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

mutation · 2.6monte carlo tree search · 2.6extreme bandit allocation · 2.6crossover · 2.6projection · 1.7all-to-all communication · 1.7mixture-of-experts · 0.9mixture of experts · 0.9
YearPublicationVenuePosition
2026 An Efficient IID-Based Routing Table Aggregation Algorithm for Large-Scale Load-Sharing Networks
Haoran Pang, Zhenyu Ming
ICC3
2025 BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference
abstract
The Mixture-of-Experts (MoE) structure scales the Transformer-based large language models (LLMs) and improves their performance with only the sub-linear increase in computation resources. Recently, a fine-grained DeepSeekMoE structure is proposed, which can further improve the computing efficiency of MoE without performance degradation. However, the All-to-All communication introduced by MoE has become a bottleneck, especially for the fine-grained structure, which typically involves and activates more experts, hence contributing to heavier communication overhead. In this paper, we propose a novel MoE structure named BigMac, which is also fine-grained but with high communication efficiency. The innovation of BigMac is mainly due to that we abandon the Communicate-Descend-Ascend-Communicate (CDAC) manner used by fine-grained MoE, which leads to the All-to-All communication always taking place at the highest dimension. Instead, BigMac designs an efficient Descend-Communicate-Communicate-Ascend (DCCA) manner. Specifically, we add a descending and ascending projection at the entrance and exit of the expert, respectively, which enables the communication to perform at a very low dimension. Furthermore, to adapt to DCCA, we re-design the structure of small experts, ensuring that the expert in BigMac has enough complexity to address tokens. Experimental results show that BigMac achieves comparable or even better model quality than fine-grained MoEs with the same number of experts and a similar number of total parameters. Equally importantly, BigMac reduces the end-to-end latency by up to 3.09 x for training and increases the throughput by up to 3.11 x for inference on state-of-the-art AI computing frameworks including Megatron, Tutel, and DeepSpeed-Inference.
Zewen Jin, Jiaan Zhu, Hongrui Zhan, Youhui Bai, Zhenyu Ming
AAAI7
2025 Accelerated Dropout: A Bitmask Approach to Speed Up Model Training
abstract
Dropout [1], a standard regularization technique in Large Language Models, incurs extra computational overhead, particularly due to its repeated application during the training process. To address this issue, we propose Accelerated Dropout, an algorithm that revolutionizes the traditional dropout method by employing a bitmask, rather than a float mask, to significantly expedite the training process. In the theoretical proof, we demonstrate that accelerated dropout is exponentially convergent to the probability of retaining neurons. Our extensive experimental analysis, conducted on 13 benchmark datasets and 9 deep learning models, confirms that accelerated dropout outperforms traditional dropout in terms of training efficiency and generalization performance. The experimental results indicate that, compared to torch dropout, our accelerated dropout achieves 8x speedup on a single dropout operator. With the use of accelerated dropout, the average training speed of each model increased by 1.0624x. The training speed improvements ranged from a 1.046x increase for ProphetNet on the CNN/DailyMail dataset to a 1.077x increase for ViT-B on the ImageNet dataset.
Jincheng Xie, Sicheng Xu, Zhenyu Ming
IJCNN5
2025 Improving Monte Carlo Tree Search for Symbolic Regression
abstract
Symbolic regression aims to discover concise, interpretable mathematical expressions that satisfy desired objectives, such as fitting data, posing a highly combinatorial optimization problem. While genetic programming has been the dominant approach, recent efforts have explored reinforcement learning methods for improving search efficiency. Monte Carlo Tree Search (MCTS), with its ability to balance exploration and exploitation through guided search, has emerged as a promising technique for symbolic expression discovery. However, its traditional bandit strategies and sequential symbol construction often limit performance. In this work, we propose an improved MCTS framework for symbolic regression that addresses these limitations through two key innovations: (1) an extreme bandit allocation strategy tailored for identifying globally optimal expressions, with finite-time performance guarantees under polynomial reward decay assumptions; and (2) evolution-inspired state-jumping actions such as mutation and crossover, which enable non-local transitions to promising regions of the search space. These state-jumping actions also reshape the reward landscape during the search process, improving both robustness and efficiency. We conduct a thorough numerical study to the impact of these improvements and benchmark our approach against existing symbolic regression methods on a variety of datasets, including both ground-truth and black-box datasets. Our approach achieves competitive performance with state-of-the-art libraries in terms of recovery rate, attains favorable positions on the Pareto frontier of accuracy versus model complexity.
Zhengyao Huang, Daniel Zhengyu Huang, Tiannan Xiao, Dina Ma, Zhenyu Ming, Yuanhui Wen
NeurIPS5
2023 LBFF: Load-Balancing First Fit Algorithm for Tenant Placement Problem
abstract
To meet the prevalence of cloud services, it is of great significance for operators to design a cost-effective tenant placement strategy on physical machines while the services are still guaranteed. Moreover, workload balancing between physical machines is critical for high performance. We characterize the problem of tenant placement as a mixed-integer programming problem that efficiently minimizes the number of active physical machines and balances the workloads among them. To address the NP-hardness of the optimization problem, we first investigate the global lower bound and other inherent properties of the proposed model, and then design two efficient heuristic algorithms. Through numerous simulations at different scales and settings, we demonstrate the superiority of our proposed algorithms over state-of-the-art works in terms of the objective function values and computational time.
Zhenyu Ming, Xi Peng 0006, Liping Zhang 0008
ICC2
2022 An accurate and practical algorithm for internet traffic recovery problem
Zhenyu Ming, Liping Zhang 0008, Hao Wu 0060, Yanwei Xu 0004, Mayank Bakshi, Bo Bai 0001, Gong Zhang 0001
Neurocomputing1