Daning Cheng

dblp:163/1978 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-9217-6980ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 FP=XINT: Representing Neural Networks via Low-Bit Series Basis Functions
abstract
Deep neural networks are often over-parameterized, resulting in prohibitive storage and computational costs. A fundamental question is whether a complex network can be re-expressed in terms of a compact set of basis functions without sacrificing accuracy. Motivated by this perspective, we aim to approximate a dense model by decomposing it into a small number of lightweight components that capture the essential functional structure of the network. To this end, we propose a series expansion framework that rewrites a neural network as a linear combination of low-bit basis models. Within the post-training quantization setting, the full-precision model is expanded hierarchically at the tensor, layer, and model levels into a structured set of basis functions. We theoretically prove that this expansion converges exponentially to the original model. Furthermore, we design AbelianAdd and AbelianMul operations between isomorphic basis models, endowing the expansion with an Abelian group structure that naturally supports commutative and parallel computation. Experimental results across diverse architectures show that our series expansion method leverages a set of ultra-low-bit basis functions, not only preserving full-precision performance without the need for calibration data or fine-tuning, but also featuring a parallel-friendly design that enables efficient and scalable deployment.
Daning Cheng, Yunquan Zhang, Jiake Tian, Fangming Liu
AAAI2
2026 BCD-Megatron: A Cost-Effective Training System for Large Language Models
Yunquan Zhang, Guoyong Jiang, Daning Cheng
ICDCS9
2026 Data Scaling Laws for Block-Sparse Training
Zhenfeng Zhang, Yunquan Zhang, Daning Cheng
ICPR (4)4
2023 Asynch-SGBDT: Train Stochastic Gradient Boosting Decision Trees in an Asynchronous Parallel Manner
abstract
Gradient Boosting Decision Tree (GBDT) is a costly machine learning model. Current parallel GBDT algorithms generally follow a synchronous parallel design: Fork-join parallel manner, like MapReduce. Fork-join parallel manner needs considerable time. Thus, we propose whether synchronization is necessary for GBDT training and is asynchronous training manner efficient. In this paper, we solve the above problem by offering an asynchronous algorithm. We try to build a stochastic optimization problem by sampling, which shares the same output with original GBDT training problem and use asynchronous parallel SGD manner to train Gradient step GBDT. We name our algorithm as asynch-SGBDT. Our theoretical and experimental results indicate that compared with the serial GBDT training process, when the datasets’ high sample diversity is high and using Gradient step training GBDT, asynch-SGBDT does not slow down convergence speed on the epoch, and the sample diversity of current high-dimensional sparse datasets is usually high. We conduct experiments on a 32-node cluster using four different datasets. The results show that with LightGBM using a single worker as the baseline, LightGBM (the state-of-the-art synchronous parallel algorithm implement) on 32 workers achieves 5x-7x speedup, while our asynch-SGBDT on 32 workers increases the speedup to 11x-15x.
Daning Cheng, Shigang Li 0002, Yunquan Zhang
IPDPS1
2021 Why Dataset Properties Bound the Scalability of Parallel Machine Learning Training Algorithms
abstract
As the training dataset size and the model size of machine learning increase rapidly, more computing resources are consumed to speedup the training process. However, the scalability and performance reproducibility of parallel machine learning training, which mainly uses stochastic optimization algorithms, are limited. In this paper, we demonstrate that the sample difference in the dataset plays a prominent role in the scalability of parallel machine learning algorithms. We propose to use statistical properties of dataset to measure sample differences. These properties include the variance of sample features, sample sparsity, sample diversity, and similarity in sampling sequences. We choose four types of parallel training algorithms as our research objects: (1) the asynchronous parallel SGD algorithm (Hogwild! algorithm), (2) the parallel model average SGD algorithm (minibatch SGD algorithm), (3) the decentralization optimization algorithm, and (4) the dual coordinate optimization (DADM algorithm). Our results show that the statistical properties of training datasets determine the scalability upper bound of these parallel training algorithms.
Daning Cheng, Shigang Li 0002, Hanping Zhang, Fen Xia, Yunquan Zhang
IEEE Trans. Parallel Distributed Syst.1
2020 WP-SGD: Weighted parallel SGD for distributed unbalanced-workload training system
Daning Cheng, Shigang Li 0002, Yunquan Zhang
J. Parallel Distributed Comput.1
2019 Using Gradient Based Multikernel Gaussian Process and Meta-Acquisition Function to Accelerate SMBO
abstract
Automatic machine learning (automl) is a crucial technology in machine learning. Sequential model-based optimisation algorithms (SMBO) (e.g., SMAC, TPE) are state-of-the-art hyperparameter optimisation methods in automl. However, SMBO does not consider known information, like the best hyperparameters high possibility range and gradients. In this paper, we accelerate the traditional SMBO method and name our method as accSMBO. In accSMBO, we build a gradient-based multikernel Gaussian process with a good generalisation ability and we design meta-acquisition function which encourages that SMBO puts more attention on the best hyperparameters high possibility range. In L2 norm regularised logistic loss function experiments, our method exhibited state-of-the-art performance.
Daning Cheng, Hanping Zhang, Fen Xia, Shigang Li 0002, Yunquan Zhang
ICTAI1