VLDB 2026 Research / reviewers in the wild / expert
Eduin E. Hernandez
dblp:277/9559
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2024
0009-0008-2812-5311ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Communication-Efficient Federated DNN Training: Convert, Compress, CorrectabstractIn the federated training of a deep neural network (DNN), model updates are transmitted from the remote users to the parameter server (PS). In many scenarios of practical relevance, one is interested in reducing the communication overhead to enhance training efficiency. To address this challenge, we introduce$\textsf {CO}_{3}$.$\textsf {CO}_{3}$takes its name from three processing applied which reduce the communication load when transmitting the local DNN gradients from the remote users to the PS. Namely, 1) gradient quantization through floating-point conversion; 2) lossless compression of the quantized gradient; and 3) correction of quantization error. We carefully design each of the steps above to ensure good training performance under a constraint on the communication rate. In particular, in steps 1) and 2), we adopt the assumption that DNN gradients are distributed according to a generalized normal distribution, which is validated numerically in this article. For step 3), we utilize an error feedback with a memory decay mechanism to correct the quantization error introduced in step 1). We argue that the memory decay coefficient –similar to the learning rate—can be optimally tuned to improve convergence. A rigorous convergence analysis of the proposed$\textsf {CO}_{3}$with stochastic gradient descent (SGD) is provided. Moreover, with extensive simulations, we show that$\textsf {CO}_{3}$offers improved performance as compared with existing gradient compression schemes proposed in the literature which employ sketching and nonuniform quantization of the local gradients. Zhong-Jing Chen, Eduin E. Hernandez, Yu-Chih Huang, Stefano Rini |
IEEE Internet Things J. | 2 |
| 2022 | DNN gradient lossless compression: Can GenNorm be the answer?abstractIn this paper, the problem of optimal gradient lossless compression in Deep Neural Network (DNN) training is considered. Gradient compression is relevant in many distributed DNN training scenarios, including the recently popular federated learning (FL) scenario in which each remote users are connected to the parameter server (PS) through a noiseless but rate limited channel. In distributed DNN training, if the underlying gradient distribution is available, classical lossless compression approaches can be used to reduce the number of bits required for communicating the gradient entries. Mean field analysis has suggested that gradient updates can be considered as independent random variables, while Laplace approximation can be used to argue that gradient has a distribution approximating the normal (Norm) distribution in some regimes. In this paper we argue that, for some networks of practical interest, the gradient entries can be well modelled as having a generalized normal (GenNorm) distribution. We provide numerical evaluations to validate that the hypothesis GenNorm modelling provides a more accurate prediction of the DNN gradient tail distribution. Additionally, this modeling choice provides concrete improvement in terms of lossless compression of the gradients when applying classical fix-to-variable lossless coding algorithms, such as Huffman coding, to the quantized gradient updates. This latter results indeed provides an effective compression strategy with low memory and computational complexity that has great practical relevance in distributed DNN training scenarios. Zhong-Jing Chen, Eduin E. Hernandez, Yu-Chih Huang, Stefano Rini |
ICC | 2 |
| 2022 | Straggler Mitigation Through Unequal Error Protection for Distributed Approximate Matrix MultiplicationabstractLarge-scale machine learning and data mining methods routinely distribute computations across multiple agents to parallelize processing. The time required for the computations at the agents is affected by the availability of local resources and/or poor channel conditions, thus giving rise to the “straggler problem.” In this paper, we address this problem for distributed approximate matrix multiplication. In particular, we employ Unequal Error Protection (UEP) codes to obtain an approximation of the matrix product to provide higher protection for the blocks with a higher effect on the multiplication outcome. We characterize the performance of the proposed approach from a theoretical perspective by bounding the expected reconstruction error for matrices with uncorrelated entries. We also apply the proposed coding strategy to the computation of the back-propagation step in the training of a Deep Neural Network (DNN) for an image classification task in the evaluation of the gradients. Our numerical experiments show that it is indeed possible to obtain significant improvements in the overall time required to achieve DNN training convergence by producing approximation of matrix products using UEP codes in the presence of stragglers. Busra Tegin, Eduin E. Hernandez, Stefano Rini, Tolga M. Duman |
IEEE J. Sel. Areas Commun. | 2 |
| 2021 | Straggler Mitigation through Unequal Error Protection for Distributed Matrix Multiplication
Busra Tegin, Eduin E. Hernandez, Stefano Rini, Tolga M. Duman |
ICC | 2 |