Shayan Mohajer Hamidi

dblp:307/5037 · DBLP profile ↗
← Back
16ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0001-8321-7130ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Entropy-Adaptive Federated Learning with Efficient Bit Allocation over Wireless Channels
Shayan Mohajer Hamidi, Ben Liang 0001
INFOCOM1
2025 Rate-Constrained Quantization for Communication-Efficient Federated Learning
abstract
Quantization is a common approach to mitigate the communication cost of federated learning (FL). In practice, the quantized local parameters are further encoded via an entropy coding technique, such as Huffman coding, for efficient data compression. In this case, the exact communication overhead is determined by the bit rate of the encoded gradients. Recognizing this fact, this work deviates from the existing approaches in the literature and develops a novel quantized FL framework, called rate-constrained federated learning (RC-FED), in which we deploy the conventional entropy-constrained scalar quantization technique to quantize the gradients subject to both fidelity and data rate constraints. Particularly, we formulate this scheme, as a joint optimization in which the quantization distortion is minimized while the rate of encoded gradients is kept below a target threshold. This enables for a tunable trade-off between quantization distortion and communication cost. We analyze the convergence behavior of RC-FED, and show its superior performance against baseline quantized FL schemes on several datasets.
Shayan Mohajer Hamidi, Ali Bereyhi
ICASSP1
2025 Second-Order Wireless Federated Leaning via Nonparametric Hessian Estimation
abstract
Quasi-Newton algorithms estimate the second-order information of loss landscape from its first-order information. They hence propose a promising solution for communication-efficient federated learning (FL), as they reduce the required number of training rounds while avoiding the necessity of exchanging local Hessians over the network. Despite that, the quasi-Newton approaches prove less effective in wireless FL, as the noisy aggregation in this case causes bias in the estimate of the Newton direction. This paper proposes a novel second-order wireless FL algorithm. The pivotal innovation lies in the server’s ability to estimate the global Hessian based on a window of noisy aggregations. The server acquires this ability by computing a stochastic estimator of the global Hessian under a Gaussian prior belief. Numerical experiments show that the proposed scheme can compute a less-biased estimator of the Newton direction, and hence a superior learning performance, as compared to the baseline.
Shayan Mohajer Hamidi, Ali Bereyhi
ICASSP1
2025 Score-Based Manifold Projection for Diffusion-Based Inverse Problems
abstract
Inverse problems such as inpainting, deblurring, and super-resolution benefit significantly from generative diffusion models, which serve as powerful learned priors. However, incorporating measurement-consistency gradients naively can push the intermediate solutions off the high-likelihood manifold encoded by the diffusion model, leading to artifacts or suboptimal reconstructions. In this paper, we propose a scorebased manifold projection framework that leverages the internal score function of the diffusion model itself to preserve manifold fidelity. Specifically, we exploit the fact that the diffusion score$\boldsymbol{s}_{\theta}(\boldsymbol{x}, t) \approx \nabla_{\boldsymbol{x}} \log p_{t}(\boldsymbol{x})$is ideally orthogonal to manifolds of constant log-likelihood at time$t$. By removing the component of the measurement gradient parallel to$\boldsymbol{s}_{\theta}$, our method constrains each update to remain (to first order) tangent to the learned data manifold. Theoretically, we prove that our approach preserves proximity to the manifold more effectively than an un-projected update, and empirically, we demonstrate improved robustness and quality on various inverse problems, including deblurring, inpainting, and super-resolution. Our results show that scorebased manifold projection not only reduces artifacts but also maintains the fidelity of the reconstructions, offering a simple yet effective enhancement to measurement-guided diffusion solvers.
Shayan Mohajer Hamidi, En-Hui Yang
ISIT1
2025 Conditional Mutual Information Based Diffusion Posterior Sampling for Solving Inverse Problems
abstract
Inverse problems are prevalent across various disciplines in science and engineering. In the field of computer vision, tasks such as inpainting, deblurring, and super-resolution are commonly formulated as inverse problems. Recently, diffusion models (DMs) have emerged as a promising approach for addressing noisy linear inverse problems, offering effective solutions without requiring additional task-specific training. Specifically, with the prior provided by DMs, one can sample from the posterior by finding the likelihood. Since the likelihood is intractable, it is often approximated in the literature. However, this approximation compromises the quality of the generated images. To overcome this limitation and improve the effectiveness of DMs in solving inverse problems, we propose an information-theoretic approach. Specifically, we maximize the conditional mutual information$\mathrm{I}\left(x_{0}; y \mid x_{t}\right)$, where$x_{0}$represents the reconstructed signal,$y$is the measurement, and$x_{t}$is the intermediate signal at stage$t$. This ensures that the intermediate signals$x_{t}$are generated in a way that the final reconstructed signal$x_{0}$retains as much information as possible about the measurement$y$. We demonstrate that this method can be seamlessly integrated with recent approaches and, once incorporated, enhances their performance both qualitatively and quantitatively.
Shayan Mohajer Hamidi, En-Hui Yang
ISIT1
2025 Coded Deep Learning: Framework and Preliminary Results
abstract
Deep learning (DL) often achieves success at the cost of large model sizes and high computational complexity, making training and inference challenging in resource-limited environments. To address this, we introduce coded deep learning (CDL), a framework that integrates information-theoretic coding concepts into DL to compress model weights and activations, reduce computational complexity, and enable efficient model/data parallelism. Specifically, CDL: (i) introduces a probabilistic quantization method for model weights and activations, including a differentiable variant for gradient computation; (ii) executes both forward and backward passes on quantized values, significantly reducing floating-point operations and training complexity; (iii) enforces entropy constraints on weights and activations, ensuring compressibility throughout training and lowering communication costs in distributed settings; and (iv) produces a quantized model by default, reducing post-training inference and storage complexity. Extensive experiments demonstrate that CDL outperforms state-of-the-art DNN compression methods.
En-Hui Yang, Shayan Mohajer Hamidi
ISIT2
2025 Coupled Data and Measurement Space Dynamics for Enhanced Diffusion Posterior Sampling
abstract
Inverse problems, where the goal is to recover an unknown signal from noisy or incomplete measurements, are central to applications in medical imaging, remote sensing, and computational biology. Diffusion models have recently emerged as powerful priors for solving such problems. However, existing methods either rely on projection-based techniques that enforce measurement consistency through heuristic updates, or they approximate the likelihood $p(\boldsymbol{y} \mid \boldsymbol{x})$, often resulting in artifacts and instability under complex or high-noise conditions. To address these limitations, we propose a novel framework called coupled data and measurement space diffusion posterior sampling (C-DPS), which eliminates the need for constraint tuning or likelihood approximation. C-DPS introduces a forward stochastic process in the measurement space $\{\boldsymbol{y}_t\}$, evolving in parallel with the data-space diffusion $\{\boldsymbol{x}_t\}$, which enables the derivation of a closed-form posterior $p(\boldsymbol{x}_{t-1} \mid \boldsymbol{x}_t, \boldsymbol{y}_{t-1})$. This coupling allows for accurate and recursive sampling based on a well-defined posterior distribution. Empirical results demonstrate that C-DPS consistently outperforms existing baselines, both qualitatively and quantitatively, across multiple inverse problem benchmarks.
Shayan Mohajer Hamidi, Ben Liang 0001, En-Hui Yang
NeurIPS1
2025 A coded knowledge distillation framework for image classification based on adaptive JPEG encoding
abstract
In knowledge distillation (KD), a lightweight student model yields enhanced test accuracy by mimicking the behaviour of a pre-trained large model (teacher). However, the cumbersome teacher model often makes over-confident responses, resulting in poor generalization when presented with unseen data. Consequently, a student trained by such a teacher also inherits this problem. To mitigate this issue, in this paper, we present a new framework of KD dubbed coded knowledge distillation (CKD) in which the student is trained to mimic instead the behaviour of a coded teacher. Compared to the teacher in KD, the coded teacher in CKD has an additional adaptive encoding layer in the front, which adaptively encodes an input image into a compressed version (using JPEG encoding for instance) and then feeds the compressed input image to the pre-trained teacher. Comprehensive experimental results show the effectiveness of CKD over KD. In addition, we extend the deployment of a coded teacher to other knowledge transfer methods, showcasing its ability to enhance test accuracy across these methods.
Ahmed H. Salamah, Shayan Mohajer Hamidi, En-Hui Yang
Pattern Recognit.2
2025 Coded Deep Learning: Framework and Algorithm
abstract
The success of deep learning (DL) is often achieved at the expense of large model sizes and high computational complexity during both training and post-training inferences, making it difficult to train and run large models in a resource-limited environment. To alleviate these issues, this paper introduces a new framework dubbed “coded deep learning” (CDL), which integrates information-theoretic coding concepts into the inner workings of DL, aiming to substantially compress model weights and activations, reduce computational complexity at both training and post-training inference stages, and enable efficient model/data parallelism. Specifically, within CDL, (i) we first propose a novel probabilistic method for quantizing both model weights and activations, and its soft differentiable variant which offers an analytic formula for gradient calculation during training; (ii) both the forward and backward passes during training are executed over quantized weights and activations, which eliminates a majority of floating-point operations and reduces the training computation complexity; (iii) during training, both weights and activations are entropy constrained so that they are compressible in an information-theoretic sense at any stage of training, which in turn reduces communication costs in cases where model/data parallelism is adopted; and (iv) the trained model in CDL is by default in a quantized format with compressible quantized weights, reducing post-training inference complexity and model storage complexity. Additionally, a variant of CDL, namely relaxed CDL (R-CDL), is presented to further improve the trade-off between validation accuracy and compression at the disadvantage of full precision operation involved in forward and backward passes during training with other advantageous features of CDL intact. Extensive empirical results show that CDL and R-CDL outperform the state-of-the-art algorithms in DNN compression in the literature.
En-Hui Yang, Shayan Mohajer Hamidi
IEEE Trans. Inf. Theory2
2025 Conditional Mutual Information Constrained Deep Learning for Classification
abstract
The concepts of conditional mutual information (CMI) and normalized CMI (NCMI) are introduced to measure the concentration and separation performance of a classification deep neural network (DNN) in the output probability distribution space of the DNN, where CMI and the ratio between CMI and NCMI represent the intraclass concentration and interclass separation of the DNN, respectively. By using NCMI to evaluate popular DNNs pretrained over CIFAR-100 and ImageNet in the literature, it is shown that their validation accuracies are more or less inversely proportional to their NCMI values. Based on this observation, the standard deep learning (DL) framework is further modified to minimize the standard cross entropy (CE) function subject to an NCMI constraint, yielding CMI constrained DL (CMIC-DL). A novel alternating learning algorithm is proposed to solve such a constrained optimization problem. Extensive experimental results show that DNNs trained within CMIC-DL outperform the state-of-the-art models trained within the standard DL and other loss functions in the literature in terms of both accuracy and robustness against adversarial attacks. In addition, visualizing the evolution of the learning process through the lens of CMI and NCMI is also advocated.
En-Hui Yang, Shayan Mohajer Hamidi, Linfeng Ye, Renhao Tan, Beverly Yang
IEEE Trans. Neural Networks Learn. Syst.2
2024 How to Train the Teacher Model for Effective Knowledge Distillation
Shayan Mohajer Hamidi, Xizhen Deng, Renhao Tan, Linfeng Ye, Ahmed H. Salamah
ECCV (89)1
2024 Robustness Against Adversarial Attacks Via Learning Confined Adversarial Polytopes
abstract
Deep neural networks (DNNs) could be deceived by generating human-imperceptible perturbations of clean samples. Therefore, enhancing the robustness of DNNs against adversarial attacks is a crucial task. In this paper, we aim to train robust DNNs by limiting the set of outputs reachable via a norm-bounded perturbation added to a clean sample. We refer to this set as adversarial polytope, and each clean sample has a respective adversarial polytope. Indeed, if the respective polytopes for all the samples are compact such that they do not intersect the decision boundaries of the DNN, then the DNN is robust against adversarial samples. Hence, the inner-working of our algorithm is based on learning confined adversarial polytopes (CAP). By conducting a thorough set of experiments, we demonstrate the effectiveness of CAP over existing adversarial robustness methods in improving the robustness of models against state-of-the-art attacks including AutoAttack.
Shayan Mohajer Hamidi, Linfeng Ye
ICASSP1
2024 Bayes Conditional Distribution Estimation for Knowledge Distillation Based on Conditional Mutual Information
abstract
It is believed that in knowledge distillation (KD), the role of the teacher is to provide an estimate for the unknown Bayes conditional probability distribution (BCPD) to be used in the student training process. Conventionally, this estimate is obtained by training the teacher using maximum log-likelihood (MLL) method. To improve this estimate for KD, in this paper we introduce the concept of conditional mutual information (CMI) into the estimation of BCPD and propose a novel estimator called the maximum CMI (MCMI) method. Specifically, in MCMI estimation, both the log-likelihood and CMI of the teacher are simultaneously maximized when the teacher is trained. In fact, maximizing the teacher's CMI value ensures that the teacher can effectively capture the contextual information within the images, and for visualizing this information, we deploy Eigen-CAM. Via conducting a thorough set of experiments, we show that by employing a teacher trained via MCMI estimation rather than one trained via MLL estimation in various state-of-the-art KD frameworks, the student's classification accuracy consistently increases, with the gain of up to 3.32\%. This suggests that the teacher's BCPD estimate provided by MCMI method is more accurate than that provided by MLL method. In addition, we show that such improvements in the student's accuracy are more drastic in zero-shot and few-shot settings. Notably, the student's accuracy increases with the gain of up to 5.72\% when 5\% of the training samples are available to student (few-shot), and increases from 0\% to as high as 84\% for an omitted class (zero-shot).
Linfeng Ye, Shayan Mohajer Hamidi, Renhao Tan, En-Hui Yang
ICLR2
2024 Fed-IT: Addressing Class Imbalance in Federated Learning through an Information- Theoretic Lens
abstract
Federated learning (FL) is a promising technology wherein edge devices/clients collaboratively train a machine learning model under the orchestration of a central server. However, due to the inherent data heterogeneity among clients, local datasets on individual clients often exhibit class imbalance, i.e., samples from majority classes vastly outnumber those from minority classes. This imbalance significantly diminishes the performance of the trained model. To understand why, we first closely examine the output probability distribution clusters of the local deep neural networks (DNNs) in the probability space over the label set, and observe that for class imbalanced datasets, FL has two interesting phenomena: (1) dispersion problem-clusters corresponding to minority classes tend to disperse; and (2) gravity problem-clusters corresponding to minority classes are drawn toward those of majority classes. To overcome these two problems, we then introduce information quantities into FL, propose a new information theoretic loss function for FL, and develop a new FL framework called Fed-IT. It is shown that Fed-IT significantly outperforms previous counterparts, while maintaining client privacy.
Shayan Mohajer Hamidi, Renhao Tan, Linfeng Ye, En-Hui Yang
ISIT1
2024 Conditional Mutual Information Constrained Deep Learning: Framework and Preliminary Results
abstract
In this paper, we introduce the notions of conditional mutual information (CMI) and normalized conditional mutual information (NCMI) for classification deep neural networks (DNNs). In particular, CMI and the ratio between CMI and NCMI quantify the intra-class concentration and inter-class separation of a DNN in its output probability distribution space, respectively. Utilizing NCMI to assess widely recognized DNNs pre-trained on ImageNet reveals a notable inverse relationship between their validation accuracies and NCMI values on the ImageNet validation dataset. Building upon this insight, the conventional deep learning (DL) framework is modified by minimizing the standard cross-entropy function while imposing an NCMI constraint. This refinement results in a novel approach known as CMI-constrained deep learning (CMIC-DL). Comprehensive experimental findings demonstrate that DNNs trained using CMIC-DL outperform state-of-the-art models trained within standard DL and other loss functions in the literature. In addition, some semantic meaning of CMI is also discovered.
En-Hui Yang, Shayan Mohajer Hamidi, Linfeng Ye, Renhao Tan, Beverly Yang
ISIT2
2024 Training Neural Networks on Remote Edge Devices for Unseen Class Classification
abstract
Conventionally, training a deep neural network (DNN) involves minimizing an empirical risk over a training dataset that comprises a certain number of classes. However, for training more versatile DNNs on edge devices, the training datasets are often updated to contain new classes that were not present in the original dataset. To this end, a naive approach could be to share the training samples corresponding to the newly-added classes with the edge devices. However, this comes at a huge communication cost. To tackle this issue, in this paper, we introduce a training method through which a parameter server (PS), having access to all the training samples including those for the newly-added classes, is able to train remote edge devices that lack access to the training samples for the new classes. To realize this, the PS sends an estimate of Bayes conditional probability distribution (BCPD) of the labels to the edge devices using which they train their local models. Via conducting some experiments, we demonstrate the effectiveness of the proposed method.
Shayan Mohajer Hamidi
IEEE Signal Process. Lett.1