Ali Karkehabadi

dblp:344/8239 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0001-6075-0880ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 CLS-LCR: Classification Subspace Learning with Learnable Categorical Regularization in Forward Forward Networks
abstract
The Forward–Forward (FF) algorithm provides a biologically motivated alternative to backpropagation by relying on layer-wise local updates computed through forward passes only. Despite its conceptual appeal, FF exhibits limited scalability in deeper networks, where rigid goodness aggregation over the full activation space couples representation learning with discrimination and leads to unstable behavior as depth increases. In this work, we introduce Classification Subspace Learning with Learnable Categorical Regularization in Forward–Forward Networks (CLS-LCR), a structural modification that explicitly decouples feature propagation from goodness computation within each layer. The proposed method partitions activations into a feature subspace and a dedicated CLS subspace for discrimination, and employs a learnable, depth-aware routing mechanism to regulate neuron contributions across layers. By improving the separation between positive and negative goodness signals and mitigating early saturation of discrimination neurons, CLS-LCR enables deeper forward-only architectures to maintain stable learning dynamics. Experiments on MNIST and Fashion-MNIST show that CLS-LCR improves with depth, achieving 96.28% accuracy at depth 8 compared to 89.66% for vanilla FF, while preserving the strictly forward, locally trained nature of the algorithm.
Ali Karkehabadi, Zuxiong Tan, Tooraj Nikoubin, Houman Homayoun, Avesta Sasan
ACM Great Lakes Symposium on VLSI1
2026 PRISM: Pruning via Rectified-gradient Importance and Saliency Mapping - making models sparse for execution on edge
abstract
Deep vision models routinely exceed the memory and latency budgets of edge devices, making pruning a practical necessity. However, existing approaches face a three-way trade-off: methods tailored to specific architectures lack generality, hardware-friendly structured sparsity can hurt accuracy, and accurate importance estimates are often computationally expensive. We present PRISM, a saliency-based pruning framework that resolves this tension by using gated (rectified) gradients to denoise per-sample signals and produce reliable weight-level importance in a single backward pass. These scores accumulate over data and can be aggregated along structural axes—channels, neurons, attention heads, or fixed N: M blocks—so the same criterion supports both unstructured and structured sparsity with linear-time scoring. On ImageNet-1K, PRISM prunes ResNet-50, reducing parameters by 41.9% and MACs by 51.2% while improving Top-1 by +0.66 percentage points; under 2:4 sparsity it reaches 78.2% Top-1. On Transformers, PRISM matches or surpasses strong baselines, e.g., 74.1% Top-1 on DeiT-Tiny with 2: 4 sparsity, and outperforms prior structured methods. By coupling rectified-gradient saliency with lightweight aggregation, PRISM delivers an architecture-agnostic, hardware-aligned, and interpretable route to efficient deep learning across CNNs and ViTs.
Zuxiong Tan, Ali Karkehabadi, Houman Homayoun, Tooraj Nikoubin, Avesta Sasan
ACM Great Lakes Symposium on VLSI2
2025 Unified Gravity Loss for Robust Neural Networks Through Feature Space Optimization
Ali Karkehabadi, Houman Homayoun, Avesta Sasan
ACM Great Lakes Symposium on VLSI1
2025 Energy-Efficient Quantization-Aware Training with Dynamic Bit-Width Optimization
Ali Karkehabadi, Avesta Sasan
ACM Great Lakes Symposium on VLSI1
2024 FFCL: Forward-Forward Net with Cortical Loops, Training and Inference on Edge Without Backpropogation
abstract
The Forward-Forward Learning (FFL) algorithm is a recently proposed solution for training neural networks without needing memory-intensive backpropagation. During training, labels accompany input data, classifying them as positive or negative inputs. Each layer learns its response to these inputs independently. In this study, we enhance the FFL with the following contributions: 1) We optimize label processing by segregating label and feature forwarding between layers, enhancing learning performance. 2) By revising label integration, we enhance the inference process, reduce computational complexity, and improve performance. 3) We introduce feedback loops akin to cortical loops in the brain, where information cycles through and returns to earlier neurons, enabling layers to combine complex features from previous layers with lower-level features, enhancing learning efficiency.
Ali Karkehabadi, Houman Homayoun, Avesta Sasan
ACM Great Lakes Symposium on VLSI1