Edward Humes

dblp:328/8564 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0002-3945-0116ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021
YearPublicationVenuePosition
2026 MaGrIP: Magnitude and Gradient-Informed Pruning for Task-Agnostic Large Language Models
abstract
Large Language Models (LLMs) have become foundational tools in natural language processing, achieving state-of-the-art performance across a variety of tasks. However, their immense size and computational requirements make them impractical for deployment in resource-constrained environments, such as edge devices and embedded systems. In this work, we introduce Magnitude and Gradient-Informed Pruning (MaGrIP) , a novel framework for task-agnostic pruning and compression of LLMs. MaGrIP employs a dual-threshold strategy combining magnitude- and gradient-based saliency measures to efficiently prune redundant neurons while retaining task performance. Our results demonstrate the effectiveness of MaGrIP in compressing state-of-the-art models. The compression reduced the total computational complexity of the FFN layers from \(\mathcal {O}(d \cdot h)\) to \(\mathcal {O}((d - q) \cdot h)\) . In terms of model size, our pruning approach significantly reduces both model parameters and storage requirements while maintaining competitive perplexity scores evaluated on WikiText-2. For the Gemma 7B model, our method reduces the total size from 28 GB to 5 GB, while for Gemma 2B, MaGrIP achieves a size reduction from 8 GB to 1.5 GB. MaGrIP furthermore exhibits robust performance across multiple benchmarks, such as BOOLQ, ARC-E, and CSQA. Specifically, the pruned Gemma 7B model at 50% pruning achieved 59.26% accuracy on ARC-E compared to 81.06% for the baseline, and 64.74% accuracy on BoolQ compared to 59.98% for the baseline. Similarly, the pruned Llama 3 8B at 50% pruning achieved 46.76% accuracy on ARC-E compared to 77.57% for the baseline, reflecting the tradeoff between compression and accuracy. LLMs compressed using MaGrIP, when deployed on the Nvidia Jetson Orin Nano, achieved a 2.16× improvement in throughput and a 2.3× improvement in performance compared to baseline LLMs.
Uttej Kallakuri, Edward Humes, Hasib-Al Rashid, Tinoosh Mohsenin
ACM Trans. Embed. Comput. Syst.2
2025 Invited Paper: BitMedViT: Ternary-Quantized Vision Transformer for Medical AI Assistants on the Edge
abstract
Vision Transformers (ViTs) have demonstrated strong capabilities in interpreting complex medical imaging data. However, their significant computational and memory demands pose challenges for deployment in real-time, resource-constrained mobile and wearable devices used in clinical environments. We introduce, BitMedVit, a new class of Edge ViTs serving as medical AI assistants that perform structured analysis of medical images directly on the edge. BitMedVit utilizes ternary-quantized linear layers tailored for medical imaging and combines a training procedure with multi-query attention, preserving stability under ternary weights with low-precision activations. Furthermore, BitMedVit employs task-aware distillation from a high-capacity teacher to recover accuracy lost due to extreme quantization. Lastly, we also present a pipeline that maps the ternarized ViTs to a custom CUDA kernel for efficient memory bandwidth utilization and latency reduction on the Jetson Orin Nano. Finally, BitMedVit achieves 86% diagnostic accuracy (89% SOTA) on MedMNIST across 12 datasets, while reducing model size by 43×, memory traffic by 39×, and enabling 16.8 ms inference at an energy efficiency up to 41× that of SOTA models at 183.62 GOPs/J on the Orin Nano. Our results demonstrate a practical and scientifically grounded route for extreme-precision medical imaging ViTs deployable on the edge, narrowing the gap between algorithmic advances and deployable clinical tools.
Mikolaj Walczak, Uttej Kallakuri, Edward Humes, Xiaomin Lin 0002, Tinoosh Mohsenin
ICCAD3
2025 E2AR: An Energy-Efficient Augmented Reality Framework for Collaborative Multi-Drone Systems
abstract
The safety, energy efficiency, and small size of smart drones have led to the broad use of autonomous Unmanned Aerial Vehicles (UAVs) across various applications, creating opportunities for human-machine collaboration. Machine Learning (ML) algorithms like Neural Networks (NNs) offer promising solutions for vision-based navigation and autonomous systems. However, these algorithms are computationally intensive, making it challenging to deploy them on robots expanded by resource-constrained edge devices with limited computational power and low energy consumption requirements. In this paper, we propose an Energy-Efficient Framework for Video Streaming and Augmented Reality called E2AR to enable ML-based multi-edge device video streaming to a AR device while applying augmented reality to enhance human-machine teaming. For this aim, a YOLO is deployed on the edge device for energy-efficient computation and higher performance. Moreover, video streaming to HoloLens is optimized to improve communication latency and power consumption. To evaluate the proposed method, we implemented it on edge devices such as Crazyflie drone while streaming video to the HoloLens. Crazyflie drones with LiDAR sensors and the GAP8 processor has consisted of an octa-core RISC-V. We measured the power consumption, latency, and core usage of the GAP8 processor while implementing the proposed approach. View a video demonstration of the E2AR concept at: Video.
Mozhgan Navardi, Edward Humes, Tinoosh Mohsenin
SEC2
2024 Resource-Aware Saliency-Guided Differentiable Pruning for Deep Neural Networks
abstract
The increasing demand for efficient deep learning model deployment on Tiny Machine Learning (tinyML) and Edge platforms necessitates the development of methods that enable automated and effective network pruning, tailored to tinyML hardware constraints. In this paper, we present a novel differentiable pruning method that accepts total available memory on a tinyML hardware and employs saliency based measurements to identify and prune less significant connections within a deep neural network (DNN). Our approach integrates network compression within the training process, adapting resource utilization to the specific constraints of FPGAs, particularly focusing on on-chip memory. By leveraging a custom tinyML accelerator, we enable an efficient hardware-software co-design. Our framework further quantizes the model to int-8 to optimize the balance between model size and accuracy, crucial for tinyML applications. The efficacy of our approach is examined for the compression of LeNet and VGG16 DNNs. When compared to similar state of the art pruning techniques, our approach for no drop in accuracy further compresses LeNet by 1.15 ×. In the case of VGG16, compared to the baseline implementation, for a 4% drop in accuracy we compress the model up to 55 ×. A comparative analysis of our FPGA hardware accelerator against leading image classification accelerators emphasizes the merits of our approach with a marked improvement in throughput by 1.46 × for LeNet and energy efficiency by 1.7 × for VGG16.
Uttej Kallakuri, Edward Humes, Tinoosh Mohsenin
ACM Great Lakes Symposium on VLSI2
2022 E2EdgeAI: Energy-Efficient Edge Computing for Deployment of Vision-Based DNNs on Autonomous Tiny Drones
abstract
Artificial Intelligence (AI) and Deep Neural Networks (DNNs) have attracted attention as a solution within autonomous systems fields as they enable applications such as visual perception and navigation. Although cloud-based approaches have already been highly addressed, there is a growing interest in using both AI and DNNs on the edge as this allows for lower latency and avoids the potential security concerns of transmitting data to a remote server. However, deploying DNNs on edge devices is challenging due to the limited computational power available, as well as energy efficiency being of the utmost importance. In this work, we introduce an approach named E2EdgeAI for Energy-Efficient Edge computing that takes advantage of AI for autonomous tiny drones. This approach optimizes the energy efficiency of DNNs by considering the effects of memory access and core utilization on the energy consumption of tiny UAVs. To perform the experiment, we used a tiny drone named Crazyflie with the AI -deck expansion, which includes an octa-core RISC-V processor. The experimental results show the proposed approach reduces the model size by up to 14.4x, improves energy per inference by 78%, and increases energy efficiency by 5.6x. A recorded video for the proposed approach can be found here: Video.
Mozhgan Navardi, Edward Humes, Tinoosh Mohsenin
SEC2