EDBT 2026 Demo / reviewers in the wild / expert
Seyyed Hasan Mozafari
dblp:163/0019
· DBLP profile ↗
9ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-0360-1747ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Latency-Aware Pruning and Quantization of Self-Supervised Speech Transformers for Edge DevicesabstractThe growing adoption of self-supervised learning transformers for speech (speech SSL) is constrained by their significant computational and memory demands, making deployment on resource-constrained edge devices challenging. We propose a latency-aware compression framework that integrates structured pruning and quantization to address these challenges. Guided by a latency model that considers the combined effects of pruning and quantization, our method dynamically identifies and removes less critical blocks while maintaining task performance, avoiding the inefficiencies of over-pruning and under-pruning seen in prior approaches. Unlike prior methods specialized in either post-training compression without fine-tuning data or in cases where fine-tuning data is available, our method is effective in both settings. Experimental results show that, in task-agnostic compression, our method achieves a 4.2× speedup on the Hikey970 edge development platform, outperforming previous task-agnostic pruning methods in most tasks, while requiring only 21–24 GPU hours—a 3× reduction compared to prior methods. Additionally, our method achieves a lower word error rate of 7.8% using task-specific pruning, while reducing computational overhead by approximately 19.4% in terms of GFLOPs compared to previous task-specific methods. Finally, our method consistently achieves higher accuracy than the state-of-the-art post-training compression approach across various latency speedup constraints, even without fine-tuning data. Seyed Milad Ebrahimipour, Seyyed Hasan Mozafari, James J. Clark, Warren J. Gross, Brett H. Meyer |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | Efficient 1D Grouped Convolution for PyTorch a Case Study: Fast On-Device Fine-Tuning for SqueezeBERTabstractGrouped convolution has been observed to be an effective approximation for convolution in many DNN applications. For example, SqueezeBERT, which is a light and fast BERT language processing model, utilizes 1D grouped convolutions. Though SqueezeBERT is well-optimized for inference on edge devices, it suffers from poor memory management during fine-tuning (training). This results in longer fine-tuning time on resource-limited GPUs compared to the original BERT model, BERT-base, despite being specifically designed for edge devices. We study this behavior and show that this poor memory management originates from the use of 1D grouped convolutions in SqueezeBERT. We re-implement 1D grouped convolutions using fully-connected layers, addressing the poor memory allocation and data locality of 1D grouped convolutions. We show that our method is well-suited for edge devices with limited memory; further, it has a negligible effect on inference speed. When utilizing our method, we observe a 42 % reduction in fine-tuning time for SqueezeBERT on edge devices. Seyyed Hasan Mozafari, James J. Clark, Warren J. Gross, Brett H. Meyer |
ASAP | 1 |
| 2023 | High-Throughput Edge Inference for BERT Models via Neural Architecture Search and PipelineabstractThere has been growing interest in improving the BERT inference throughput on resource-constrained edge devices for a satisfactory user experience. One methodology is to employ heterogeneous computing, which utilizes multiple processing elements to accelerate inference. Another methodology is to deploy Neural Architecture Search (NAS) to find optimal solutions in accuracy-throughput design space. In this paper, for the first time, we incorporate NAS with pipelining for BERT models. We show that performing NAS with pipelining achieves on average 53% higher throughput, compared to NAS with a homogeneous system. Hung-Yang Chang, Seyyed Hasan Mozafari, James J. Clark, Brett H. Meyer, Warren J. Gross |
ACM Great Lakes Symposium on VLSI | 2 |
| 2023 | Training Acceleration of Frequency Domain CNNs Using Activation CompressionabstractReducing the complexity of training convolutional neural networks results in lower energy consumption expended during training, or higher accuracy by admitting a greater number of training epochs within a training time budget. During backpropagation, a considerable amount of temporary data is offloaded from GPU memory to CPU memory, increasing training time. In this paper, we address this training time overhead by introducing an activation compression technique for frequency domain convolutional neural networks. Applying this compression technique on frequency domain AlexNet results in activation compression of 57.7%, and a reduction of training time by 23%, with a negligible effect on classification accuracy. Seyyed Hasan Mozafari, James J. Clark, Warren J. Gross, Brett H. Meyer |
ISCAS | 1 |
| 2022 | Fast Heterogeneous Task Mapping for Reducing Edge DNN LatencyabstractTo meet DNN inference latency constraints on resource-constrained edge devices, we employ heterogeneous computing, utilizing multiple processing elements (e.g. CPU + GPU) accelerate inference. This leads to the challenge of efficiently mapping DNN operations to heterogeneous processing elements. For this task, we introduce a novel genetic algorithm (GA) optimizer. Through intelligent initialization and a customized mutation operation, we are able to evaluate 20x fewer generations while finding superior configurations compared with a baseline GA. Using our mapping optimizer, we find device placement configurations that achieve 15%, 24%, and 31% inference speed-up for BERT, SqueezeBERT, and InceptionV3,respectively. Murray L. Kornelsen, Seyyed Hasan Mozafari, James J. Clark, Brett H. Meyer, Warren J. Gross |
ASAP | 2 |
| 2022 | Work-in-Progress: Utilizing latency and accuracy predictors for efficient hardware-aware NASabstractWith the increased size and complexity of state-of-the-art language models such as BERT, deploying them on resource-constrained devices has become challenging. Latency-aware Neural Architecture Search (NAS) is an effective solution for finding an efficient implementation of complex models that satisfy hardware limitations. However, collecting on-device accuracy and latency feedback would significantly slow down the search process, making NAS impractical. To address this, we propose a low-cost method that models both accuracy and latency of BERT-based models on the target device, NVIDIA Jetson TX2, and removes the hardware-related delays from the search loop. Using a Random Forest regressor, our predictors outperform the state-of-the-art and achieve up to 57x speedup while finding a set of near-optimal models. Negin Firouzian, Seyyed Hasan Mozafari, James J. Clark, Warren J. Gross, Brett H. Meyer |
CODES+ISSS | 2 |
| 2019 | Characterizing the Effectiveness of Hot Sparing on Cost and Performance-per-Watt in Application Specific SIMT
Seyyed Hasan Mozafari, Brett H. Meyer |
Integr. | 1 |
| 2017 | Area, Throughput, and Power Trade-Offs for FPGA- and ASIC-Based Execution Stream CompressionabstractAn emerging trend in safety-critical computer system design is the use of compression—for example, using cyclic redundancy check (CRC) or Fletcher checksum (FC)—to reduce the state that must be compared to verify correct redundant execution. We examine the costs and performance of CRC and FC as compression algorithms when implemented in hardware for embedded safety-critical systems. To do so, we have developed parameterizable hardware-generation tools targeting CRC and two novel FC implementations. We evaluate the resulting designs implemented for FPGA and ASIC and analyze their efficiency. While CRC is often best, FC dominates when high throughput is needed. Maria Isabel Mera, Jonah Caplan, Seyyed Hasan Mozafari, Brett H. Meyer, Peter A. Milder |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2015 | Yield-aware Performance-Cost Characterization for Multi-Core SIMTabstractRedundancy is now routinely allocated in circuits, microarchitectural structures, or at the system level, to mitigate mounting manufacturing yield losses. In this paper, we propose spare lane sharing, which reduces the cost of multi-core SIMT systems by allowing one of two neighboring cores to make use of a redundant lane if necessary. We have evaluated the performance-cost trade-offs of core-, lane-, and shared-lane-sparing under a variety of benchmarks, and found that for nearly all applications shared-lane-sparing outperforms lane-sparing, reducing cost by up to 20%. Seyyed Hasan Mozafari, Brett H. Meyer, Kevin Skadron |
ACM Great Lakes Symposium on VLSI | 1 |