Helya Hosseini

dblp:392/3341 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0003-4628-9006ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 39% GPUs and heterogeneous computing · 29% Reconfigurable computing and FPGAs · 12%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression › transformer compression › attention compression
KV cache pruning
0.912025
MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference · NeurIPS 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression › sparsity
unstructured sparsity
0.912025
MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference · NeurIPS 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
Coruscant: Co-Designing GPU Kernel and Sparse Tensor Core to Advocate Unstructured Sparsity in Efficient LLM Inference · MICRO 2025
GPUs and heterogeneous computing › GPU computing › tensor cores
sparse tensor core
0.912025
Coruscant: Co-Designing GPU Kernel and Sparse Tensor Core to Advocate Unstructured Sparsity in Efficient LLM Inference · MICRO 2025
Reconfigurable computing and FPGAs
FPGA accelerator
0.812024
Acamar: A Dynamically Reconfigurable Scientific Computing Accelerator for Robust Convergence and Minimal Resource Underutilization · MICRO 2024
Electronic design automation
high-level synthesis
0.812024
Acamar: A Dynamically Reconfigurable Scientific Computing Accelerator for Robust Convergence and Minimal Resource Underutilization · MICRO 2024
Hardware accelerators and domain-specific architectures
scientific computing accelerator
0.812024
Acamar: A Dynamically Reconfigurable Scientific Computing Accelerator for Robust Convergence and Minimal Resource Underutilization · MICRO 2024
Hardware accelerators and domain-specific architectures
sparse computation
0.812024
Acamar: A Dynamically Reconfigurable Scientific Computing Accelerator for Robust Convergence and Minimal Resource Underutilization · MICRO 2024
High-performance computing › scientific computing systems
partial differential equation solver
0.212024
Acamar: A Dynamically Reconfigurable Scientific Computing Accelerator for Robust Convergence and Minimal Resource Underutilization · MICRO 2024
High-performance computing
scientific computing
0.212024
Acamar: A Dynamically Reconfigurable Scientific Computing Accelerator for Robust Convergence and Minimal Resource Underutilization · MICRO 2024

Methods — techniques the papers use, named apart from their topics

magnitude-based pruning · 0.9bitmap-based sparse format · 0.9Vitis HLS · 0.8
YearPublicationVenuePosition
2026 Situla: Studying the Interplay of Sparse Formats and CPU/GPU Libraries
abstract
The abundance of sparse matrix formats, vendor libraries, and heterogeneous hardware architectures creates a complex optimization landscape in which performance depends on the interplay among data representation, software implementation, and hardware capabilities. To address this complexity, we present Situla, a comprehensive characterization of sparse-dense and sparse-sparse matrix multiplication across Intel Xeon CPUs, NVIDIA H100 GPUs, and many-core ARM Neoverse processors. Our systematic evaluation of CSR, CSC, BSR, and COO formats demonstrates that hardware superiority is not absolute but is instead contingent on software efficiency and matrix structure. Notably, a custom OpenMP implementation on ARM CPUs can outperform vendor-optimized libraries and achieve GPU-level performance for a significant subset of workloads. By analyzing latency, throughput, and power efficiency across diverse matrices from the SuiteSparse collection, Situla challenges conventional assumptions regarding accelerator dominance and highlights critical trade-offs between format selection, platform choice, and energy consumption. Situla’s findings can serve as a blueprint for developing sparsity-aware runtime systems capable of intelligent, dynamic workload scheduling in largescale heterogeneous environments.
Amirmahdi Namjoo, Sanjali Yadav, Helya Hosseini, Bahar Asgari
ISPASS3
2025 Coruscant: Co-Designing GPU Kernel and Sparse Tensor Core to Advocate Unstructured Sparsity in Efficient LLM Inference
Donghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar Asgari
MICRO2
2025 MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference
abstract
We demonstrate that unstructured sparsity significantly improves KV cache compression for LLMs, enabling sparsity levels up to 70\% without compromising accuracy or requiring fine-tuning. We conduct a systematic exploration of pruning strategies and find per-token magnitude-based pruning as highly effective for both Key and Value caches under unstructured sparsity, surpassing prior structured pruning schemes. The Key cache benefits from prominent outlier elements, while the Value cache surprisingly benefits from a simple magnitude-based pruning despite its uniform distribution. KV cache size is the major bottleneck in decode performance due to high memory overhead for large context lengths. To address this, we use a bitmap-based sparse format and a custom attention kernel capable of compressing and directly computing over compressed caches pruned to arbitrary sparsity patterns, significantly accelerating memory-bound operations in decode computations and thereby compensating for the overhead of runtime pruning and compression. Our custom attention kernel coupled with the bitmap-based format delivers substantial compression of KV cache up to 45\% of dense inference and thereby enables longer context lengths and increased tokens/sec throughput of up to 2.23$\times$ compared to dense inference. Our pruning mechanism and sparse attention kernel is available at https://github.com/dhjoo98/mustafar.
Donghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar Asgari
NeurIPS2
2024 Acamar: A Dynamically Reconfigurable Scientific Computing Accelerator for Robust Convergence and Minimal Resource Underutilization
abstract
Although modern supercomputers are capable of delivering Exaflops now, they do not always achieve their peak performance. For instance, even today's high-end supercom-puters achieve only less than 5% of their peak FLOPS when running HPCG, a benchmark designed to represent real-world scientific computing programs. To improve the efficiency of the key kernels in scientific computing, such as those used in solving partial differential equations, computer architects have begun to expand the applications of domain-specific architectures (DSAs) to scientific computing. However, DSAs that often have a fixed design are not likely to be practical solutions, as one specialized solution cannot fit all the diverse scientific computing workloads, making them less effective. The challenges of hardware inef-ficiency in today's supercomputers and the ineffectiveness of DSAs are further exacerbated by sparsity, a key characteristic of scientific computing workloads. While prior studies have proposed DSA solutions for sparse computations, they too are static and not adaptable to variations in the patterns and levels of sparsity across different scientific workloads. To address these challenges and target not only the diversity of computations in such workloads but also variations in sparsity, we propose Acamar11Acamar /’ rckomarr/ is a binary star system in the constellation of Eridanus., a dynamically reconfigurable accelerator. Acamar is adaptable to various solvers across different workloads and dynamically optimizes the trade-off between resource utilization and latency for sparse computations. The adaptable design also enables selecting a solver that guarantees convergence. We evaluate Acamar based on its Vitis HLS implementation on Xilinx Alveo u55c. Our experiments show a resource utilization and latency improvement up to 3.5 x and 6 x as well as improved performance efficiency and achieved throughput over a static design and Nvidia GTX 1650 Super.
Ubaid Bakhtiar, Helya Hosseini, Bahar Asgari
MICRO2