EDBT 2026 Demo / reviewers in the wild / expert
Helya Hosseini
dblp:392/3341
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0003-4628-9006ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 39% GPUs and heterogeneous computing · 29% Reconfigurable computing and FPGAs · 12% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression › transformer compression › attention compression
KV cache pruning |
0.9 | 1 | 2025 | MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression › sparsity
unstructured sparsity |
0.9 | 1 | 2025 | MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference · NeurIPS 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 1 | 2025 | Coruscant: Co-Designing GPU Kernel and Sparse Tensor Core to Advocate Unstructured Sparsity in Efficient LLM Inference · MICRO 2025 |
GPUs and heterogeneous computing › GPU computing › tensor cores
sparse tensor core |
0.9 | 1 | 2025 | Coruscant: Co-Designing GPU Kernel and Sparse Tensor Core to Advocate Unstructured Sparsity in Efficient LLM Inference · MICRO 2025 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.8 | 1 | 2024 | Acamar: A Dynamically Reconfigurable Scientific Computing Accelerator for Robust Convergence and Minimal Resource Underutilization · MICRO 2024 |
Electronic design automation
high-level synthesis |
0.8 | 1 | 2024 | Acamar: A Dynamically Reconfigurable Scientific Computing Accelerator for Robust Convergence and Minimal Resource Underutilization · MICRO 2024 |
Hardware accelerators and domain-specific architectures
scientific computing accelerator |
0.8 | 1 | 2024 | Acamar: A Dynamically Reconfigurable Scientific Computing Accelerator for Robust Convergence and Minimal Resource Underutilization · MICRO 2024 |
Hardware accelerators and domain-specific architectures
sparse computation |
0.8 | 1 | 2024 | Acamar: A Dynamically Reconfigurable Scientific Computing Accelerator for Robust Convergence and Minimal Resource Underutilization · MICRO 2024 |
High-performance computing › scientific computing systems
partial differential equation solver |
0.2 | 1 | 2024 | Acamar: A Dynamically Reconfigurable Scientific Computing Accelerator for Robust Convergence and Minimal Resource Underutilization · MICRO 2024 |
High-performance computing
scientific computing |
0.2 | 1 | 2024 | Acamar: A Dynamically Reconfigurable Scientific Computing Accelerator for Robust Convergence and Minimal Resource Underutilization · MICRO 2024 |
Methods — techniques the papers use, named apart from their topics
magnitude-based pruning · 0.9bitmap-based sparse format · 0.9Vitis HLS · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Situla: Studying the Interplay of Sparse Formats and CPU/GPU LibrariesabstractThe abundance of sparse matrix formats, vendor libraries, and heterogeneous hardware architectures creates a complex optimization landscape in which performance depends on the interplay among data representation, software implementation, and hardware capabilities. To address this complexity, we present Situla, a comprehensive characterization of sparse-dense and sparse-sparse matrix multiplication across Intel Xeon CPUs, NVIDIA H100 GPUs, and many-core ARM Neoverse processors. Our systematic evaluation of CSR, CSC, BSR, and COO formats demonstrates that hardware superiority is not absolute but is instead contingent on software efficiency and matrix structure. Notably, a custom OpenMP implementation on ARM CPUs can outperform vendor-optimized libraries and achieve GPU-level performance for a significant subset of workloads. By analyzing latency, throughput, and power efficiency across diverse matrices from the SuiteSparse collection, Situla challenges conventional assumptions regarding accelerator dominance and highlights critical trade-offs between format selection, platform choice, and energy consumption. Situla’s findings can serve as a blueprint for developing sparsity-aware runtime systems capable of intelligent, dynamic workload scheduling in largescale heterogeneous environments. Amirmahdi Namjoo, Sanjali Yadav, Helya Hosseini, Bahar Asgari |
ISPASS | 3 |
| 2025 | Coruscant: Co-Designing GPU Kernel and Sparse Tensor Core to Advocate Unstructured Sparsity in Efficient LLM Inference
Donghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar Asgari |
MICRO | 2 |
| 2025 | MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM InferenceabstractWe demonstrate that unstructured sparsity significantly improves KV cache compression for LLMs, enabling sparsity levels up to 70\% without compromising accuracy or requiring fine-tuning. We conduct a systematic exploration of pruning strategies and find per-token magnitude-based pruning as highly effective for both Key and Value caches under unstructured sparsity, surpassing prior structured pruning schemes. The Key cache benefits from prominent outlier elements, while the Value cache surprisingly benefits from a simple magnitude-based pruning despite its uniform distribution. KV cache size is the major bottleneck in decode performance due to high memory overhead for large context lengths. To address this, we use a bitmap-based sparse format and a custom attention kernel capable of compressing and directly computing over compressed caches pruned to arbitrary sparsity patterns, significantly accelerating memory-bound operations in decode computations and thereby compensating for the overhead of runtime pruning and compression. Our custom attention kernel coupled with the bitmap-based format delivers substantial compression of KV cache up to 45\% of dense inference and thereby enables longer context lengths and increased tokens/sec throughput of up to 2.23$\times$ compared to dense inference. Our pruning mechanism and sparse attention kernel is available at https://github.com/dhjoo98/mustafar. Donghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar Asgari |
NeurIPS | 2 |
| 2024 | Acamar: A Dynamically Reconfigurable Scientific Computing Accelerator for Robust Convergence and Minimal Resource UnderutilizationabstractAlthough modern supercomputers are capable of delivering Exaflops now, they do not always achieve their peak performance. For instance, even today's high-end supercom-puters achieve only less than 5% of their peak FLOPS when running HPCG, a benchmark designed to represent real-world scientific computing programs. To improve the efficiency of the key kernels in scientific computing, such as those used in solving partial differential equations, computer architects have begun to expand the applications of domain-specific architectures (DSAs) to scientific computing. However, DSAs that often have a fixed design are not likely to be practical solutions, as one specialized solution cannot fit all the diverse scientific computing workloads, making them less effective. The challenges of hardware inef-ficiency in today's supercomputers and the ineffectiveness of DSAs are further exacerbated by sparsity, a key characteristic of scientific computing workloads. While prior studies have proposed DSA solutions for sparse computations, they too are static and not adaptable to variations in the patterns and levels of sparsity across different scientific workloads. To address these challenges and target not only the diversity of computations in such workloads but also variations in sparsity, we propose Acamar11Acamar /’ rckomarr/ is a binary star system in the constellation of Eridanus., a dynamically reconfigurable accelerator. Acamar is adaptable to various solvers across different workloads and dynamically optimizes the trade-off between resource utilization and latency for sparse computations. The adaptable design also enables selecting a solver that guarantees convergence. We evaluate Acamar based on its Vitis HLS implementation on Xilinx Alveo u55c. Our experiments show a resource utilization and latency improvement up to 3.5 x and 6 x as well as improved performance efficiency and achieved throughput over a static design and Nvidia GTX 1650 Super. Ubaid Bakhtiar, Helya Hosseini, Bahar Asgari |
MICRO | 2 |