Noushin Azami

dblp:323/5606 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0002-7771-6480ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Efficient Lossless Compression of Scientific Floating-Point Data on CPUs and GPUs
abstract
The amount of scientific data being produced, transferred, and processed increases rapidly. Whereas GPUs have made faster processing possible, storage limitations and slow data transfers remain key bottlenecks. Data compression can help, but only if it does not create a new bottleneck. This paper presents four new lossless compression algorithms for single- and double-precision data that compress well and are fast even though they are fully compatible between CPUs and GPUs. Averaged over many SDRBench inputs, our implementations outperform most of the 18 compressors from the literature we compare to in compression ratio, compression throughput, and decompression throughput. Moreover, they outperform all of them in either throughput or compression ratio on the two CPUs and two GPUs we used for evaluation. For example, on an RTX 4090 GPU, our fastest code compresses and decompresses at over 500 GB/s while delivering one of the highest compression ratios.
Noushin Azami, Alex Fallin, Martin Burtscher
ASPLOS (1)1
2025 Fast and Effective Lossy Compression on GPUs and CPUs with Guaranteed Error Bounds
abstract
High-throughput data compression is increasingly important for large scientific projects. This paper presents PFPL, a lossy floating-point data compressor with guaranteed error bounds that is fully compatible between CPUs and GPUs. Despite this compatibility, PFPL delivers some of the highest compression and decompression speeds and compression ratios on both single- and double-precision data. For example, using an absolute error bound of$1 \mathrm{E}-3$, it yields a single-precision compression throughput on the SDRBench inputs of$5 \text{GB} / \mathrm{s}$on a Ryzen 2950X CPU and 423 GB/s on an RTX 4090 GPU. This is at least 4.6 times higher than the throughput of seven leading compressors on both devices. Moreover, PFPL's compression ratio is higher than that of all tested GPU codes.
Alex Fallin, Noushin Azami, Sheng Di, Franck Cappello, Martin Burtscher
IPDPS2
2025 Identifying Important Data Transformations for Synthesizing Effective Lossless Compressors
abstract
Efficient data compression is essential in high-performance computing, particularly when managing large-scale datasets. Lossless compression algorithms generally incorporate multiple stages, i.e., data transformations. Our LC synthesis framework contains an extensive library of such transformations that it can combine in any order. We have used some of these transformations to build state-of-the-art CPU/GPU-compatible compressors such as SPspeed and SPratio. In this paper, we evaluate the effectiveness of these transformations in achieving high compression ratios on single- and double-precision datasets. We present a comprehensive ranking of these transformations and explore their impact across different datasets. Our analysis reveals that not all transformations are equally effective across datasets, even across datasets produced by the same simulation code, underscoring the need for customized compressors to maximize compression ratios. Using this ranking, we are able to eliminate unimportant transformations, enabling us to more quickly search for and find well-compressing algorithms with more stages. Additionally, we study the importance of these transformations in each compression-algorithm stage and show that the various stages prefer distinct transformations.
Noushin Azami, Martin Burtscher
ISPASS1
2024 LICO: An Effective, High-Speed, Lossless Compressor for Images
abstract
Due to the large and growing number of photographs taken with progressively higher resolution, the volume of image data being generated, stored, and processed increases steadily. Compressing images losslessly is important for enhancing storage efficiency, expediting transmission, and reducing energy consumption while preserving image quality. This paper presents LICO, a new lossless compression algorithm for color images. On the 30 CLIC’24 images, our serial and parallel implementations of LICO deliver a compression speed that is 13× to 46× and a decompression speed that is 14× to 135× higher than JPEG2000. Furthermore, LICO is both faster and compresses more than PNG, TIFF, BZIP2, GZIP, and Zstandard.
Noushin Azami, Rain Lawson, Martin Burtscher
DCC1
2023 Choosing the Best Parallelization and Implementation Styles for Graph Analytics Codes: Lessons Learned from 1106 Programs
abstract
Graph analytics has become a major workload in recent years. The underlying core algorithms tend to be irregular and data dependent, making them challenging to parallelize. Yet, these algorithms can be implemented and parallelized in many ways for CPUs and even more ways for GPUs. We took 6 key graph algorithms and created hundreds of CUDA, OpenMP, and parallel C++ versions of each of them, most of which have never been described or studied. To determine which parallelization and implementation styles work well and under what circumstances, we evaluated the resulting 1106 programs on 2 GPUs and 2 CPUs using 5 input graphs. Our results show which styles and combinations perform well and which ones should be avoided. We found that choosing the wrong implementation style can yield over a 10× slowdown on average. The worst combinations of styles can cost 6 orders of magnitude in performance.
Yiqian Liu, Noushin Azami, Avery Vanausdal, Martin Burtscher
SC2
2022 The Indigo Program-Verification Microbenchmark Suite of Irregular Parallel Code Patterns
abstract
Irregular programs are found in many domains and tend to exhibit input-dependent control flow and memory accesses. This paper introduces the Indigo suite of important irregular parallel code patterns for testing verification and other tools. We studied many irregular CPU and GPU programs and extracted the key code patterns. Then, we methodically built variations of these patterns to alter the control-flow and memory-access behavior and/or introduce bugs, yielding the thousands of OpenMP and CUDA microbenchmarks in the suite. Indigo includes a set of generators to systematically create an unbounded number of inputs for each microbenchmark, which is essential to exercise the wide range of possible behaviors of input-dependent codes. To manage the millions of code and input combinations, Indigo provides the flexibility to generate user-defined subsets of the suite. Experiments with a subset of buggy and bug-free codes illustrate that irregular programs pose a significant challenge to both static and dynamic program verification tools. Moreover, such tools can perform quite differently across code patterns that contain the same bug.
Yiqian Liu, Noushin Azami, Corbin Walters, Martin Burtscher
ISPASS2