VLDB 2026 Research / reviewers in the wild / expert
Nader Al Awar
dblp:254/1685
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-7390-5834ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Speeding up the Local C++ Development Cycle with Header SubstitutionabstractC++ remains one of the most widely used languages in various computing fields, from embedded programming to high-performance computing. While new features are constantly being added to C++, an important aspect of the language that is often overlooked is its compilation time. Merely including a few header files can cause compilation time to increase significantly. An alternative to including header files is using forward declarations; however, the rules for forward declaring classes and functions are non obvious and confusing to most developers. Additionally, forward declaring methods, as well as functions that accept lambdas as arguments, is not possible. In this paper, we present a novel technique, termed Header Substitution, to automatically detect opportunities for forward declarations with the goal of replacing includes of header files and improving compilation time. Header Substitution also introduces function wrappers as an alternative to forward declaring methods and functions with lambda arguments. We implemented Header Substitution in a tool, dubbed Yalla, and applied it to various C++ projects in order to speed up the development cycle, i.e., the debugging, editing, compiling, and rerunning loop, achieving up to a 24.5x speedup when compiling C++ files and a 4.68x speedup of the development cycle. Nader Al Awar, Zijian Yi, George Biros, Milos Gligoric 0001 |
CGO | 1 |
| 2022 | A Multi-GPU Python Solver for Low-Temperature Non-Equilibrium PlasmasabstractThe collisional Boltzmann kinetic equations for low-temperature plasmas find important applications in industry, for example semiconductor processing. Particle-in-cell (PIC) methods are the state-of-the art solvers for such problems, but they can be quite expensive. We present GPU acceleration of PIC codes and ways to increase programming productivity for rapid prototyping and algorithmic exploration. First, we present algorithms that minimize data movement and take advantage of modern GPU architectures. Second, we discuss their HPC implementation using Python-based productivity tools: CuPy, Numba, and PyKokkos. We analyze their performance, interoperability, portability, and overheads. We present performance analysis, comparing different algorithms for the main computational kernels. On a single GPU we observe 1.4ns/particle/time step. We also report scaling results on up to 16 NVIDIA Volta V100 GPUs using MPI. James Almgren-Bell, Nader Al Awar, Dilip S. Geethakrishnan, Milos Gligoric 0001, George Biros |
SBAC-PAD | 2 |
| 2021 | A performance portability framework for PythonabstractKokkos is a programming model for writing performance portable applications for all major high performance computing platforms. It provides abstractions for data management and common parallel operations, allowing developers to write portable high performance code with minimal knowledge of architecture-specific details. Kokkos is implemented as a heavily-templated C++ library. However, C++ is not ideal for rapid prototyping and quick algorithmic exploration. An increasing number of developers use Python for scientific computing, machine learning, and data analytics. In this paper, we present a new Python framework, dubbed PyKokkos, for writing performance portable applications entirely in Python. PyKokkos provides Kokkos-like abstractions that are easier to use and more concise than the C++ interface. We implemented PyKokkos by building a translator from a subset of Python to C++ Kokkos and bridging necessary function calls via automatically generated Python bindings. PyKokkos is also compatible with NumPy, a widely-used high performance Python library. By porting several existing Kokkos applications to PyKokkos, including ExaMiniMD (∼3k lines of code in C++), we show that the latter can achieve efficient execution with low performance overhead. Nader Al Awar, Steven Zhu, George Biros, Milos Gligoric 0001 |
ICS | 1 |
| 2021 | Dynamic Generation of Python Bindings for HPC KernelsabstractTraditionally, high performance kernels (HPKs) have been written in statically typed languages, such as C/C++ and Fortran. A recent trend among scientists—prototyping applications in dynamic languages such as Python—created a gap between the applications and existing HPKs. Thus, scientists have to either reimplement necessary kernels or manually create a connection layer to leverage existing kernels. Either option requires substantial development effort and slows down progress in science. We present a technique, dubbed WayOut, which automatically generates the entire connection layer for HPKs invoked from Python and written in C/C++. WayOut performs a hybrid analysis: it statically analyzes header files to generate Python wrapper classes and functions, and dynamically generates bindings for those kernels. By leveraging the type information available at run-time, it generates only the necessary bindings. We evaluate WayOut by rewriting dozens of existing examples from C/C++ to Python and leveraging HPKs enabled by WayOut. Our experiments show the feasibility of our technique, as well as negligible performance overhead on HPKs performance. Steven Zhu, Nader Al Awar, Mattan Erez, Milos Gligoric 0001 |
ASE | 2 |
| 2021 | Programming and execution models for parallel bounded exhaustive testingabstractBounded-exhaustive testing (BET), which exercises a program under test for all inputs up to some bounds, is an effective method for detecting software bugs. Systematic property-based testing is a BET approach where developers write test generation programs that describe properties of test inputs. Hybrid test generation programs offer the most expressive way to write desired properties by freely combining declarative filters and imperative generators. However, exploring hybrid test generation programs, to obtain test inputs, is both computationally demanding and challenging to parallelize. We present the first programming and execution models, dubbed Tempo, for parallel exploration of hybrid test generation programs. We describe two different strategies for mapping the computation to parallel hardware and implement them both for GPUs and CPUs. We evaluated Tempo by generating instances of various data structures commonly used for benchmarking in the BET domain. Additionally, we generated CUDA programs to stress test CUDA compilers, finding four bugs confirmed by the developers. Nader Al Awar, Kush Jain, Christopher J. Rossbach, Milos Gligoric 0001 |
Proc. ACM Program. Lang. | 1 |
| 2019 | How effective are existing Java API specifications for finding bugs during runtime verification?
Owolabi Legunsen, Nader Al Awar, Wajih Ul Hassan, Grigore Rosu, Darko Marinov |
Autom. Softw. Eng. | 2 |