EDBT 2026 Demo / reviewers in the wild / expert
Mohammad Almasri
dblp:242/8208
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2023
0000-0003-3154-4433ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Parallelizing Maximal Clique Enumeration on GPUsabstractWe present a GPU solution for exact maximal clique enumeration (MCE) that performs a search tree traversal following the Bron-Kerbosch algorithm. Prior works on parallelizing MCE on GPUs perform a breadth-first traversal of the tree, which has limited scalability because of the explosion in the number of tree nodes at deep levels. We propose to parallelize MCE on GPUs by performing depth-first traversal of independent subtrees in parallel. Since MCE suffers from high load imbalance and memory capacity requirements, we propose a worker list for dynamic load balancing, as well as partial induced subgraphs and a compact representation of excluded vertex sets to regulate memory consumption. Our evaluation shows that our GPU implementation on a single GPU outperforms the state-of-the-art parallel CPU implementation by a geometric mean of 4.9× (up to 16.7×), and scales efficiently to multiple GPUs. Our code has been open-sourced to enable further research on accelerating MCE. Mohammad Almasri, Yen-Hsiang Chang 0001, Izzat El Hajj, Rakesh Nagi, Jinjun Xiong, Wen-Mei W. Hwu |
PACT | 1 |
| 2023 | BEEP: Balanced Efficient subgraph Enumeration in ParallelabstractBEEP is a state-of-the-art subgraph enumerator that delivers high performance through a combination of balanced, parallel GPU processing and novel algorithmic improvements. With a rapidly increasing demand for fast tools on large graphs, GPU-based subgraph enumerators are of growing interest. Most existing GPU enumerators are based on Breadth First Search (BFS), which often impose limitations on hardware resources due to excessive memory requirements. PARSEC [12] was the first GPU enumerator to adopt Depth First Search (DFS) that demonstrated impressive speedups and its adaptability to hardware with limited memory resources. However, PARSEC’s DFS implementation suffers from computational inefficiencies and load imbalances. BEEP introduces novel search space reduction techniques and load balancing strategies to tackle these challenges in DFS-based parallelization and achieves exceptional performance and scalability. Experimental results indicate that BEEP outperforms PARSEC with geometric mean speedups of up to 10.52 × across disparate data graphs and up to 7.28 × across various queries with maximum speedups of 33.46 ×. This makes BEEP the fastest subgraph enumerator to date. Furthermore, a multi-GPU implementation is developed that exhibits almost linear scalability with the number of devices. Samiran Kawtikwar, Mohammad Almasri, Wen-Mei W. Hwu, Rakesh Nagi, Jinjun Xiong |
ICPP | 2 |
| 2022 | Parallel K-clique counting on GPUsabstractCounting k-cliques in a graph is an important problem in graph analysis with many applications such as community detection and graph partitioning. Counting k-cliques is typically done by traversing search trees starting at each vertex in the graph. Parallelizing k-clique counting has been well-studied on CPUs and many solutions exist. However, there are no performant solutions for k-clique counting on GPUs. Mohammad Almasri, Izzat El Hajj, Rakesh Nagi, Jinjun Xiong, Wen-Mei W. Hwu |
ICS | 1 |
| 2022 | PARSEC: PARallel Subgraph Enumeration in CUDAabstractSubgraph enumeration is an important problem in the field of Graph Analytics with numerous applications. The problem is provably NP-complete and requires sophisticated heuristics and highly efficient implementations to be feasible on problem sizes of realistic scales. Parallel solutions have shown a lot of promise on CPUs and distributed environments. Recently, GPU-based parallel solutions have also been proposed to take advantage of the massive execution resources in modern GPUs. Subgraph enumeration involves traversing a search tree for each vertex of the data graph to find matches of a query in a graph. Most GPU-based solutions traverse the tree in breadth-first manner that exploits parallelism at the cost of high memory requirement and presents a formidable challenge for processing large graphs with high-degree vertices since the memory capacity of GPUs is significantly lower than that of CPUs. In this work, we propose a novel GPU solution based on a hybrid BFS and DFS approach where the top level(s) of the search trees are traversed in a fully parallel, breadth-first manner while each subtree is traversed in a more space-efficient, depth-first manner. The depth-first traversal of subtrees requires less memory but presents more challenges for parallel execution. To overcome the less parallel nature of depth-first traversal, we exploit fine-grained parallelism in each step of the depth-first traversal of sub-trees. We further identify and implement various optimizations to efficiently utilize memory and compute resources of the GPUs. We evaluate our performance in comparison with the state-of-the-art GPU and CPU implementations. We outperform the GPU and CPU implementations with a geometric mean speedup of 9.47× (up to 92.01×) and 2.37× (up to 12.70×), respectively. We also show that the proposed approach can efficiently process the graphs that previously cannot be processed by the state-of-the-art GPU solutions due to their excessive memory requirement. Vibhor Dodeja, Mohammad Almasri, Rakesh Nagi, Jinjun Xiong, Wen-Mei W. Hwu |
IPDPS | 2 |
| 2022 | An efficient GPU implementation and scaling for higher-order 3D stencils
Omer Anjum, Mohammad Almasri, Simon Garcia de Gonzalo, Wen-Mei W. Hwu |
Inf. Sci. | 2 |
| 2021 | PhraseScope: An Effective and Unsupervised Framework for Mining High Quality PhrasesabstractPhrase mining is one of the fundamental NLP tasks that can have significant impact on the efficacy of many downstream applications.Many supervised and unsupervised phrase mining approaches have been proposed.Some rely on linguistic analyzers, and others are language agnostic.A daunting challenge in this task is to distinguish quality phrases from noise phrases, which tightly coexists with quality phrases in the entire frequency spectrum.Most existing approaches to phrase mining, however, rely on frequency-based statistics, hence suffer from quality loss.In this paper, we propose an unsupervised phrase mining framework, "PhraseScope", which consists of a sequence of filters, namely cohesion, domain, and graph filters, to remove noise phrase.Each filter is responsible for removing noise phrase of particular characteristics.Collectively, our proposed filters are capable of detecting and removing noise phrases effectively while preserving quality phrases.Our results show significant improvement in both recall and precision over state-of-the-art frameworks when tested on three different domains of datasets. Omer Anjum, Mohammad Almasri, Jinjun Xiong, Wen-Mei W. Hwu |
SDM | 2 |
| 2020 | CCF: An efficient SpMV storage format for AVX512 platforms
Mohammad Almasri, Walid A. Abu-Sufah |
Parallel Comput. | 1 |