Muhammed Nufail Farooqi

dblp:181/4472 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
0since 2021 · last 2018
0000-0002-1609-5847ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 78% Parallel and multicore computing · 22%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › scientific computing systems
adaptive mesh refinement
0.622018
Phase asynchronous AMR execution for productive and performant astrophysical flows · SC 2018
Perilla: metadata-based optimizations of an asynchronous runtime for adaptive mesh refinement · SC 2016
High-performance computing › scientific computing
scientific computing application
0.312018
Phase asynchronous AMR execution for productive and performant astrophysical flows · SC 2018

Methods — techniques the papers use, named apart from their topics

metadata-based optimization · 0.2
YearPublicationVenuePosition
2018 Phase asynchronous AMR execution for productive and performant astrophysical flows
Muhammed Nufail Farooqi, Tan Nguyen 0001, Weiqun Zhang, Ann S. Almgren, John Shalf, Didem Unat
SC1
2017 Nonintrusive AMR Asynchrony for Communication Optimization
Muhammed Nufail Farooqi, Didem Unat, Tan Nguyen 0001, Weiqun Zhang, Ann S. Almgren, John Shalf
Euro-Par1
2016 Perilla: metadata-based optimizations of an asynchronous runtime for adaptive mesh refinement
abstract
Hardware architecture is increasingly complex, urging the development of asynchronous runtime systems with advance resource and locality management supports. However, these supports may come at the cost of complicating the user interface while programming remains one of the major constraints to wide adoption of asynchronous runtimes in practice. In this paper, we propose a solution that leverages application metadata to enable challenging optimizations as well as to facilitate the task of transforming legacy code to an asynchronous representation. We develop Perilla, a task graph-based runtime system that requires only modest programming effort. Perilla utilizes metadata of an AMR software framework to enable various optimizations at the communication layer without complicating its API. Experimental results with different applications on up to 24K processor cores show that Perilla can realize up to 1.44x speedup over the synchronous code variant. The metadata enabled optimizations account for 25% to 100% of the performance improvement.
Tan Nguyen 0001, Didem Unat, Weiqun Zhang, Ann S. Almgren, Muhammed Nufail Farooqi, John Shalf
SC5