EDBT 2026 Demo / reviewers in the wild / expert
M. Suhail Rehman
dblp:05/7063 · also Mohammed Suhail Rehman
· DBLP profile ↗
7ranked-venue papers
5as first author
1since 2021 · last 2021
0000-0002-1765-1753ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data integration and cleaning · 82% Data mining · 18% | |
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% |
Topics — the 2 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Empirical software engineering
data science workflows |
0.4 | 1 | 2019 | Towards Understanding Data Analysis Workflows using a Large Notebook Corpus · SIGMOD Conference 2019 |
Data mining
exploratory data analysis |
0.1 | 1 | 2019 | Towards Understanding Data Analysis Workflows using a Large Notebook Corpus · SIGMOD Conference 2019 |
Methods — techniques the papers use, named apart from their topics
retrospective lineage inference · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | A Demonstration of Relic: A System for REtrospective Lineage InferenCe of Data WorkflowsabstractThe ad-hoc, heterogeneous process of modern data science typically involves loading, cleaning, and mutating dataset(s) into multiple versions recorded as artifacts by various tools within a single data science workflow. Lineage information, including the source datasets, data transformation programs or scripts, or manual annotations, is rarely captured, making it difficult to infer the relationships between artifacts in a given workflow retrospectively. We demonstrate Relic, a tool to retrospectively infer the lineage of data artifacts generated as a result of typical data science workflows, with an interactive demonstration that allows users to input artifact files and visualize the inferred lineage in a web-based setting. M. Suhail Rehman, Silu Huang, Aaron J. Elmore |
Proc. VLDB Endow. | 1 |
| 2019 | Towards Understanding Data Analysis Workflows using a Large Notebook CorpusabstractThe advent of big data analysis as a profession as well as a hobby has brought an increase in novel forms of data exploration and analysis, particularly ad-hoc analysis. Analysis of raw datasets using frameworks such as pandas and R have become very popular [8]. Typically these types of workflows are geared towards ingesting and transforming data in an exploratory fashion in order to derive knowledge while minimizing time-to-insight. However, there exists very little work studying usability and performance concerns of such unstructured workflows. M. Suhail Rehman |
SIGMOD Conference | 1 |
| 2015 | A Cloud Computing Course: From Systems to ServicesabstractWe have designed, developed and administered a course on cloud computing that was taught to over 700 students at our institution over two years. The goal of this project-based course is to provide students with foundational systems concepts as well as experience in developing the required skills to design and deploy viable, robust and elastic web-services within performance and budgetary constraints. We present our objectives, learning outcomes, projects, learning model, outcomes and lessons learned. So far, for this demanding course, our student retention rate is above 80% and enrollment is doubling every year. M. Suhail Rehman, Jason Boles, Mohammad Hammoud, Majd F. Sakr |
SIGCSE | 1 |
| 2012 | Center-of-Gravity Reduce Task Scheduling to Lower MapReduce Network TrafficabstractMapReduce is by far one of the most successful realizations of large-scale data-intensive cloud computing platforms. MapReduce automatically parallelizes computation by running multiple map and/or reduce tasks over distributed data across multiple machines. Hadoop is an open source implementation of MapReduce. When Hadoop schedules reduce tasks, it neither exploits data locality nor addresses partitioning skew present in some MapReduce applications. This might lead to increased cluster network traffic. In this paper we investigate the problems of data locality and partitioning skew in Hadoop. We propose Center-of-Gravity Reduce Scheduler (CoGRS), a locality-aware skew-aware reduce task scheduler for saving MapReduce network traffic. In an attempt to exploit data locality, CoGRS schedules each reduce task at its center-of-gravity node, which is computed after considering partitioning skew as well. We implemented CoGRS in Hadoop-0.20.2 and tested it on a private cloud as well as on Amazon EC2. As compared to native Hadoop, our results show that CoGRS minimizes off-rack network traffic by averages of 9.6% and 38.6% on our private cloud and on an Amazon EC2 cluster, respectively. This reflects on job execution times and provides an improvement of up to 23.8%. Mohammad Hammoud, M. Suhail Rehman, Majd F. Sakr |
IEEE CLOUD | 2 |
| 2010 | Initial Findings for Provisioning Variation in Cloud ComputingabstractCloud computing offers a paradigm shift in management of computing resources for large-scale applications. Using the Infrastructure-as-a-service (IaaS) cloud computing model, users today can request dynamically provisioned, virtualized resources such as CPU, memory, disk, and network access in the form of virtualized resources. The client typically requests resources based on computational needs and pays for resource instances based on their capacity and time utilized. Mapping these virtual resource requests to physical hardware could vary for identical requests. This can potentially cause variations in the performance of applications deployed on such resources. The performance of the application can vary according to the physical layout of the provisioned hardware (the number of virtual machines (VMs), the size/configuration of the VMs and the inter-VM locality). In this paper, we study the effects of this “provisioning variation” and its impact on application performance using suitable benchmarks as well as demonstrate their effect on a few MapReduce workloads. Our initial findings indicate that provisioning variation can impact performance by a factor of 5 primarily due to I/O contention. M. Suhail Rehman, Majd F. Sakr |
CloudCom | 1 |
| 2009 | A performance prediction model for the CUDA GPGPU platformabstractThe significant growth in computational power of modern Graphics Processing Units (GPUs) coupled with the advent of general purpose programming environments like NVIDIA's CUDA, has seen GPUs emerging as a very popular parallel computing platform. Till recently, there has not been a performance model for GPGPUs. The absence of such a model makes it difficult to definitively assess the suitability of the GPU for solving a particular problem and is a significant impediment to the mainstream adoption of GPUs as a massively parallel (super)computing platform. In this paper we present a performance prediction model for the CUDA GPGPU platform. This model encompasses the various facets of the GPU architecture like scheduling, memory hierarchy, and pipelining among others. We also perform experiments that demonstrate the effects of various memory access strategies. The proposed model can be used to analyze pseudo code for a CUDA kernel to obtain a performance estimate, in a way that is similar to performing asymptotic analysis. We illustrate the usage of our model and its accuracy with three case studies: matrix multiplication, list ranking, and histogram generation. Kishore Kothapalli, Rishabh Mukherjee, M. Suhail Rehman, Suryakant Patidar, P. J. Narayanan, K. Srinathan 0001 |
HiPC | 3 |
| 2009 | Fast and scalable list ranking on the GPUabstractGeneral purpose programming on the graphics processing units (GPGPU) has received a lot of attention in the parallel computing community as it promises to offer the highest performance per dollar. The GPUs have been used extensively on regular problems that can be easily parallelized. In this paper, we describe two implementations of List Ranking, a traditional irregular algorithm that is difficult to parallelize on such massively multi-threaded hardware. We first present an implementation of Wyllie's algorithm based on pointer jumping. This technique does not scale well to large lists due to the suboptimal work done. We then present a GPU-optimized, Recursive Helman-JaJa (RHJ) algorithm. Our RHJ implementation can rank a random list of 32 million elements in about a second and achieves a speedup of about 8-9 over a CPU implementation as well as a speedup of 3-4 over the best reported implementation on the Cell Broadband engine. We also discuss the practical issues relating to the implementation of irregular algorithms on massively multi-threaded architectures like that of the GPU. Regular or coalesced memory accesses pattern and balanced load are critical to achieve good performance on the GPU. M. Suhail Rehman, Kishore Kothapalli, P. J. Narayanan |
ICS | 1 |