EDBT 2026 Demo / reviewers in the wild / expert
Md. Musfiqur Rahman Sanim
dblp:372/3077
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0009-8431-3699ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsabstractThe increasing size and complexity of modern deep neural networks (DNNs) pose significant challenges for on-device inference on mobile GPUs, with limited memory and computational resources. Existing DNN acceleration frameworks primarily deploy a weight preloading strategy, where all model parameters are loaded into memory before execution on mobile GPUs. We posit that this approach is not adequate for modern DNN workloads that comprise very large model(s) and possibly execution of several distinct models in succession. In this work, we introduce FlashMem, a memory streaming framework designed to efficiently execute large-scale modern DNNs and multi-DNN workloads while minimizing memory consumption and reducing inference latency. Instead of fully preloading weights, FlashMem statically determines model loading schedules and dynamically streams them on demand, leveraging 2.5D texture memory to minimize data transformations and improve execution efficiency. Experimental results on 11 models demonstrate that FlashMem achieves 2.0× to 8.4× memory reduction and 1.7× to 75.0× speedup compared to existing frameworks, enabling efficient execution of large-scale models and multi-DNN support on resource-constrained mobile GPUs. Zhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu, Miao Yin, Gagan Agrawal, Wei Niu 0002 |
ASPLOS (2) | 2 |
| 2025 | Poster: HeteroSched: Co-Optimizing Scheduling and Parallelization for Deep Learning WorkloadsabstractAs Deep Learning (DL) training has emerged as a dominant workload on clusters in recent years, there has been significant interest in parallelizing individual models and scheduling them efficiently. However, to date, no prior work has taken an integrated approach to addressing both problems while simultaneously accounting for heterogeneity in the execution environment. This paper introduces, demonstrating that the integrated optimization of parallelization and scheduling is both feasible and highly effective. The main insight of our work is to generate a set of possible target resources for each job and to develop corresponding parallelization strategies for each. Each jobstrategy pair is then treated as a scheduling candidate. Scheduling decisions are made by extending prior Linear Programming (LP)based formulations and incorporating new heuristics to reduce decision overhead. Bahram Afsharmanesh, Md. Musfiqur Rahman Sanim, AmirAli Mirian, Gagan Agrawal |
PACT | 2 |
| 2025 | Optimizing 3D Gaussian Splattering for Mobile GPUsabstractImage-based 3D scene reconstruction, which transforms multi-view images into a structured 3D representation of the surrounding environment, is a common task across many modern applications. 3D Gaussian Splatting (3DGS) is a new paradigm to address this problem and offers considerable efficiency as compared to the previous methods. Motivated by this, and considering various benefits of mobile device deployment (data privacy, operating without internet connectivity, and potentially faster responses), this paper develops Texture3dgs, an optimized mapping of 3DGS for a mobile GPU. A critical challenge in this area turns out to be optimizing for the twodimensional (2D) texture cache, which needs to be exploited for faster executions on mobile GPUs. As a sorting method dominates the computations in 3DGS on mobile platforms, the core of Texture3dgs is a novel sorting algorithm where the processing, data movement, and placement are highly optimized for 2D memory. The properties of this algorithm are analyzed in view of a cost model for the texture cache. In addition, we accelerate other steps of the 3DGS algorithm through improved variable layout design and other optimizations. End-to-end evaluation shows that Texture 3 dgs delivers up to $\mathbf{4. 1} \times$ and $\mathbf{1. 7} \times$ speedup for the sorting and overall 3D scene reconstruction, respectively while also reducing memory usage by up to $1.6 \times-$ demonstrating the effectiveness of our design for efficient mobile 3D scene reconstruction. Md. Musfiqur Rahman Sanim, Zhihao Shu, Bahram Afsharmanesh, AmirAli Mirian, Jiexiong Guan, Wei Niu 0002, Bin Ren 0002, Gagan Agrawal |
PACT | 1 |
| 2024 | SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on MobileabstractThis work is motivated by recent developments in Deep Neural Networks, particularly the Transformer architectures underlying applications such as ChatGPT, and the need for performing inference on mobile devices. Focusing on emerging transformers (specifically the ones with computationally efficient Swin-like architectures) and large models (e.g., Stable Diffusion and LLMs) based on transformers, we observe that layout transformations between the computational operators cause a significant slowdown in these applications. This paper presents SmartMem, a comprehensive framework for eliminating most layout transformations, with the idea that multiple operators can use the same tensor layout through careful choice of layout and implementation of operations. Our approach is based on classifying the operators into four groups, and considering combinations of producer-consumer edges between the operators. We develop a set of methods for searching such layouts. Another component of our work is developing efficient memory layouts for 2.5 dimensional memory commonly seen in mobile devices. Our experimental results show that SmartMem outperforms 5 state-of-the-art DNN execution frameworks on mobile devices across 18 varied neural networks, including CNNs, Transformers with both local and global attention, as well as LLMs. In particular, compared to DNNFusion, SmartMem achieves an average speedup of 2.8×, and outperforms TVM and MNN with speedups of 6.9× and 7.9×, respectively, on average. Wei Niu 0002, Md. Musfiqur Rahman Sanim, Zhihao Shu, Jiexiong Guan, Xipeng Shen, Miao Yin, Gagan Agrawal, Bin Ren 0002 |
ASPLOS (3) | 2 |