VLDB 2026 Research / reviewers in the wild / expert
Apan Qasem
dblp:10/6778
· DBLP profile ↗
22ranked-venue papers
10as first author
12since 2021 · last 2026
0009-0000-6213-829XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 6 first-author · 4 since 2021Software engineering, systems software and programming languages · 8 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Building Flexible, Scalable AI + X Pathways for STEM-Adjacent Disciplines
Apan Qasem, Cindy Royal, Barbara Hewitt, Elise V. Lambert, Pratheesh Omana Sudhakaran |
COMPSAC | 1 |
| 2026 | Does Semantic Heterogeneity Matter? Investigating Multi-Modal Feature Fusion for HPC Performance Modeling
Alex Ford Schneider, Apan Qasem |
COMPSAC | 2 |
| 2025 | Improving Energy Efficiency of Irregular Workloads with Transformers and Tabular Data DiffusionabstractThis J2C paper is an extension of prior work, "Uncovering Input-Sensitive Energy Bottlenecks in Oversubscribed GPU Workloads" which was published in the Journal of Sustainable Computing (SUSCOM22). The previous paper conducted a study to analyze the energy impact of GPU oversubscription in graph algorithms and developed methods for identifying energy bottlenecks in irregular applications. The proposed work outlined in this paper, extends the previous work by adding the following components to our framework: (i) a data diffusion strategy for generating robust training samples for modeling energy behavior (ii) enhanced prediction with contextualized information using transformer models and (iii) cross-platform optimization for both AMD and NVIDIA platforms. Zarif Sadman, Apan Qasem |
COMPSAC | 3 |
| 2025 | Autotuning CNN Workloads on the Edge: A Hybrid Approach with Cross-Domain EmbeddingsabstractLarge-scale convolution neural networks (CNN) are increasingly being deployed on Edge devices to perform inference tasks in applications from mobile robotics to autonomous driving vehicles. Edge devices, however, are often constrained by limited computational resources and power budgets, making autotuning methods a necessity for CNN deployment. Autotuning CNNs for the Edge poses a unique set of challenges which include reasoning about disjoint search spaces, making accurate trade-offs between power and performance, and reducing tuning overhead.This paper presents the design and implementation of an end-to-end autotuning framework that addresses the aforementioned challenges. The framework leverages state-of-the-art machine learning techniques, including a novel feature fusion method, to explore the combined space of the CNN architecture and compiler optimization parameters under different execution environments. Evaluation on three well-known CNN networks show that our autotuning method can yield as much as a factor of 4.8× improvement in performance/watt over state-of-the-art search-based methods, at substantially lower overheads. Experimental results also provide new insight about the intricate trade-offs between power and performance of CNN workloads deployed on the Edge. Trevor Hanz, Zarif Sadman, Apan Qasem |
COMPSAC | 4 |
| 2025 | A Multi-Tiered Autotuner for Portable Heterogeneous-Compute InterfacesabstractWriting performance portable code has been a longstanding challenge in the high-performance computing (HPC) community. This challenge has been exacerbated by the inclusion of accelerators in HPC systems from multiple competing vendors. To address performance portability issues on accelerator-based heterogeneous systems, AMD and Intel have recently introduced suites of APIs capable of generating code for different accelerators with reasonable performance. While these APIs improve programmer productivity, achieving optimal performance still requires significant manual tuning effort.This paper proposes a multi-tiered approach to autotuning CUDA applications ported to AMD architectures via the Heterogeneous-Compute Interface (HIP). We begin by deconstructing the HIP compilation process, identifying key transformations in converting CUDA to HIP. We pinpoint the native compiler optimizations that most influence the performance of CUDA codes when executed on AMD architectures with the HIP runtime interface. We then construct a combined search space around these transformations and their parameters and explore it using a Genetic Algorithm. Our autotuner is the first system to investigate the HIP autotuning space and includes a regression-based strategy to quickly narrow down the search space to the most promising regions. Experimental results with the Rodinia benchmark suite demonstrate that our autotuning strategy achieves an average speedup of 1.7, with performance improvements of up to a factor of seven on certain applications. Shahriar Ahmed Zisan, Mohammad Nooruddin, Apan Qasem |
COMPSAC | 3 |
| 2025 | Accelerated Autotuning of Deep Learning Workloads with Pretrained Performance ModelsabstractAutotuning has proven effective in improving the execution performance of Deep Neural Networks (DNNs). Notwithstanding, its practicality is often limited by the substantial overhead incurred when exploring large optimization spaces, an issue that becomes especially critical for inference workloads on Edge devices with constrained computational resources and realtime requirements. This paper presents a novel machine learning-driven autotuning framework designed to accelerate the tuning process while maintaining high performance. The proposed approach features a two-stage autotuning pipeline that combines predictive modeling with runtime refinement, significantly reducing search costs. In addition, we introduce an attention-based feature fusion mechanism that integrates semantically distinct feature sets, such as task-level and optimization-level features, enabling more accurate performance prediction across domains. Experimental results demonstrate that our framework can effectively explore search spaces of up to 10126configurations, identifying highperforming points that deliver multi-fold speedups, all within a runtime budget of two minutes and thirty seconds. Trevor Hanz, Apan Qasem |
HPCC | 2 |
| 2022 | YODA: A Pedagogical Tool for Teaching Systems ConceptsabstractComputer science undergraduates often struggle in hardware-oriented courses like Computer Organization and Computer Architecture. Active learning instruments can improve student performance in these classes. Regrettably, few tools exist today to support the creation of active learning teaching material for such courses. Apan Qasem |
SIGCSE (1) | 1 |
| 2022 | Heterogeneous Computing for Undergraduates: Introducing the ToUCH Module RepositoryabstractThe need for increased performance per watt, coupled with the demands of processing diverse workloads, has triggered an industry shift towards heterogeneous computing systems. Integration of high-performance CPUs with energy-efficient GPUs is now common in HPC. Architectural heterogeneity has also permeated other domains such as mobile processing, cloud computing, and the Internet of Things. Machine learning practitioners routinely use accelerators in both training and inference. The move towards heterogeneity presents a significant educational challenge since few current curricula include much about heterogeneous computing except possibly in upper-division electives. The NSF-funded initiative ToUCH: Teaching Undergraduates Collaborative and Heterogeneous computing was conceived to confront this impending challenge (https://touch.cs.txstate.edu). The ToUCH project has several ongoing initiatives to promote and encourage teaching of heterogeneous computing. These include summer bootcamps, faculty training workshops and the design, implementation, and integration of a collection of teaching modules on heterogenous computing. In this workshop, we present modules from the ToUCH repository to incorporate heterogeneous computing into core CS courses taken by all majors (e.g., CS 1, CS 2, Computer Organization, Operating Systems). Attendees will have time to work through lab exercises, assignments and tutorials associated with the modules while we assist. We will provide post-workshop support for instructors interested in adopting the modules. In addition, we will solicit feedback from them to help guide our future module development. Apan Qasem, David P. Bunde |
SIGCSE (2) | 1 |
| 2021 | Characterizing Input-sensitivity in Tightly-Coupled Collaborative Graph AlgorithmsabstractThis paper conducts a study of input-sensitivity in collaborative graph algorithms for CPU-GPU systems with support for Unified Memory. The study, conducted on an extensive set of real-world graphs from the Koblenz Network, identifies three main sources of performance inefficiencies that are influenced by characteristics of the input graph. We develop autotuning methods to specifically address these inefficiencies. We then explore machine learning approaches to characterize the relationship between input graph properties, performance, and optimization parameters. In applying our learned models to a test dataset of 70 real-world graphs on the problems of breadth-first search (BFS) and single-source shortest path (SSSP), we are able to attain 96.33% of the peak performance on BFS, and 99.40% on SSSP when using the top-3 predictions of a neural network. We also attain 95.63% of the peak performance on BFS when using three decision trees trained on different categories of graphs. The performance of the learned models is superior to selecting the most frequent optimal configuration, indicating that the machine learning models were able to successfully correlate our selected configuration parameters with graph attributes. Jacob M. Hope, Mikel Gjergji, Johana Di Girolamo, Marco A. Alvarez, Apan Qasem |
CCGRID | 5 |
| 2021 | Migrating Software from x86 to ARM Architecture: An Instruction Prediction ApproachabstractFor decades, the x86 architecture supported by Intel and AMD has been the dominate target for software development. Recently, ARM has solidified itself as a highly competitive and promising CPU architecture by exhibiting both high performance and low power consumption simultaneously. In the foreseeable future, a copious amount of software will be fully migrated to the ARM architecture or support both x86 and ARM simultaneously. Nevertheless, software ports from x86 to ARM are not trivial for a number of reasons. First, it is time consuming to write code that resolves all compatibility issues for a new architecture. Second, specific hardware (e.g. ARM chips) and supporting toolkits (e.g. libraries and compilers) may not be readily available for developers, which will delay the porting process. Third, it is hard to predict the performance of software before testing it on production chips. In this paper, we strive to tackle these challenges by proposing an instruction prediction method that can automatically generate AARCH64 code from existing x86-64 executables. Although the generated code might not be directly executable, it provides a cheap and efficient solution for developers to estimate certain runtime metrics before actually building, deploying and testing code on an ARM-based CPU. Our experimental results show that AARCH64 instructions derived using prediction can achieve a high Bilingual Evaluation Understudy (BLEU) Score. This indicates a quality match between generated executables and natively ported AARCH64 software. Blake W. Ford, Apan Qasem, Jelena Tesic, Ziliang Zong |
NAS | 2 |
| 2021 | Teaching about Heterogeneous ComputingabstractCS faculty have spent the last several years adding parallel computing to their curricula since essentially all processors sold today have multiple cores. A typical target system is a multicore processor with identical cores. This is currently the main configuration for desktop and laptop systems, but the technology continues to evolve and systems are incorporating several kinds of heterogeneity. Many phone processors include cores of different sizes, with high-performance "fat cores" and lower-performance "thin cores", allowing the phone to vary its power and performance profile over time. Other processors incorporate low-power modes or special instructions for specialized computations. Meanwhile, high-end systems make heavy use of accelerators such as graphics cards. We are at a stage where heterogeneous computing concepts should pervade the curriculum rather than being limited to upper-level courses. David P. Bunde, Apan Qasem, Philip J. Schielke |
SIGCSE | 2 |
| 2021 | A module-based introduction to heterogeneous computing in core courses
Apan Qasem, David P. Bunde, Philip J. Schielke |
J. Parallel Distributed Comput. | 1 |
| 2020 | Intelligent Data Placement on Discrete GPU Nodes with Unified MemoryabstractWith increasing heterogeneity, the importance of data organization within a compute node has grown immensely. Recently, industry vendors have introduced technology that can present a unified shared address space for multiple physical pools of memory. In this paper, we leverage unified memory technology and characterize the performance trade-offs of host and device placement across a range of hybrid application design patterns. We perform a Roofline analysis to establish fundamental performance bounds in collaborative applications and then develop an analytical model that makes profitable placement decisions at the individual data structure level. We integrate the placement model into a runtime system and enable transparent data placement in CUDA/C++ applications. Preliminary experiments yield the following results: (i) placement policies have significant performance impact across hybrid application design paradigms (ii) placement decisions are impacted by the sparsity of data access, page re-migration, amount of latency hiding opportunities and design specific attributes such as the number of pipeline stages, and (iii) intelligent data placement can improve node performance by up to 5x on applications with sparse access patterns. Tanzima Sultana, Blake Allen, Apan Qasem |
PACT | 3 |
| 2017 | Characterizing data organization effects on heterogeneous memory architectures
Apan Qasem, Ashwin M. Aji, Gregory Rodgers |
CGO | 1 |
| 2015 | A Module-based Approach to Adopting the 2013 ACM Curricular Recommendations on Parallel ComputingabstractThe widespread deployment of multicore systems over the last decade has brought about major changes in the software and hardware landscape. The resulting importance of parallel computing is reflected in the 2013 Curriculum Guidelines developed by the joint ACM/IEEE taskforce. The document recommends increased coverage of parallel computing and describes a new Knowledge Area on this topic. These recommendations have already been adopted by several universities in the form of new parallel programming courses. Implementing the recommendations in a complete curriculum, however, poses many challenges, including deciding on existing material to be removed, complying with administrative and ABET requirements, and maintaining caps on graduation credit hours. This paper describes an alternative approach for adopting the 2013 curricular recommendations on parallel computing. Specifically, we use a module based approach that introduces parallel computing concepts and re-iterates them through a series of short, self-contained modules taught across several lower-division courses. Most of these concepts are then combined into a new senior-level capstone course on parallel programming. Each module covers parallelism aspects in the context of a conventional computer science topic, thus enabling us to include parallel computing without a major overhaul of the curriculum. Evaluations conducted during the first year show encouraging results for this early-and-often approach in terms of learning outcomes, student interest, and confidence gains. Martin Burtscher, Wuxu Peng, Apan Qasem, Hongchi Shi, Dan E. Tamir, Heather Thiry |
SIGCSE | 3 |
| 2013 | Improving TLB performance on current chip multiprocessor architectures through demand-driven superpagingabstractSUMMARY Translation Lookaside Buffers (TLBs) can play a critical role in improving the performance of emerging parallel workloads. Most current chip multiprocessor systems include multilevel TLBs and provide support for superpages both at the hardware and software level. Judicious use of superpages can significantly cut down the number of TLB misses and improve overall system performance. However, indiscriminate superpage allocation results in page fragmentation and increased application footprint, which often outweigh the benefits of reduced TLB misses. Previous research has explored policies for smart allocation of superpages from an operating system perspective. This paper presents a compiler‐based strategy for automatic and profitable memory allocation via superpages. A significant advantage of a compiler‐based approach is the availability of data‐reuse information within an application. Our strategy employs data‐locality analysis to estimate the TLB demands for both single‐threaded and multi‐threaded programs and uses this metric to apply selective superpage allocation. Apart from its obvious utility in improving TLB performance, this strategy can be used to improve the effectiveness of certain data‐layout transformations and can be a useful tool in benchmarking and automatic performance tuning. We demonstrate the effectiveness of this strategy with experiments on three multicore platforms on a workload that contains both sequential and parallel applications. Copyright © 2012 John Wiley & Sons, Ltd. Apan Qasem, Joshua Magee |
Softw. Pract. Exp. | 1 |
| 2012 | Automatic Restructuring of GPU Kernels for Exploiting Inter-thread Data Locality
Swapneela Unkule, Christopher Shaltz, Apan Qasem |
CC | 3 |
| 2010 | Exposing Tunable Parameters in Multi-threaded Numerical Code
Apan Qasem, Jichi Guo, Faizur Rahman, Qing Yi |
NPC | 1 |
| 2009 | Balancing Locality and Parallelism on Shared-cache Mulit-core SystemsabstractThe emergence of multi-core systems opens new opportunities for thread-level parallelism and dramatically increases the performance potential of applications running on these systems. However, the state of the art in performance enhancing software is far from adequate in regards to the exploitation of hardware features on this complex new architecture. As a result, much of the performance capabilities of multi-core systems are yet to be realized. This research addresses one facet of this problem by exploring the relationship between data-locality and parallelism in the context of multi-core architectures where one or more levels of cache are shared among the different cores. A model is presented for determining a profitable synchronization interval for concurrent threads that interact in a producer-consumer relationship. Experimental results suggest that consideration of the synchronization window, or the amount of work individual threads can be allowed to do between synchronizations, allows for parallelism- and locality-aware performance optimizations. The optimum synchronization window is a function of the number of threads, data reuse patterns within the workload, and the size and configuration of the last-level of cache that is shared among processing units. By considering these factors, the calculation of the optimum synchronization window incorporates parallelism and data locality issues for maximum performance. Michael Jason Cade, Apan Qasem |
HPCC | 2 |
| 2006 | Profitable loop fusion and tiling using model-driven empirical searchabstractLoop fusion and tiling are both recognized as effective transformations for improving memory performance of scientific applications. However, because of their sensitivity to the underlying cache architecture and their interaction with each other it is difficult to determine a good heuristic for applying these transformations profitably across architectures. In this paper, we present a model-guided empirical tuning strategy for profitable application of loop fusion and tiling. Our strategy consists of a detailed cost model that characterizes the interaction between the two transformations at different levels of the memory hierarchy. The novelty of our approach is in exposing key architectural parameters within the model for automatic tuning through empirical search. Preliminary experiments with a set of applications on four different platforms show that our strategy achieves significant performance improvement over fully optimized code generated by state-of-the-art commercial compilers. The time spent in searching for the best parameters is considerably less than with other search strategies. Apan Qasem, Ken Kennedy |
ICS | 1 |
| 2006 | Automatic tuning of whole applications using direct search and a performance-based transformation system
Apan Qasem, Ken Kennedy, John M. Mellor-Crummey |
J. Supercomput. | 1 |
| 2001 | Using a Swap Instruction to Coalesce Loads and Stores
Apan Qasem, David B. Whalley, Xin Yuan 0001, Robert A. van Engelen |
Euro-Par | 1 |