VLDB 2026 Research / reviewers in the wild / expert
Yifan Sun 0002
dblp:99/10261-2
· DBLP profile ↗
32ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0003-3532-6521ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 1 first-author · 12 since 2021Software engineering, systems software and programming languages · 11 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Designing AI Peers for Collaborative Mathematical Problem Solving with Middle School Students: A Participatory Design StudyabstractCollaborative problem solving (CPS) is a fundamental practice in middle-school mathematics education; however, student groups frequently stall or struggle without ongoing teacher support. Recent work has explored how Generative AI tools can be designed to support one-on-one tutoring, but little is known about how AI can be designed as peer learning partners in collaborative learning contexts. We conducted a participatory design study with 24 middle school students, who first engaged in mathematics CPS tasks with AI peers in a technology probe, and then collaboratively designed their ideal AI peer. Our findings reveal that students envision an AI peer as competent in mathematics yet explicitly deferential, providing progressive scaffolds such as hints and checks under clear student control. Students preferred a tone of friendly expertise over exaggerated personas. We also discuss design recommendations and implications for AI peers in middle school mathematics CPS. Wenhan Lyu, Murong Yue, Yifan Sun 0002, Jennifer Suh, Meredith Kier, Ziyu Yao 0002, Yixuan Zhang 0001 |
CHI | 4 |
| 2026 | QuCo: Efficient and Flexible Hardware-Driven Automatic Configuration of Tile Transfers in GPUsabstractThe growing complexity and parallelism demands of modern GPU workloads have driven architectural innovations toward asynchronous tile transfers (ATTs) to overlap computation and data movement. While ATT units such as the NVIDIA's Tensor Memory Accelerator (TMA) introduce high-throughput memory transfers, programmers must deal with wavefront specialization, select tile sizes, queue slots, and synchronization primitives, all of which are hardware-specific and workloaddependent. Existing GPU libraries fall short-offering limited ATT support and configurability-so developers still resort to manual exploration of this vast parameter space, which is laborious, error-prone, and fundamentally limits performance portability across GPUs. In this work, we present QuCo (Queue Configurator), a single lightweight hardware unit embedded in the GPU that fully automates the ATT configuration process. Inspired by Blackwell GPU design, QuCo includes a compact RISC-V processor, small memory structures for instructions and data, and a GPU Specification Table (GST) storing key architectural parameters. Using the GST and workload characteristics, along with built-in heuristics, QuCo computes optimal queue configurations at kernel launch. This relieves the programmer of the tedious, time-consuming task of tuning and offline profiling, while simultaneously increasing post-compilation performance portability. Nicolás Meseguer, Daoxuan Xu, Yifan Sun 0002, Michael Pellauer, José L. Abellán, Manuel E. Acacio |
HPCA | 3 |
| 2026 | HDPAT: Hierarchical Distributed Page Address Translation for Wafer-Scale GPUsabstractA Wafer-scale GPU connects a large number of chiplets via a high-bandwidth, low-latency interposer-based network, promising to overcome the communication bottleneck of traditional multi-GPU systems. While prior work has prototyped wafer-scale GPUs to demonstrate technical feasibility, scaling to massive chiplet counts creates new bottlenecks: virtual-tophysical address translation becomes severely constrained by massive concurrent requests and long multi-hop network latencies. We propose HDPAT, a hardware-accelerated distributed address translation system that addresses this challenge through three complementary techniques: (1) Concentric caching converts near-IOMMU chiplets into hierarchical translation caches based on their distance to the IOMMU. A lightweight rotation mechanism ensures that there is always a nearby chiplet that can provide translation caching. (2) The redirection table further reduces the burden of IOMMU by delegating translations to caching chiplets, and (3) Prefetching proactively delivers potentially needed address translation into the chiplet to improve translation cache hit rate. Experimental results on 14 representative workloads show that HDPAT improves overall performance by an average of$1.57 \times$. Daoxuan Xu, Ying Li 0049, Yuwei Sun, Jie Ren 0015, Yifan Sun 0002 |
HPCA | 5 |
| 2025 | NetCrafter: Tailoring Network Traffic for Non-Uniform Bandwidth Multi-GPU SystemsabstractMultiple Graphics Processing Units (GPUs) are being integrated into systems to meet the computing demands of emerging workloads.To continuously support more GPUs in a system, it is important to connect them efficiently and effectively.To this end, emerging multi-GPU systems are adopting a hierarchical approach -a group of GPUs with high affinity are connected with higher-bandwidth networks, while multiple groups of GPUs are connected with lowerbandwidth networks to support the scaling of GPUs.Unfortunately, such a non-uniform bandwidth configuration leads to significant performance bottlenecks, especially across lower-bandwidth networks.We present NetCrafter, a combination of novel approaches to deal with the network traffic.NetCrafter is based on three observations: a) not all flits in the network fully utilize the network bandwidth, b) not all requested flits are even necessary -they are requested in the hope that their data might be useful later, c) some flits are more latency-sensitive than others and must be prioritized in the network.NetCrafter leverages these observations to reduce the network traffic by stitching compatible flits that are partly filled, and trimming the number of flits by not fetching flits that are unnecessary.NetCrafter also effectively manages network traffic by sequencing flits such that latency-sensitive flits reach their destinations faster.Although our proposed techniques are generic and can be applied to any network, they are especially useful in alleviating the bottlenecks presented by lower-bandwidth networks connecting multiple groups of GPUs.Overall, NetCrafter significantly improves multi-GPU performance, thereby contributing to efficient scaling of GPU-based systems. Amel Fatima, Yifan Sun 0002, Rachata Ausavarungnirun, Adwait Jog |
ISCA | 3 |
| 2025 | TrioSim: A Lightweight Simulator for Large-Scale DNN Workloads on Multi-GPU SystemsabstractDeep Neural Networks (DNNs) have become increasingly capable of performing tasks ranging from image recognition to content generation.The training and inference of DNNs heavily rely on GPUs, as GPUs' massively parallel architecture delivers extremely high computing capability.With the growing complexity of DNNs and the size of training datasets, training DNNs with a large number of GPUs is becoming a prevalent strategy.Researchers have been exploring how to design software and hardware systems for GPU farms to achieve the best utilization, efficiency, and DNN accuracy during training or inference.However, when designing and deploying such systems, designers usually rely on testing on physical hardware platforms equipped with many GPUs, incurring high costs that are almost prohibitive for system designers to test different configurations and designs, even for highly resourceful companies.While an alternative solution is to test on GPU simulators, they are often too slow for these large-scale systems and depend on profiling details collected from real distributed systems to initiate the simulation.To address these challenges, we present TrioSim, a novel lightweight simulator for DNNs on multi-GPU systems.TrioSim combines performance modeling techniques and simulation methods to achieve high flexibility, high simulation speed, and * Part of this work was done while Yuhui Bao and Pranav Vaid were interns at Lightmatter. Ying Li 0049, Yuhui Bao, Gongyu Wang, Xinxin Mei, Pranav Vaid, Anandaroop Ghosh, Adwait Jog, Darius Bunandar, Ajay Joshi, Yifan Sun 0002 |
ISCA | 10 |
| 2025 | The Sparsity-Aware LazyGPU Architecture
Changxi Liu, Miao Yu 0009, Yifan Sun 0002, Trevor E. Carlson |
ISCA | 3 |
| 2025 | Luthier: A Dynamic Binary Instrumentation Framework Targeting AMD GPUsabstractDynamic Binary instrumentation (DBI) is a widely used technique for collecting detailed, fine-grained information from program execution without requiring recompilation or access to the program's source code. DBI provides several benefits over static instrumentation, including full code discovery and the ability to selectively toggle profiling during runtime. Luthier is an open-source DBI framework targeting AMD GPUs, designed to integrate and run seamlessly on the ROCm software stack. During runtime, Luthier allows inspection of loaded GPU code objects and carries out instrumentation by either manually modifying instructions or inserting calls to special device functions (i.e., 'hooks'') at user-specified locations in the program. Luthier hooks allow inspection and modification of the device visible state, and can communicate with the host via host-accessible device memory buffers. Luthier also supports switching between instrumented and un-instrumented versions of a kernel. In this paper, we describe some of the key design challenges we encountered when developing this open-source DBI framework. We then showcase Luthier's user-facing APIs and internal components, providing example usecases implemented using our framework. While Luthier incurs a 50X runtime overhead (on average) when running an instrumented application, this overhead is 10 times lower as compared to the state-of-the-art GPU-based DBI framework, when running equivalent tools on the same workload written in CUDA. Matin Raayai Ardakani, Andrew Nguyen, Ivan Rosales, Daoxuan Xu, Yuwei Sun, Yifan Sun 0002, David R. Kaeli, Norman Rubin |
ISPASS | 6 |
| 2025 | Will Your Next Pair Programming Partner Be Human? An Empirical Evaluation of Generative AI as a Collaborative Teammate in a Semester-Long Classroom SettingabstractGenerative AI (GenAI), especially Large Language Models (LLMs), is rapidly reshaping both programming workflows and computer science education. Many programmers now incorporate GenAI tools into their workflows, including for collaborative coding tasks such as pair programming. While prior research has demonstrated the benefits of traditional pair programming and begun to explore GenAI-assisted coding, the role of LLM-based tools as collaborators in pair programming remains underexamined. In this work, we conducted a mixed-methods study with 39 undergraduate students to examine how GenAI influences collaboration, learning, and performance in pair programming. Specifically, students completed six in-class assignments under three conditions: Traditional Pair Programming (PP), Pair Programming with GenAI (PAI), and Solo Programming with GenAI (SAI). They used both LLM-based inline completion tools (e.g., GitHub Copilot) and LLM-based conversational tools (e.g., ChatGPT). Our results show that students in the PAI condition achieved the highest assignment scores, whereas those in the SAI condition attained the lowest. Additionally, students' attitudes toward LLMs' programming capabilities improved significantly after collaborating with LLM-based tools, and preferences were largely shaped by the perceived usefulness for completing assignments and learning programming skills, as well as the quality of collaboration. Our qualitative findings further reveal that while students appreciated LLM-based tools as valuable pair programming partners, they also identified limitations (e.g., contextual constraints and possibly outdated knowledge bases), and had different expectations compared to human teammates. Students in our study primarily relied on LLM-based tools for syntax clarification and conceptual guidance, while turning to human partners for idea exchanges. Our study provides one of the first empirical evaluations of GenAI as a pair programming collaborator through a comparison of three conditions (PP, PAI, and SAI). We also discuss the design implications and pedagogical considerations for future GenAI-assisted pair programming approaches. Wenhan Lyu, Yifan Sun 0002, Yixuan Zhang 0001 |
L@S | 3 |
| 2024 | Evaluating the Effectiveness of LLMs in Introductory Computer Science Education: A Semester-Long Field StudyabstractThe integration of AI assistants, especially through the development of Large Language Models (LLMs), into computer science education has sparked significant debate, highlighting both their potential to augment student learning and the risks associated with their misuse. An emerging body of work has looked into using LLMs in education, primarily focusing on evaluating the performance of existing models or conducting short-term human subject studies. However, very little work has examined the impacts of LLM-powered assistants on students in entry-level programming courses, particularly in real-world contexts and over extended periods. To address this research gap, we conducted a semester-long, between-subjects study with 50 students using CodeTutor, an LLM-powered assistant developed by our research team. Our study results show that students who used CodeTutor (the "CodeTutor group" as the experimental group) achieved statistically significant improvements in their final scores compared to peers who did not use the tool (the "control group"). Within the CodeTutor group, those without prior experience with LLM-powered tools demonstrated significantly greater performance gain than their counterparts. We also found that students expressed positive feedback regarding CodeTutor's capability to comprehend their queries and assist in learning programming language syntax. However, they had concerns about CodeTutor's limited role in developing critical thinking skills. Over the course of the semester, students' agreement with CodeTutor's suggestions decreased, with a growing preference for support from traditional human teaching assistants. Our findings also show that students turned to CodeTutor for different tasks, including programming task completion, syntax comprehension, and debugging, particularly seeking help for programming assignments. Our analysis further reveals that the quality of user prompts was significantly correlated with CodeTutor's response effectiveness. Building upon these results, we discuss the implications of our findings for the need to integrate Generative AI literacy into curricula to foster critical thinking skills, and turn to examining the temporal dynamics of user engagement with LLM-powered tools. We further discuss the discrepancy between the anticipated functions of tools and students' actual capabilities, which sheds light on the need for tailored strategies to improve educational outcomes. Wenhan Lyu, Tingting (Rachel) Chung, Yifan Sun 0002, Yixuan Zhang 0001 |
L@S | 4 |
| 2024 | Looking into the Black Box: Monitoring Computer Architecture Simulations in Real-Time with AkitaRTMabstractComputer architecture simulators are essential for validating novel chip designs. However, they often provide little transparency during execution. This opaqueness limits the ability of users to identify issues during the simulation, leading to both wasted computational and human time. We address these issues by providing an intuitive user experience during simulation execution. Particularly, we reveal the status of executions and allow users to control the execution of a computer architecture simulator through AkitaRTM, an interactive web-based tool for real-time monitoring of computer architecture simulations. We based its design on the design workflow inefficiencies experienced by computer architects when using simulations. We demonstrate AkitaRTM's utility through two case studies, the second leading to a patch in the simulator. Additionally, we conducted a user study with computer architects, aiming to validate AkitaRTM. We found that, in addition to solving the observed problems, AkitaRTM also provided an educational benefit by making simulators more transparent to users. Based on these findings, we reflect upon the design of AkitaRTM and provide guidance for future human-centered tools in this space. Ali Mosallaei, Katherine E. Isaacs, Yifan Sun 0002 |
MICRO | 3 |
| 2024 | Visual Exploratory Analysis for Designing Large-Scale Network-on-Chip Architectures: A Domain Expert-Led Design StudyabstractVisualization design studies bring together visualization researchers and domain experts to address yet unsolved data analysis challenges stemming from the needs of the domain experts. Typically, the visualization researchers lead the design study process and implementation of any visualization solutions. This setup leverages the visualization researchers' knowledge of methodology, design, and programming, but the availability to synchronize with the domain experts can hamper the design process. We consider an alternative setup where the domain experts take the lead in the design study, supported by the visualization experts. In this study, the domain experts are computer architecture experts who simulate and analyze novel computer chip designs. These chips rely on a Network-on-Chip (NOC) to connect components. The experts want to understand how the chip designs perform and what in the design led to their performance. To aid this analysis, we develop Vis4Mesh, a visualization system that provides spatial, temporal, and architectural context to simulated NOC behavior. Integration with an existing computer architecture visualization tool enables architects to perform deep-dives into specific architecture component behavior. We validate Vis4Mesh through a case study and a user study with computer architecture researchers. We reflect on our design and process, discussing advantages, disadvantages, and guidance for engaging in a domain expert-led design studies. Hang Yan 0003, Katherine E. Isaacs, Yifan Sun 0002 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | What Do We Mean When We Talk about Trust in Social Media? A Systematic ReviewabstractDo people trust social media? If so, why, in what contexts, and how does that trust impact their lives? Researchers, companies, and journalists alike have increasingly investigated these questions, which are fundamental to understanding social media interactions and their implications for society. However, trust in social media is a complex concept, and there is conflicting evidence about the antecedents and implications of trusting social media content, users, and platforms. More problematic is that we lack basic agreement as to what trust means in the context of social media. Addressing these challenges, we conducted a systematic review to identify themes and challenges in this field. Through our analysis of 70 papers, we contribute a synthesis of how trust in social media is defined, conceptualized, and measured, a summary of trust antecedents in social media, an understanding of how trust in social media impacts behaviors and attitudes, and directions for future work. Yixuan Zhang 0001, Joseph D. Gaggiano, Nutchanon Yongsatianchot, Nurul Suhaimi, Miso Kim, Yifan Sun 0002, Jacqueline A. Griffin, Andrea G. Parker |
CHI | 6 |
| 2023 | A Regression-based Model for End-to-End Latency Prediction for DNN Execution on GPUsabstractDeep neural networks (DNNs) have become increasingly popular in many domains as they reduce the requirement for human effort. However, today’s DNN applications suffer from high computational complexity and sub-optimal device utilization. To solve this problem, researchers have been proposing new system design solutions, which require performance models to help them with pre-product concept validation. This paper discusses how to build a simple, yet accurate, performance model for DNNs on GPUs. Our observations demonstrate prevalent linear relationships between the GPU execution times and operation counts of DNNs layers. Our proposed linear-regression-based execution time predictor can make predictions with an error rate of 28%.11This material is based upon work supported in part by the Google Research Scholar Award and William & Mary. This work was performed in part using the computing facilities at William & Mary and Google Cloud. This work was done while Jog was with William & Mary. Jog is currently with the University of Virginia. Ying Li 0049, Yifan Sun 0002, Adwait Jog |
ISPASS | 2 |
| 2023 | Path Forward Beyond Simulators: Fast and Accurate GPU Execution Time Prediction for DNN WorkloadsabstractToday, DNNs’ high computational complexity and sub-optimal device utilization present a major roadblock to democratizing DNNs. To reduce the execution time and improve device utilization, researchers have been proposing new system design solutions, which require performance models (especially GPU models) to help them with pre-product concept validation. Currently, researchers have been utilizing simulators to predict execution time, which provides high flexibility and acceptable accuracy, but at the cost of a long simulation time. Simulators are becoming increasingly impractical to model today’s large-scale systems and DNNs, urging us to find alternative lightweight solutions. Ying Li 0049, Yifan Sun 0002, Adwait Jog |
MICRO | 2 |
| 2023 | Photon: A Fine-grained Sampled Simulation Methodology for GPU WorkloadsabstractGPUs, due to their massively-parallel computing architectures, provide high performance for data-parallel applications. However, existing GPU simulators are too slow to enable architects to quickly evaluate their hardware designs and software analysis studies. Sampled simulation methodologies are one common way to speed up CPU simulation. However, GPUs apply drastically different execution models that challenge the sampled simulation methods designed for CPU simulations. Recent GPU sampled simulation methodologies do not fully take advantage of the GPU’s special architecture features, such as limited types of basic blocks or warps. Moreover, these methods depend on up-front analysis via profiling tools or functional simulation, making them difficult to use. Changxi Liu, Yifan Sun 0002, Trevor E. Carlson |
MICRO | 2 |
| 2023 | Visualization Design Practices in a Crisis: Behind the Scenes with COVID-19 Dashboard CreatorsabstractDuring the COVID-19 pandemic, a number of data visualizations were created to inform the public about the rapidly evolving crisis. Data dashboards, a form of information dissemination used during the pandemic, have facilitated this process by visualizing statistics regarding the number of COVID-19 cases over time. Prior work on COVID-19 visualizations has primarily focused on the design and evaluation of specific visualization systems from technology-centered perspectives. However, little is known about what occurs behind the scenes during the visualization creation processes, given the complex sociotechnical contexts in which they are embedded. Yet, such ecological knowledge is necessary to help characterize the nuances and trajectories of visualization design practices in the wild, as well as generate insights into how creators come to understand and approach visualization design on their own terms and for their own situated purposes. In this research, we conducted a qualitative interview study among dashboard creators from federal agencies, state health departments, mainstream news media outlets, and other organizations that created (often widely-used) COVID-19 dashboards to answer the following questions: how did visualization creators engage in COVID-19 dashboard design, and what tensions, conflicts, and challenges arose during this process? Our findings detail the trajectory of design practices-from creation to expansion, maintenance, and termination-that are shaped by the complex interplay between design goals, tools and technologies, labor, emerging crisis contexts, and public engagement. We particularly examined the tensions between designers and the general public involved in these processes. These conflicts, which often materialized due to a divergence between public demands and standing policies, centered around the type and amount of information to be visualized, how public perceptions shape and are shaped by visualization design, and the strategies utilized to deal with (potential) misinterpretations and misuse of visualizations. Our findings and lessons learned shed light on new ways of thinking in visualization design, focusing on the bundled activities that are invariably involved in human and nonhuman participation throughout the entire trajectory of design practice. Yixuan Zhang 0001, Yifan Sun 0002, Joseph D. Gaggiano, Neha Kumar 0001, Clio Andris, Andrea G. Parker |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | NaviSim: A Highly Accurate GPU Simulator for AMD RDNA GPUsabstractAs GPUs continue to grow in popularity for accelerating demanding applications, such as high-performance computing and machine learning, GPU architects need to deliver more powerful devices with updated instruction set architectures (ISAs) and new microarchitectural features. The introduction of the AMD RDNA architecture is one example where the GPU architecture was dramatically changed, modifying the underlying programming model, the core architecture, and the cache hierarchy. To date, no publicly-available simulator infrastructure can model the AMD RDNA GPU, preventing researchers from exploring new GPU designs based on the state-of-the-art RDNA architecture. Yuhui Bao, Yifan Sun 0002, Zlatan Feric, Michael Tian Shen, Micah Weston, José L. Abellán, Trinayan Baruah, John Kim 0001, Ajay Joshi, David R. Kaeli |
PACT | 2 |
| 2022 | Shifting Trust: Examining How Trust and Distrust Emerge, Transform, and Collapse in COVID-19 Information SeekingabstractDuring crises like COVID-19, individuals are inundated with conflicting and time-sensitive information that drives a need for rapid assessment of the trustworthiness and reliability of information sources and platforms. This parallels evolutions in information infrastructures, ranging from social media to government data platforms. Distinct from current literature, which presumes a static relationship between the presence or absence of trust and people’s behaviors, our mixed-methods research focuses on situated trust, or trust that is shaped by people’s information-seeking and assessment practices through emerging information platforms (e.g., social media, crowdsourced systems, COVID data platforms). Our findings characterize the shifts in trustee (what/who people trust) from information on social media to the social media platform(s), how distrust manifests skepticism in issues of data discrepancy, the insufficient presentation of uncertainty, and how this trust and distrust shift over time. We highlight the deep challenges in existing information infrastructures that influence trust and distrust formation. Yixuan Zhang 0001, Nurul Suhaimi, Nutchanon Yongsatianchot, Joseph D. Gaggiano, Miso Kim, Shivani A. Patel, Yifan Sun 0002, Stacy Marsella, Jacqueline A. Griffin, Andrea G. Parker |
CHI | 7 |
| 2021 | Mapping the Landscape of COVID-19 Crisis VisualizationsabstractIn response to COVID-19, a vast number of visualizations have been created to communicate information to the public. Information exposure in a public health crisis can impact people’s attitudes towards and responses to the crisis and risks, and ultimately the trajectory of a pandemic. As such, there is a need for work that documents, organizes, and investigates what COVID-19 visualizations have been presented to the public. We address this gap through an analysis of 668 COVID-19 visualizations. We present our findings through a conceptual framework derived from our analysis, that examines who, (uses) what data, (to communicate) what messages, in what form, under what circumstances in the context of COVID-19 crisis visualizations. We provide a set of factors to be considered within each component of the framework. We conclude with directions for future crisis visualization research. Yixuan Zhang 0001, Yifan Sun 0002, Lace M. K. Padilla, Sumit Barua, Enrico Bertini, Andrea G. Parker |
CHI | 2 |
| 2021 | GNNMark: A Benchmark Suite to Characterize Graph Neural Network Training on GPUsabstractGraph Neural Networks (GNNs) have emerged as a promising class of Machine Learning algorithms to train on non-euclidean data. GNNs are widely used in recommender systems, drug discovery, text understanding, and traffic forecasting. Due to the energy efficiency and high-performance capabilities of GPUs, GPUs are a natural choice for accelerating the training of GNNs. Thus, we want to better understand the architectural and system-level implications of training GNNs on GPUs. Presently, there is no benchmark suite available designed to study GNN training workloads. In this work, we address this need by presenting GNNMark, a feature-rich benchmark suite that covers the diversity present in GNN training workloads, datasets, and GNN frameworks. Our benchmark suite consists of GNN workloads that utilize a variety of different graph-based data structures, including homogeneous graphs, dynamic graphs, and heterogeneous graphs commonly used in a number of application domains that we mentioned above. We use this benchmark suite to explore and characterize GNN training behavior on GPUs. We study a variety of aspects of GNN execution, including both compute and memory behavior, highlighting major bottlenecks observed during GNN training. At the system level, we study various aspects, including the scalability of training GNNs across a multi-GPU system, as well as the sparsity of data, encountered during training. The insights derived from our work can be leveraged by both hardware and software developers to improve both the hardware and software performance of GNN training on GPUs. Trinayan Baruah, Kaustubh Shivdikar, Shi Dong 0002, Yifan Sun 0002, Saiful A. Mojumder, Kihoon Jung, José L. Abellán, Yash Ukidave, Ajay Joshi, John Kim 0001, David R. Kaeli |
ISPASS | 4 |
| 2021 | Daisen: A Framework for Visualizing Detailed GPU ExecutionabstractAbstract Graphics Processing Units (GPUs) have been widely used to accelerate artificial intelligence, physics simulation, medical imaging, and information visualization applications. To improve GPU performance, GPU hardware designers need to identify performance issues by inspecting a huge amount of simulator‐generated traces. Visualizing the execution traces can reduce the cognitive burden of users and facilitate making sense of behaviors of GPU hardware components. In this paper, we first formalize the process of GPU performance analysis and characterize the design requirements of visualizing execution traces based on a survey study and interviews with GPU hardware designers. We contribute data and task abstraction for GPU performance analysis. Based on our task analysis, we propose Daisen, a framework that supports data collection from GPU simulators and provides visualization of the simulator‐generated GPU execution traces. Daisen features a data abstraction and trace format that can record simulator‐generated GPU execution traces. Daisen also includes a web‐based visualization tool that helps GPU hardware designers examine GPU execution traces, identify performance bottlenecks, and verify performance improvement. Our qualitative evaluation with GPU hardware designers demonstrates that the design of Daisen reflects the typical workflow of GPU hardware designers. Using Daisen, participants were able to effectively identify potential performance bottlenecks and opportunities for performance improvement. The open‐sourced implementation of Daisen can be found at gitlab.com/akita/vis . Supplemental materials including a demo video, survey questions, evaluation study guide, and post‐study evaluation survey are available at osf.io/j5ghq . Yifan Sun 0002, Yixuan Zhang 0001, Ali Mosallaei, Michael D. Shah, Cody Dunne, David R. Kaeli |
Comput. Graph. Forum | 1 |
| 2021 | Spartan: A Sparsity-Adaptive Framework to Accelerate Deep Neural Network Training on GPUsabstractDeep Neural Networks (DNNs) have emerged as an important class of machine learning algorithms, providing accurate solutions to a broad range of applications. Sparsity in activation maps in DNN training presents an opportunity to reduce computations. However, exploiting activation sparsity presents two major challenges: i) profiling activation sparsity during training comes with significant overhead due to computing the degree of sparsity and the data movement; ii) the dynamic nature of activation maps requires dynamic dense-to-sparse conversion during training, leading to significant overhead. In this article, we present Spartan, a lightweight hardware/software framework to accelerate DNN training on a GPU. Spartan provides a cost-effective and programmer-transparent microarchitectural solution to exploit activation sparsity detected during training. Spartan provides an efficient sparsity monitor, a tile-based sparse GEMM algorithm, and a novel compaction engine designed for GPU workloads. Spartan can reduce sparsity profiling overhead by 52.5× on average. For the most compute-intensive layers, i.e., convolutional layers, we can speedup AlexNet by 3.4×, VGGNet-16 by 2.14×, and ResNet-18 by 2.02×, when training on the ImageNet dataset. Shi Dong 0002, Yifan Sun 0002, Nicolas Bohm Agostini, Elmira Karimi, Daniel Lowell, José Cano 0001, José L. Abellán, David R. Kaeli |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | Valkyrie: Leveraging Inter-TLB Locality to Enhance GPU PerformanceabstractProgramming on a GPU has been made considerably easier with the introduction of Virtual Memory features, which support common pointer-based semantics between the CPU and the GPU. However, supporting virtual memory on a GPU comes with some additional costs and overhead, with the largest being from the support for address translation. The fact that a massive number of threads run concurrently on a GPU means that the translation lookaside buffers (TLBs) are oversubscribed most of the time. Our investigation into a diverse set of GPU workloads shows that TLB misses can be extremely high (up to 99%), which inevitably leads to significant performance degradation due to long-latency page-table walks. Our profiling of TLB-sensitive workloads reveals a high degree of page sharing across the different cores of a GPU. In many applications, a page can be accessed in temporal proximity by multiple cores, following similar memory access patterns. To support the inherent sharing present in GPU workloads, we propose Valkyrie, an integrated cooperative TLB prefetching mechanism and an inter L1-TLB probing scheme that can efficiently reduce TLB bottlenecks in GPUs. Our evaluation using a diverse set of GPU workloads reveals that Valkyrie is able to achieve an average speedup of 1.95x, while adding modest hardware overhead. Trinayan Baruah, Yifan Sun 0002, Saiful A. Mojumder, José L. Abellán, Yash Ukidave, Ajay Joshi, Norman Rubin, John Kim 0001, David R. Kaeli |
PACT | 2 |
| 2020 | Introducing Gamettes: A Playful Approach for Capturing Decision-Making for Informing Behavioral ModelsabstractAgent-based simulations are widely used for modeling human behavior in various contexts. However, such simulations may oversimplify human decision-making. We propose the use of Gamettes to extract rich data on human decision-making and help in improving the human behavioral aspects of models underlying agent-based simulations. We show how Gamettes are designed and provide empirical validation for using Gamettes in an experimental supply chain setting to study human decision-making. Our results show that Gamettes are successful in capturing the expected behaviors and patterns in supply chain decisions, and, thus, we find evidence for the capability of Gamettes to inform behavioral models. Omid Mohaddesi, Yifan Sun 0002, Rana Azghandi, Rozhin Doroudi, Sam Snodgrass, Özlem Ergun, Jacqueline A. Griffin, David R. Kaeli, Stacy Marsella, Casper Harteveld |
CHI | 2 |
| 2020 | Griffin: Hardware-Software Support for Efficient Page Migration in Multi-GPU SystemsabstractAs transistor scaling becomes increasingly more difficult to achieve, scaling the core count on a single GPU chip has also become extremely challenging. As the volume of data to process in today's increasingly parallel workloads continues to grow unbounded, we need to find scalable solutions that can keep up with this increasing demand. To meet the need of modern-day parallel applications, multi-GPU systems offer a promising path to deliver high performance and large memory capacity. However, multi-GPU systems suffer from performance issues associated with GPU-to-GPU communication and data sharing, which severely impact the benefits of multi-GPU systems. Programming multi-GPU systems has been made considerably simpler with the advent of Unified Memory which enables runtime migration of pages to the GPU on demand. Current multi-GPU systems rely on a first-touch Demand Paging scheme, where memory pages are migrated from the CPU to the GPU on the first GPU access to a page. The data sharing nature of GPU applications makes deploying an efficient programmer-transparent mechanism for inter-GPU page migration challenging. Therefore following the initial CPU-to-GPU page migration, the page is pinned on that GPU. Future accesses to this page from other GPUs happen at a cache-line granularity - pages are not transferred between GPUs without significant programmer intervention. We observe that this mechanism suffers from two major drawbacks: 1) imbalance in the page distribution across multiple GPUs, and 2) inability to move the page to the GPU that uses it most frequently. Both of these problems lead to load imbalance across GPUs, degrading the performance of the multi-GPU system. To address these problems, we propose Griffin, a holistic hardware-software solution to improve the performance of NUMA multi-GPU systems. Griffin introduces programmer-transparent modifications to both the IOMMU and GPU architecture, supporting efficient runtime page migration based on locality information. In particular, Griffin employs a novel mechanism to detect and move pages at runtime between GPUs, increasing the frequency of resolving accesses locally, which in turn improves the performance. To ensure better load balancing across GPUs, Griffin employs a Delayed First-Touch Migration policy that ensures pages are evenly distributed across multiple GPUs. Our results on a diverse set of multi-GPU workloads show that Griffin can achieve up to a 2.9× speedup on a multi-GPU system, while incurring low implementation overhead. Trinayan Baruah, Yifan Sun 0002, Ali Tolga Dinçer, Saiful A. Mojumder, José L. Abellán, Yash Ukidave, Ajay Joshi, Norman Rubin, John Kim 0001, David R. Kaeli |
HPCA | 2 |
| 2019 | Exploiting Adaptive Data Compression to Improve Performance and Energy-Efficiency of Compute Workloads in Multi-GPU SystemsabstractGraphics Processing Unit (GPU) performance has relied heavily on our ability to scale of number of transistors on chip, in order to satisfy the ever-increasing demands for more computation. However, transistor scaling has become extremely challenging, limiting the number of transistors that can be crammed onto a single die. Manufacturing large, fast and energy-efficient monolithic GPUs, while growing the number of stream processing units on-chip, is no longer a viable solution to scale performance. GPU vendors are aiming to exploit multi-GPU solutions, interconnecting multiple GPUs in the single node with a high bandwidth network (such as NVLink), or exploiting Multi-Chip-Module (MCM) packaging, where multiple GPU modules are integrated in a single package. The inter-GPU bandwidth is an expensive and critical resource for designing multi-GPU systems. The design of the inter-GPU network can impact performance significantly. To address this challenge, in this paper we explore the potential of hardware-based memory compression algorithms to save bandwidth and improve energy efficiency in multi-GPU systems. Specifically, we propose an adaptive inter-GPU data compression scheme to efficiently improve both performance and energy efficiency. Our evaluation shows that the proposed optimization on multi-GPU architectures can reduce the interGPU traffic up to 62%, improve system performance by up to 33%, and save energy spent powering the communication fabric by 45%, on average. Mohammad Khavari Tavana, Yifan Sun 0002, Nicolas Bohm Agostini, David R. Kaeli |
IPDPS | 2 |
| 2019 | MGPUSim: enabling multi-GPU performance modeling and optimizationabstractThe rapidly growing popularity and scale of data-parallel workloads demand a corresponding increase in raw computational power of Graphics Processing Units (GPUs). As single-GPU platforms struggle to satisfy these performance demands, multi-GPU platforms have started to dominate the high-performance computing world. The advent of such systems raises a number of design challenges, including the GPU microarchitecture, multi-GPU interconnect fabric, runtime libraries, and associated programming models. The research community currently lacks a publicly available and comprehensive multi-GPU simulation framework to evaluate next-generation multi-GPU system designs. Yifan Sun 0002, Trinayan Baruah, Saiful A. Mojumder, Shi Dong 0002, Shane Treadway, Yuhui Bao, Spencer Hance, Carter McCardwell, Vincent Zhao, Harrison Barclay, Amir Kavyan Ziabari, Zhongliang Chen, Rafael Ubal, José L. Abellán, John Kim 0001, Ajay Joshi, David R. Kaeli |
ISCA | 1 |
| 2018 | Airavat: Improving energy efficiency of heterogeneous applicationsabstractAn emerging class of applications attempt to make use of both the CPU and GPU in a heterogeneous system. The peak performance for these applications is achieved when both the CPU and GPU are used collaboratively. However, along with this increased gain in performance, power and energy management is a larger challenge. In this paper we address the issue of executing applications that utilize both the CPU and GPU in an energy efficient way. Towards this end, we propose a power management framework named Airavat that tunes the CPU, GPU and memory frequencies, synergestically, in order to improve the energy efficiency of collaborative CPU-GPU applications. Airavat uses machine learning-based prediction models, combined with feedback based Dynamic Voltage and Frequency Scaling to improve the energy efficiency of such applications. We demonstrate our framework on the NVIDIA Jetson TX1 and observe an improvement in terms of Energy Delay Product (EDP) by 24% with negligible performance loss. Trinayan Baruah, Yifan Sun 0002, Shi Dong 0002, David R. Kaeli, Norman Rubin |
DATE | 2 |
| 2018 | Evaluating Performance Tradeoffs on the Radeon Open Compute PlatformabstractGPUs have been shown to deliver impressive computing performance, while also providing high energy efficiency, across a wide range of high-performance and embedded system workloads. However, limited support for efficient communication and synchronization between the CPU and the GPU impacts our ability to fully exploit the benefits of heterogeneous systems. Recently, the Heterogeneous System Architecture (HSA) was introduced to address these issues with synchronization and communication, but given the low-level nature of HSA, it was not easily adopted by the broader programming community. In 2016, AMD described the Radeon Open Compute (ROC) platform that brings high-level programming frameworks such as OpenCL, HC++, and HIP to end users. These high-level programming frameworks offer a simpler programming experience by wrapping complex HSA APIs, while still delivering the power of HSA. To date, there has been little evaluation of the potential performance benefits and trade-offs of leveraging the ROC platform. In this work, we evaluate the performance of the ROC platform using the Hetero-Mark and DNNMark benchmark suites. Equipped with Hetero-Mark, we compare the performance of different programming frameworks, including OpenCL, HC++, and HIP on both integrated APUs and discrete GPUs. We also present three new CPU-GPU collaborative patterns and employ three new benchmarks to evaluate system-level atomics. With DNNMark and a new DNN Face Detection benchmark, we evaluate the performance of ROC libraries including rocBLAS and MIOpen. We also provide guidance on best practices to programmers when developing applications leveraging the ROC platform. Yifan Sun 0002, Saoni Mukherjee, Trinayan Baruah, Shi Dong 0002, Julian Gutierrez 0002, Prannoy Mohan, David R. Kaeli |
ISPASS | 1 |
| 2018 | Characterizing the Microarchitectural Implications of a Convolutional Neural Network (CNN) Execution on GPUsabstractGPUs have become a very popular platform for accelerating the processing involved in deep learning applications. One class of popular variants, Convolutional Neural Networks (CNNs), have been widely deployed to run on GPUs. In many application settings, a GPU has sufficient computing power and memory space to accommodate the dense matrix operations performed during CNN training. However, few characterization studies have considered how CNNs can impact microarchitectural structures in a GPU. In this paper, we perform a characterization of one selected CNN workload as run on two different NVIDIA GPUs from distinct microarchitecture families, highlighting the impact that microarchitecture plays on this important class of workload. First, we analyze the performance implications of a CNN model using microarchitectural details on a layer-by-layer basis, and characterize the memory access behavior in the context of a typical GPU memory hierarchy, considering hardware resource utilization associated with each primitive in the CNN model. We identify major bottlenecks by considering the potential limits of using a single GPU. Additionally, we evaluate a number of optimization approaches, such as L1 cache bypassing and kernel fusion. L1 cache bypassing can achieve up to a 6.2% speedup for a single layer, but manipulating L1 cache provides very limited benefits in terms of application speedup, while kernel fusion provides an overall application speedup of 4.0%, on average. Shi Dong 0002, Yifan Sun 0002, Trinayan Baruah, David R. Kaeli |
ICPE | 3 |
| 2016 | A comprehensive performance analysis of HSA and OpenCL 2.0abstractHeterogeneous systems, that marry CPUs and GPUs together in a range of configurations, are quickly becoming the design paradigm for today's platforms because of their impressive parallel processing capabilities. However, in many existing heterogeneous systems, the GPU is only treated as an accelerator by the CPU, working as a slave to the CPU master. But recently we are starting to see the introduction of a new class of devices and changes to the system runtime model, which enable accelerators to be treated as first-class computing devices. To support programmability and efficiency of heterogeneous programming, the HSA foundation introduced the Heterogeneous System Architecture (HSA), which defines a platform and runtime architecture that provides rich support for OpenCL 2.0 features including shared virtual memory, dynamic parallelism, and improved atomic operations. In this paper, we provide the first comprehensive study of OpenCL 2.0 and HSA 1.0 execution, considering OpenCL 1.2 as the baseline. For workloads, we develop a suite of OpenCL micro-benchmarks designed to highlight the features of these emerging standards and also utilize real-world applications to better understand their impact at an application level. To fully exercise the new features provided by the HSA model, we experiment with a producer-consumer algorithm and persistent kernels. We find that by using HSA signals, we can remove 92% of the overhead due to synchronous kernel launches. In our real-world applications, the OpenCL 2.0 runtime achieves up to a 1.2X speedup, while the HSA 1.0 runtime achieves a 2.7X speedup over OpenCL 1.2. Saoni Mukherjee, Yifan Sun 0002, Paul Blinzer, Amir Kavyan Ziabari, David R. Kaeli |
ISPASS | 2 |
| 2016 | UMH: A Hardware-Based Unified Memory Hierarchy for Systems with Multiple Discrete GPUsabstractIn this article, we describe how to ease memory management between a Central Processing Unit (CPU) and one or multiple discrete Graphic Processing Units (GPUs) by architecting a novel hardware-based Unified Memory Hierarchy (UMH). Adopting UMH, a GPU accesses the CPU memory only if it does not find its required data in the directories associated with its high-bandwidth memory, or the NMOESI coherency protocol limits the access to that data. Using UMH with NMOESI improves performance of a CPU-multiGPU system by at least 1.92 × in comparison to alternative software-based approaches. It also allows the CPU to access GPUs modified data by at least 13 × faster. Amir Kavyan Ziabari, Yifan Sun 0002, Yenai Ma, Dana Schaa, José L. Abellán, Rafael Ubal, John Kim 0001, Ajay Joshi, David R. Kaeli |
ACM Trans. Archit. Code Optim. | 2 |