EDBT 2026 Demo / reviewers in the wild / expert
Qiumin Xu
dblp:151/5435
· DBLP profile ↗
12ranked-venue papers
7as first author
2since 2021 · last 2024
0000-0003-1391-3397ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 6 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 3 first-authorArtificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Memory systems · 40% GPUs and heterogeneous computing · 26% Performance modeling and evaluation · 9% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Ubiquitous computing and smart environments · 77% Health and well-being technologies · 23% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache coherence |
0.8 | 1 | 2024 | Cross-core Data Sharing for Energy-efficient GPUs · ACM Trans. Archit. Code Optim. 2024 |
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy |
0.8 | 1 | 2024 | Cross-core Data Sharing for Energy-efficient GPUs · ACM Trans. Archit. Code Optim. 2024 |
Memory systems › cache › CPU cache
l1 data cache |
0.8 | 1 | 2024 | Cross-core Data Sharing for Energy-efficient GPUs · ACM Trans. Archit. Code Optim. 2024 |
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
device placement |
0.4 | 1 | 2020 | Transferable Graph Optimizers for ML Compilers · NeurIPS 2020 |
Machine learning › Efficient and distributed learning
distributed training |
0.4 | 1 | 2020 | Transferable Graph Optimizers for ML Compilers · NeurIPS 2020 |
Compilers and program optimization › compiler optimization
computation graph optimization |
0.4 | 1 | 2020 | Transferable Graph Optimizers for ML Compilers · NeurIPS 2020 |
Compilers and program optimization › compiler optimization
machine learning for compiler optimization |
0.4 | 1 | 2020 | Transferable Graph Optimizers for ML Compilers · NeurIPS 2020 |
Ubiquitous computing and smart environments › smart home
smart home sensing |
0.2 | 1 | 2016 | AirSense: an intelligent home-based sensing system for indoor air quality analytics · UbiComp 2016 |
Cloud and datacenter computing › resource allocation
dynamic resource allocation |
0.2 | 1 | 2016 | Warped-Slicer: Efficient Intra-SM Slicing through Dynamic Resource Partitioning for GPU Multiprogramming · ISCA 2016 |
GPUs and heterogeneous computing
GPU sharing |
0.2 | 1 | 2016 | Warped-Slicer: Efficient Intra-SM Slicing through Dynamic Resource Partitioning for GPU Multiprogramming · ISCA 2016 |
Processor architecture and microarchitecture › many-core architecture
streaming multiprocessor |
0.2 | 1 | 2016 | Warped-Slicer: Efficient Intra-SM Slicing through Dynamic Resource Partitioning for GPU Multiprogramming · ISCA 2016 |
Energy-efficient computing › energy-efficient architecture
GPU energy efficiency |
0.2 | 1 | 2024 | Cross-core Data Sharing for Energy-efficient GPUs · ACM Trans. Archit. Code Optim. 2024 |
Storage systems
flash and SSD |
0.2 | 1 | 2015 | Performance Characterization of Hyperscale Applicationson on NVMe SSDs · SIGMETRICS 2015 |
Performance modeling and evaluation
workload characterization |
0.2 | 1 | 2015 | Performance Characterization of Hyperscale Applicationson on NVMe SSDs · SIGMETRICS 2015 |
Performance modeling and evaluation
analytical modeling |
0.1 | 1 | 2016 | Warped-Slicer: Efficient Intra-SM Slicing through Dynamic Resource Partitioning for GPU Multiprogramming · ISCA 2016 |
Performance modeling and evaluation
storage performance evaluation |
0.1 | 1 | 2015 | Performance Characterization of Hyperscale Applicationson on NVMe SSDs · SIGMETRICS 2015 |
Methods — techniques the papers use, named apart from their topics
sequential attention · 0.9graph neural network · 0.9deep reinforcement learning · 0.9predictor-based cache block estimation · 0.8sensing system · 0.2online profiling · 0.2forecasting · 0.2deployment study · 0.2analytical modeling · 0.2benchmarking · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Cross-core Data Sharing for Energy-efficient GPUsabstractGraphics Processing Units (GPUs) are the accelerator of choice in a variety of application domains, because they can accelerate massively parallel workloads and can be easily programmed using general-purpose programming frameworks such as CUDA and OpenCL. Each Streaming Multiprocessor (SM) contains an L1 data cache (L1D) to exploit the locality in data accesses. L1D misses are costly for GPUs for two reasons. First, L1D misses consume a lot of energy as they need to access the L2 cache (L2) via an on-chip network and the off-chip DRAM in case of L2 misses. Second, L1D misses impose performance overhead if the GPU does not have enough active warps to hide the long memory access latency. We observe that threads running on different SMs share 55% of the data they read from the memory. Unfortunately, as the L1Ds are in the non-coherent memory domain, each SM independently fetches data from the L2 or the off-chip memory into its L1D, even though the data may be currently available in the L1D of another SM. Our goal is to service L1D read misses via other SMs, as much as possible, to cut down costly accesses to the L2 or the off-chip DRAM. To this end, we propose a new data-sharing mechanism, called Cross-Core Data Sharing (CCDS) . CCDS employs a predictor to estimate whether the required cache block exists in another SM. If the block is predicted to exist in another SM’s L1D, then CCDS fetches the data from the L1D that contain the block. Our experiments on a suite of 26 workloads show that CCDS improves average energy and performance by 1.30× and 1.20×, respectively, compared to the baseline GPU. Compared to the state-of-the-art data-sharing mechanism, CCDS improves average energy and performance by 1.37× and 1.11×, respectively. Hajar Falahati, Mohammad Sadrosadati, Qiumin Xu, Juan Gómez-Luna, Banafsheh S. Latibari, Hyeran Jeon, Shaahin Hessabi, Hamid Sarbazi-Azad, Onur Mutlu, Murali Annavaram, Massoud Pedram |
ACM Trans. Archit. Code Optim. | 3 |
| 2022 | A Language Agnostic Multilingual Streaming On-Device ASR SystemabstractOn-device end-to-end (E2E) models have shown improvements over a conventional model on English Voice Search tasks in both quality and latency.E2E models have also shown promising results for multilingual automatic speech recognition (ASR).In this paper, we extend our previous capacity solution to streaming applications and present a streaming multilingual E2E ASR system that runs fully on device with comparable quality and latency to individual monolingual models.To achieve that, we propose an Encoder Endpointer model and an End-of-Utterance (EOU) Joint Layer for a better quality and latency trade-off.Our system is built in a language agnostic manner allowing it to natively support intersentential code switching in real time.To address the feasibility concerns on large models, we conducted on-device profiling and replaced the time consuming LSTM decoder with the recently developed Embedding decoder.With these changes, we managed to run such a system on a mobile device in less than real time. Bo Li 0028, Tara N. Sainath, Ruoming Pang, Shuo-Yiin Chang, Qiumin Xu, Trevor Strohman, Vince Chen, Qiao Liang 0001, Heguang Liu, Yanzhang He, Parisa Haghani, Sameer Bidichandani |
INTERSPEECH | 5 |
| 2020 | Transferable Graph Optimizers for ML CompilersabstractMost compilers for machine learning (ML) frameworks need to solve many correlated optimization problems to generate efficient machine code. Current ML compilers rely on heuristics based algorithms to solve these optimization problems one at a time. However, this approach is not only hard to maintain but often leads to sub-optimal solutions especially for newer model architectures. Existing learning based approaches in the literature are sample inefficient, tackle a single optimization problem, and do not generalize to unseen graphs making them infeasible to be deployed in practice. To address these limitations, we propose an end-to-end, transferable deep reinforcement learning method for computational graph optimization (GO), based on a scalable sequential attention mechanism over an inductive graph neural network. GO generates decisions on the entire graph rather than on each individual node autoregressively, drastically speeding up the search compared to prior methods. Moreover, we propose recurrent attention layers to jointly optimize dependent graph optimization tasks and demonstrate 33%-60% speedup on three graph optimization tasks compared to TensorFlow default optimization. On a diverse set of representative graphs consisting of up to 80,000 nodes, including Inception-v3, Transformer-XL, and WaveNet, GO achieves on average 21% improvement over human experts and 18% improvement over the prior state of the art with 15x faster convergence, on a device placement task evaluated in real systems. Yanqi Zhou, Sudip Roy 0002, AmirAli Abdolrashidi, Daniel Wong 0001, Peter C. Ma, Qiumin Xu, Hanxiao Liu, Mangpo Phitchaya Phothilimtha, Anna Goldie, Azalia Mirhoseini, James Laudon |
NeurIPS | 6 |
| 2019 | GPUGuard: mitigating contention based side and covert channel attacks on GPUsabstractGraphics processing units (GPUs) are moving towards supporting concurrent kernel execution where multiple kernels may be co-executed on the same GPU and even on the same streaming multiprocessor (SM) core. While concurrent kernel execution improves hardware resource utilization, it opens up vulnerabilities to covert-channel and side-channel attacks. These attacks exploit information leakage across kernels that results from contention on shared resources; they have been shown to be a dangerous threat on CPUs, and are starting to be demonstrated on GPUs. The unique micro-architectural features of GPUs, such as specialized cache structures and massive parallel thread support, create opportunities for GPU-specific channels to be formed. In this paper, we propose GPUGuard, a decision tree based detection and a hierarchical defense framework that can reliably close the covert channels. Our results show that GPUGuard can detect contention with 100% sensitivity and a small (8.5%) false positive rate. The timing channels are mitigated through Tangram, a GPU-specific contention channel elimination scheme, with only 8% to 23% overhead when there is an attack and zero performance overhead when no attacks are detected. Compared to temporal partitioning, GPUGuard is 69%-96% faster in various architectures even when active, showing that it is possible to gain substantial performance from executing concurrent kernels on a single SM while securing GPUs against these attacks. Qiumin Xu, Hoda Naghibi Jouybari, Nael B. Abu-Ghazaleh, Murali Annavaram |
ICS | 1 |
| 2017 | Docker characterization on high performance SSDsabstractDocker containers are becoming the mainstay for deploying applications in cloud platforms, having many desirable features like ease of deployment, developer friendliness and lightweight virtualization. Meanwhile, solid state disks (SSDs) have witnessed tremendous performance boost through recent innovations in industry such as Non-Volatile Memory Express (NVMe) standards. However, the performance of containerized applications on these high speed contemporary SSDs has not yet been investigated. In this paper, we present a characterization of the performance impact among a wide variety of the available storage options for deploying Docker containers and provide the configuration options to best utilize the high performance SSDs. Qiumin Xu, Manu Awasthi, Krishna T. Malladi, Janki Bhimani, Jingpei Yang, Murali Annavaram |
ISPASS | 1 |
| 2017 | Rack Level Scheduling for Containerized WorkloadsabstractHigh performance SSDs have become ubiquitous in warehouse scale computing. Increased adoptions can be attributed to their high bandwidth, low latency and excellent random I/O performance. Owing to this high performance, multiple I/O intensive services can now be co-located on the same server. SSDs also introduce periodic latency spikes due to garbage collection. This, combined with multi-tenancy increases latency unpredictability since co-located applications now compete for CPU, memory, and disk bandwidth. The combination of these latency spikes and unpredictability lead to long tail latencies that can significantly decrease the system performance at scale. In this paper, we present a rack-level scheduling algorithm, which dynamically detects and shifts workloads with long tail latencies within servers in the same rack. Different from the global resource management methods, rack-level scheduling utilizes lightweight containers to minimize data movement and message passing overheads, leading to a much more efficient solution to reduce tail latency.With the algorithms implemented in the storage driver of the containerization infrastructure, it becomes viable to deploy and migrate applications in existing server racks without extensive modifications to storage, OS and other subsystems. Qiumin Xu, Krishna T. Malladi, Manu Awasthi |
NAS | 1 |
| 2016 | AirSense: an intelligent home-based sensing system for indoor air quality analyticsabstractIn the U.S., people spend approximately 90 percent of their time indoors. Unfortunately, indoor air quality (IAQ) may be two to five times worse than the air outdoors, and is often overlooked. Existing IAQ monitoring technologies focus on IAQ measurements and visualization. However, the lack of information about the pollution sources as well as the seriousness of the pollution makes people feel powerless and frustrated, resulting in the ignorance of the polluted air at their homes. In this work, we fill this critical gap by presenting AirSense, an intelligent home-based IAQ sensing system that is able to automatically detect pollution events, identify pollution sources, estimate personal exposure to indoor air pollution, and provide actionable suggestions to help people improve IAQ. We have deployed AirSense at five homes to evaluate its performance and investigate how users interact with it. We demonstrate that AirSense can accurately detect pollution events, identify pollution sources, and forecast IAQ information within five minutes in both controlled and real-world settings. We further show the great potential of AirSense in increasing users' awareness of IAQ and helping them better manage IAQ at their homes. Biyi Fang, Qiumin Xu, Taiwoo Park, Mi Zhang 0002 |
UbiComp | 2 |
| 2016 | Understanding performance of I/O intensive containerized applications for NVMe SSDsabstractOur cloud-based IT world is founded on hyper-visors and containers. Containers are becoming an important cornerstone, which is increasingly used day-by-day. Among different available frameworks, docker has become one of the major adoptees to use containerized platform in data centers and enterprise servers, due to its ease of deploying and scaling. Further more, the performance benefits of a lightweight container platform can be leveraged even more with a fast back-end storage like high performance SSDs. However, increase in number of simultaneously operating docker containers may not guarantee an aggregated performance improvement due to saturation. Thus, understanding performance bottleneck in a multi-tenancy docker environment is critically important to maintain application level fairness and perform better resource management. In this paper, we characterize the performance of persistent storage option (through data volume) for I/O intensive, dockerized applications. Our work investigates the impact on performance with increasing number of simultaneous docker containers in different workload environments. We provide, first of its kind study of I/O intensive containerized applications operating with NVMe SSDs. We show that 1) a six times better application throughput can be obtained, just by wise selection of number of containerized instances compared to single instance; and 2) for multiple application containers running simultaneously, an application throughput may degrade upto 50% compared to a stand-alone applications throughput, if good choice of application and workload is not made. We then propose novel design guidelines for an optimal and fair operation of both homogeneous and heterogeneous environments mixed with different applications and workloads. Janki Bhimani, Jingpei Yang, Zhengyu Yang 0001, Ningfang Mi, Qiumin Xu, Manu Awasthi, Rajinikanth Pandurangan, Vijay Balakrishnan |
IPCCC | 5 |
| 2016 | Warped-Slicer: Efficient Intra-SM Slicing through Dynamic Resource Partitioning for GPU MultiprogrammingabstractAs technology scales, GPUs are forecasted to incorporate an ever-increasing amount of computing resources to support thread-level parallelism. But even with the best effort, exposing massive thread-level parallelism from a single GPU kernel, particularly from general purpose applications, is going to be a difficult challenge. In some cases, even if there is sufficient thread-level parallelism in a kernel, there may not be enough available memory bandwidth to support such massive concurrent thread execution. Hence, GPU resources may be underutilized as more general purpose applications are ported to execute on GPUs. In this paper, we explore multiprogramming GPUs as a way to resolve the resource underutilization issue. There is a growing hardware support for multiprogramming on GPUs. Hyper-Q has been introduced in the Kepler architecture which enables multiple kernels to be invoked via tens of hardware queue streams. Spatial multitasking has been proposed to partition GPU resources across multiple kernels. But the partitioning is done at the coarse granularity of streaming multiprocessors (SMs) where each kernel is assigned to a subset of SMs. In this paper, we advocate for partitioning a single SM across multiple kernels, which we term as intra-SM slicing. We explore various intra-SM slicing strategies that slice resources within each SM to concurrently run multiple kernels on the SM. Our results show that there is not one intra-SM slicing strategy that derives the best performance for all application pairs. We propose Warped-Slicer, a dynamic intra-SM slicing strategy that uses an analytical method for calculating the SM resource partitioning across different kernels that maximizes performance. The model relies on a set of short online profile runs to determine how each kernel's performance varies as more thread blocks from each kernel are assigned to an SM. The model takes into account the interference effect of shared resource usage across multiple kernels. The model is also computationally efficient and can determine the resource partitioning quickly to enable dynamic decision making as new kernels enter the system. We demonstrate that the proposed Warped-Slicer approach improves performance by 23% over the baseline multiprogramming approach with minimal hardware overhead. Qiumin Xu, Hyeran Jeon, Keunsoo Kim, Won Woo Ro, Murali Annavaram |
ISCA | 1 |
| 2015 | Performance Characterization of Hyperscale Applicationson on NVMe SSDsabstractThe storage subsystem has undergone tremendous innovation in order to keep up with the ever-increasing demand for throughput. NVMe based SSDs are the latest development in this domain, delivering unprecedented performance in terms of both latency and peak bandwidth. Given their superior performance, NVMe drives are expected to be particularly beneficial for I/O intensive applications in datacenter installations. In this paper we identify and analyze the different factors leading to the better performance of NVMe SSDs. Then, using databases as the prominent use-case, we show how these would translate into real-world benefits. We evaluate both a relational database (MySQL) and a NoSQL database (Cassandra) and demonstrate significant performance gains over best-in-class enterprise SATA SSDs: from 3.5x for TPC-C and up to 8.5x for Cassandra. Qiumin Xu, Huzefa Siyamwala, Mrinmoy Ghosh, Manu Awasthi, Tameesh Suri, Zvika Guz, Anahita Shayesteh, Vijay Balakrishnan |
SIGMETRICS | 1 |
| 2015 | Performance analysis of NVMe SSDs and their implication on real world databasesabstractThe storage subsystem has undergone tremendous innovation in order to keep up with the ever-increasing demand for throughput. Non Volatile Memory Express (NVMe) based solid state devices are the latest development in this domain, delivering unprecedented performance in terms of latency and peak bandwidth. NVMe drives are expected to be particularly beneficial for I/O intensive applications, with databases being one of the prominent use-cases. Qiumin Xu, Huzefa Siyamwala, Mrinmoy Ghosh, Tameesh Suri, Manu Awasthi, Zvika Guz, Anahita Shayesteh, Vijay Balakrishnan |
SYSTOR | 1 |
| 2014 | PATS: pattern aware scheduling and power gating for GPGPUsabstractGeneral purpose computing using graphics processing units (GPGPUs) is an attractive option to achieve power efficient throughput computing. But the power efficiency of GPGPUs can be significantly curtailed in the presence of divergence. This paper evaluates two important facets of this problem. First, we study the branch divergence behavior of various GPGPU workloads. We show that only a few branch divergence patterns are dominant in most workloads. In fact only five branch divergence patterns account for 60% of all the divergent instructions in our workloads. In the second part of this work we exploit this branch divergence pattern bias to propose a new divergence pattern aware warp scheduler, called PATS. PATS prioritizes scheduling warps with the same divergence pattern so as to create long idleness windows for any given execution lane. The long idleness windows are then exploited for efficiently power gating the unused lanes while amortizing the gating overhead. We describe the architectural implementation details of PATS and evaluate the power and performance impact of PATS. Our proposed design significantly improves power gating efficiency of GPGPUs with minimal performance overhead. Qiumin Xu, Murali Annavaram |
PACT | 1 |