VLDB 2026 Research / reviewers in the wild / expert
Zhibin Yu 0001
dblp:70/8170-1
· DBLP profile ↗
51ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0001-8067-9612ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 44 · 8 first-author · 14 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Resource Efficiency of Performance-Optimized In-Memory Big Data Analytics With a Hybrid Static Dynamic Approach
Le-Le Li, Wei-Chung Hsu, Yeh-Ching Chung, Zhibin Yu 0001 |
IEEE Trans. Computers | 4 |
| 2025 | AMALI: An Analytical Model for Accurately Modeling LLM Inference on Modern GPUsabstractLarge language model (LLM) inference applications are surging in recent years, which largely relies on modern GPUs.On the other hand, GPU analytical model is a commonly used tool for architects to precisely identify bottlenecks quickly with deep insights.However, existing GPU analytical models fall short of accurately modeling LLM inference applications on modern GPUs, because of unsuitable tensor core modeling, ignoring constant cache as well as instruction cache modeling and abstracting away important details for LLM inference applications.To address this problem, we propose a novel analytical model dubbed AMALI to accurately model LLM inference on modern GPUs with three innovations.First, we develop an instruction modifier and throughput based tensor core model by accurately capturing the math pipe throttle stalls to enhance the architecture modeling for modern GPUs.Second, we propose analytical models for constant cache and instruction cache by developing micro-benchmarks to measure CUDA kernel launching latencies.This significantly improves AMALI's accuracy compared to real GPU hardware.Finally, we design a multi-warp model by leveraging warp instruction number distribution to reflect LLM inference application characteristics.We validate AMALI on an A100 GPU by using typical LLM inference applications.The results show that AMALI reduces the MAPE (mean absolute percentage error) from 127.56% to 23.59% Shiheng Cao, Junmin Wu, Junshi Chen 0003, Hong An, Zhibin Yu 0001 |
ISCA | 5 |
| 2025 | Swift: Fast Performance Tuning with GAN-Generated Configurations
Chao Chen 0022, Shixin Huang, Xuehai Qian, Zhibin Yu 0001 |
USENIX ATC | 4 |
| 2025 | Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance CountersabstractWe find existing benchmark suites for smartphone CPU micro-architecture design such as Geekbench 5.0 fail to authentically represent the micro-architecture-level performance behavior of widely used real Android applications with interactive operations such as screen sliding. It is therefore crucial to systematically construct a benchmark suite as a supplementary to Geekbench to represent the user interaction behavior of Android applications for CPU micro-architecture design. The key is to identify a small number of representative programs from a large number of real applications. To this end, a set of features used to represent a program need to be constructed, and these features should be fair for different micro-architectures and can be collected efficiently. However, this is extremely difficult for Android applications. For example, the feature collection tools for Android applications are unavailable for benchmark selection. In this article, we propose a novel benchmark suite construction approach dubbed BEMAP to efficiently build a supplementary benchmark suite from real-world Android applications to represent their user interaction behavior. 1 BEMAP innovates four techniques. The first technique, called two-stage RFC (representative feature construction), constructs program features from performance counters (events) to represent a program for selecting benchmarks from a large number of real Android applications in two stages. The first stage identifies a set of important performance events in terms of IPC (instructions per cycle) by employing a machine learning algorithm named SGBRT (Stochastic Gradient Boosted Regression Tree). The second stage constructs representative features based on the important performance events by using ICA (independent component analysis). The second technique, named SPC-MMA (source performance counters from multiple micro-architectures), collects the performance events from multiple mobile CPUs with different micro-architectures and mixes them as the source of RFC. The goal of these two innovations is to make the program features fair to different mobile CPU micro-architectures. The third technique, called ES (Elbow-Silhouette) approach, artfully leverages the synergy between the elbow method and the silhouette method to determine an optimal K when we use K-Means to group Android applications. The fourth technique is that we design a new tool named AutoProfiler to automatically profile the micro-architecture events (e.g., IPC, L1 Icache misses) of Android applications with interactive operations. Using the proposed BEMAP methodology, 2 we constructed SPBench, a novel benchmark suite supplementary to traditional mobile benchmark suites like Geekbench, for mobile CPU micro-architecture design. It consists of 15 benchmarks selected from 100 real Android applications with 3 common user interaction operations, which is the fifth innovation of this article. The experimental results on four significantly different micro-architectures show that SPBench can represent the micro-architecture performance behaviors of the 100 real-world applications with 3 common user interactive operations on each micro-architecture with significantly higher accuracy than benchmark suites produced by the state-of-the-art approaches. Chenghao Ouyang, Jinhan Xin, Siqi Zeng 0002, Guohui Li 0001, Jianjun Li 0010, Zhibin Yu 0001 |
ACM Trans. Archit. Code Optim. | 6 |
| 2024 | WASP: Workload-Aware Self-Replicating Page-Tables for NUMA ServersabstractRecently, page-table self-replication (PTSR) has been proposed to reduce the page-table caused NUMA effect for large-memory workloads on NUMA servers. However, PTSR may improve or hurt performance of an application, depending on its characteristics and the co-located applications. This is hard for users to know, but current PTSR can only be manually enabled/disabled by users. Hongliang Qu, Zhibin Yu 0001 |
ASPLOS (2) | 2 |
| 2024 | Guser: A GPGPU Power Stressmark GeneratorabstractPower stress mark is crucial for estimating Thermal Design Power (TDP) of GPGPUs to ensure efficient power control. This paper proposes Guser, the first systematic methodology to generate GPGPU power stressmarks. It features three ideas: Instruction Power Analysis (IPA) to analyze the power behavior of PTX instructions; Pipeline-based Instruction Grouping (PIG) to classify all PTX instructions into a number of groups; and Quantifying the Importance of power impact factors (QIF) to select a small number but important adjustable knobs. We adopt Optuna, an advanced BO algorithm, to generate stressmarks that maximize power consumption. We evaluate Guser on two real GPGPUs Tesla T4 (Turing) and Tesla A10 (Ampere). The experimental results show that the Guser-generated stressmark for T4 (T4-stresser) consumes 109.3 watts and the one for A10 (A10-stresser) consumes 238.7 watts, which are 48.7% and 73% higher than those consumed by the stressmarks generated by the state-of-the-art approach for CPUs, respectively. Moreover, T4-stresser and A10-stresser consume significantly more power than that consumed by any benchmarks in three benchmark suites: Cactus, Rodinia, and Parboil. Yalong Shan, Yongkui Yang, Xuehai Qian, Zhibin Yu 0001 |
HPCA | 4 |
| 2024 | Global-State Aware Automatic NUMA BalancingabstractNon-uniform memory access (NUMA) has become a standard architecture for modern servers. However, NUMA effect (i.e., local memory access typically takes shorter time than remote memory accesses) is unavoidable. To address this issue, Automatic NUMA Balancing(Auto-NUMA) was proposed. Nevertheless, Auto-NUMA can improve or hurt performance of an application, depending on its characteristics which is difficult for end users to know. Zhibin Yu 0001 |
Internetware | 2 |
| 2024 | TIE: Fast Experiment-Driven ML-Based Configuration Tuning for In-Memory Data AnalyticsabstractRecently, experiment-driven machine-learning (ML) based configuration tuning for in-memory data analytics such as Apache Spark become popular because they can achieve high speedups. However, experiment-driven ML-based approaches naturally need alargenumber of iterations and each iteration generates a configuration with a probabilistic strategy and executes the program on a real cluster with the configuration. It therefore takes a long time to optimize the performance of an in-memory data analytics program, and thereby hinders these approaches from being widely used in practice.To address this issue, we propose a novel as well as simple approach dubbedTerminating-It-Early (TIE)to reduce the time needed to perform the experiment executions but to achieve speedups similar to those obtained by experiment-driven ML-based approaches. The key idea is that, during the process of searching for the optimal configuration which produces the shortest execution time for a program, weterminatean experiment program execution with a trial configuration as soon as possible when we find its execution time islonger than a predefined threshold(e.g., the shortest execution time thus far). In contrast, traditional experiment-driven ML-based approaches always run all experiment executions completely.We employ 19 Apache Spark programs running on a physical cluster as well as a virtual cluster to evaluate TIE. We compare thetuning timeused to find the optimal configuration of a program and theoptimized execution timeof a program obtained by TIE against those obtained byCherryPickand a reinforcement learning (RL) based approach. The experimental results show that on physical machines, TIE reduces the tuning time used byCherryPickand the RL-based approach by factors of 2.39× and 1.68× on average, respectively. On virtual machines, the corresponding factors are 2.79× and 1.71×. Moreover, the average optimized execution time of the 19 programs tuned by TIE is slightly shorter than those tuned byCherryPickand the RL-based approach. Chao Chen 0022, Jinhan Xin, Zhibin Yu 0001 |
IEEE Trans. Computers | 3 |
| 2024 | Falic: An FPGA-Based Multi-Scalar Multiplication Accelerator for Zero-Knowledge ProofabstractIn this paper, we propose Falic, a novel FPGA-based accelerator to accelerate multi-scalar multiplication (MSM), the most time-consuming phase of zk-SNARK proof generation. Falic innovates three techniques. First, it leverages globally asynchronous locally synchronous (GALS) strategy to build multiple small and lightweight MSM cores to parallelize the independent inner product computation on different portions of the scalar vector and point vector. Second, each MSM core contains just one large-integer modular multiplier (LIMM) that is multiplexed to perform the point additions (PADDs) generated during MSM. We strike a balance between the throughput and hardware cost by batching the appropriate number of PADDs and selecting the computation graph of PADD with proper parallelism degree. Finally, the performance is further improved by a simple cache structure that enables the computation reuse. We implement Falic on two different FPGAs with different hardware resources, i.e., the Xilinx U200 and Xilinx U250. Compared to the prior FPGA-based accelerator, Falic improves the MSM throughput by$3.9\boldsymbol{\times}$. Experimental results also show that Falic achieves a throughput speedup of up to$1.62\boldsymbol{\times}$and saves as much as$8.5\boldsymbol{\times}$energy compared to an RTX 2080Ti GPU. Yongkui Yang, Zhenyan Lu, Jingwei Zeng, Xingguo Liu, Xuehai Qian, Zhibin Yu 0001 |
IEEE Trans. Computers | 6 |
| 2024 | Satisfying Energy-Efficiency Constraints for Mobile SystemsabstractEnergy-efficiency is one of the most important design criteria for mobile systems, such as smartphones and tablets. But current mobile systems always over-provision resources to satisfy users. The root cause is that, we have no knowledge on how much of system performance/energy will exactly satisfy users. Psychophysics defines the quantified link between physical stimuli and human-perceived stimuli. So, we will leverage psychophysics to study the quantified correlation between computer architecture resources (i.e., physical stimuli) and user satisfaction (i.e., human-perceived stimuli). We then exploit such correlation to precisely apportion resources to operate tasks and accurately satisfy users. Benefiting from our precisely-defined user satisfaction criteria and well-designed algorithms, we can reduce energy consumption of computer architectures by up to 42.9% without harming user experience. To the best of our knowledge, we for the first time theoretically and accurately model such substantial correlation. Our work opens a new research domain for fundamentally improving mobiles’ energy-efficiency. Xueliang Li 0002, Shicong Hong, Junyang Chen 0001, Junkai Ji, Chengwen Luo 0001, Guihai Yan, Zhibin Yu 0001, Jianqiang Li 0001 |
IEEE Trans. Mob. Comput. | 7 |
| 2023 | PAC: Preference-Aware Co-location Scheduling on Heterogeneous NUMA Architectures To Improve Resource UtilizationabstractLatency-critical applications directly interact with end users and often experience the diurnal load pattern. In production, best-effort applications are often co-located with them to utilize the idle cores at the low load. Meanwhile, modern computers are evolving towards heterogeneous NUMA architecture, where the cores have different computation abilities, memory access latencies and network communication delays. Prior co-location scheduling work did not consider the NUMA architecture, and failed to maximize the throughput of best-effort applications while ensuring the required QoS of latency-critical applications. Our investigation shows that NUMA effect has complex impacts on the latency of latency-critical applications and the throughput of best-effort applications. We therefore propose PAC, a preference-aware co-location scheduling scheme that considers the NUMA effect for heterogeneous NUMA architectures. PAC has a performance monitor and a core scheduler. Specifically, the performance monitor identifies the "dangerous" latency-critical applications that require upgrading core allocations. We propose two low-overhead scheduling strategies for the scheduler. The strategies identify the bottlenecks of applications and adjust core allocations accordingly. Experimental result shows that PAC improves the throughput of best-effort applications by 3.87× while ensuring the required QoS of latency-critical applications. Pu Pang, Yaoxuan Li, Bo Liu 0122, Quan Chen 0002, Zhou Yu 0003, Zhibin Yu 0001, Deze Zeng, Jingwen Leng, Jieru Zhao, Minyi Guo |
ICS | 6 |
| 2023 | Resource scheduling techniques in cloud from a view of coordination: a holistic surveyabstractNowadays, the management of resource contention in shared cloud remains a pending problem. The evolution and deployment of new application paradigms (e.g., deep learning training and microservices) and custom hardware (e.g., graphics processing unit (GPU) and tensor processing unit (TPU)) have posed new challenges in resource management system design. Current solutions tend to trade cluster efficiency for guaranteed application performance, e.g., resource over-allocation, leaving a lot of resources underutilized. Overcoming this dilemma is not easy, because different components across the software stack are involved. Nevertheless, massive efforts have been devoted to seeking effective performance isolation and highly efficient resource scheduling. The goal of this paper is to systematically cover related aspects to deliver the techniques from the coordination perspective, and to identify the corresponding trends they indicate. Briefly, four topics are involved. First, isolation mechanisms deployed at different levels (micro-architecture, system, and virtualization levels) are reviewed, including GPU multitasking methods. Second, resource scheduling techniques within an individual machine and at the cluster level are investigated, respectively. Particularly, GPU scheduling for deep learning applications is described in detail. Third, adaptive resource management including the latest microservice-related research is thoroughly explored. Finally, future research directions are discussed in the light of advanced work. We hope that this review paper will help researchers establish a global view of the landscape of resource management techniques in shared cloud, and see technology trends more clearly. Yuzhao Wang, Junqing Yu, Zhibin Yu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2022 | LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL ApplicationsabstractSpark SQL has been widely deployed in industry but it is challenging to tune its performance. Recent studies try to employ machine learning (ML) to solve this problem, but suffer from two drawbacks. First, it takes a long time (high overhead) to collect training samples. Second, the optimal configuration for one input data size of the same application might not be optimal for others. Jinhan Xin, Kai Hwang 0001, Zhibin Yu 0001 |
SIGMOD Conference | 3 |
| 2022 | SOCA-DOM: A Mobile System-on-Chip Array System for Analyzing Big Data on the Move
Le-Le Li, Jiang-Yi Liu, Jianping Fan 0002, Xuehai Qian, Kai Hwang 0001, Yeh-Ching Chung, Zhibin Yu 0001 |
J. Comput. Sci. Technol. | 7 |
| 2022 | OSC: An Online Self-Configuring Big Data Framework for Optimization of QoSabstractBig-data frameworks such as MapReduce/Hadoop or Spark have many performance-critical configuration parameters which may interact with each other in a complex way. Their optimal values for an application on a given cluster are affected by not only the application itself but also its input data. This makes offline auto-configuration approaches hard to be used in practice because the input data of an application may change at each run. To address this issue, we propose an Online Self-Configuring (OSC) approach that automatically determines the optimal parameter values for a given application. OSC synergistically integrates three key techniques. First, OSC leveragesensemble learningto build a precise performance model for a given application. Second, it quantifies theimportanceof the parameters andinteraction intensitybetween them to accelerate the genetic algorithm for searching optimal configuration parameters. Third, OSC supports anincremental modelingapproach to achieve low overhead of the models for online needs. These techniques allow OSC to effectively learn the characteristics of an application and optimize its performance by automatically adjusting the configurations at runtime. Our implementation of OSC atop MapReduce/Hadoop 2.6 improves performance by 60 percent on average and up to 120 percent compared with the state-of-the-art approach. Lastly, the performance benefit of an application running on OSC generally increases along with its input data size. Zhendong Bei, Nam Sung Kim, Kai Hwang 0001, Zhibin Yu 0001 |
IEEE Trans. Computers | 4 |
| 2021 | Para: Harvesting CPU time fragments in Big Data AnalyticsabstractModern data analytics typically run tasks on statically reserved resources (e.g., CPU and memory), which is prone to over-provision to guarantee the Quality of Service (QoS), leading to a large amount of resource time fragments. As a result, the resource utilization of a data analytics cluster is severely under-utilized. Workload co-location on shared resources has been substantially studied, but they are unaware the sizes of resource time fragments, making them hard to improve the resource utilization and guarantee QoS at the same time. In this paper, we propose Para, an event-driven scheduling mechanism, to harvest the CPU time fragments in co-located big data analytic workloads. Para innovates three techniques: 1) identifying the Idle CPU Time Window (ICTW) associated with each CPU core by capturing the task-switch event; 2) designing a runtime communication mechanism between each task execution of a workload and the underlying resource management system; 3) designing a pull-based scheduler to schedule a workload to run in the ICTW of another workload. We implement Para based on Apache Mesos and Spark. And the experimental results show that Para improves the CPU utilization by 44% and 30% on average relative to the original Mesos and enhanced Mesos under Spark's dynamic mode (MSDM), respectively. Moreover, Para increases the averaged task throughput of Mesos and MSDM by 4.8x and 1.7x, respectively, while guaranteeing the execution time of the primary applications. Yuzhao Wang, Hongliang Qu, Junqing Yu, Zhibin Yu 0001 |
CLOUD | 4 |
| 2021 | OR-ML: Enhancing Reliability for Machine Learning Accelerator with Opportunistic RedundancyabstractReliability plays a central role in deep sub-micron and nanometre IC fabrication technology and has recently been reported to be one of the key issues affecting the inference phase of neural networks. State-of-the-art machine learning (ML) accelerators exploit massively computing parallelism observed in neural networks to achieve high energy efficiency. The topology of ML engines' computing fabric, which constitutes large arrays of processing elements (PEs), has been increasing dramatically to incorporate the huge size and heterogeneity of the rapid evolving ML algorithm. However, it is commonly observed that activations of zero value lead to reduced PE utilization. In this work, we present a novel and low-cost approach to enhance the reliability of generic ML accelerators by Qpportunistically exploring the chances of runtime Redundancy provided by neighbouring PEs, named as OR-ML. In contrast to conventional redundancy techniques, the proposed technique introduces no additional computing resources, therefore significantly reduces the implementation overhead and achieves obvious level of protection. The design prototype is evaluated using emulated fault injection on FPGA, executing mainstream neural networks for objectionclassification and detection. Zheng Wang 0027, Wenxuan Chen, Chao Chen 0022, Yongkui Yang, Zhibin Yu 0001 |
DATE | 6 |
| 2021 | CNN-DMA: A Predictable and Scalable Direct Memory Access Engine for Convolutional Neural Network with Sliding-window FilteringabstractMemory bandwidth utilization has become the key performance bottleneck for state-of-the-art variants of neural network kernels. Current structures such as depth-wise, point-wise and atrous convolutions have already introduced diverse and discontinuous memory access patterns, which impact efficient activation supply due to more frequent cache misses and consequently high-penalty DRAM pre-charging. To handle this, GPU achieves efficient parallelization with sophisticated optimization of CUDA program to reduce memory footprints, which demands high engineering efforts. In this work, we in contrast propose a programmable direct memory access engine for convolutional neural networks (CNN-DMA) supporting a fast supply of activation for independent and scalable computing units. The CNN-DMA favours a predictable activation streaming approach which completely avoids penalties by bus contention, cache misses and less carefully designed low-level programs. Furthermore, we enhance the baseline DMA with the capability of out-of-order data supply to filter out unique sliding-windows to boost the performance of the computing infrastructure. Experiments on state-of-the-art neural networks show that CNN-DMA achieves optimal DRAM access efficiency for point-wise convolution layers, while reduces 30% to 70% rounds of computation with sliding-window filtering. Zheng Wang 0027, Chao Chen 0022, Yongkui Yang, Weiguang Chen, Wenxuan Chen, Weiyu Guo, Zhibin Yu 0001 |
ACM Great Lakes Symposium on VLSI | 13 |
| 2021 | Democratic learning: hardware/software co-design for lightweight blockchain-secured on-device machine learningabstractRecently, the trending 5G technology encourages extensive applications of on-device machine learning , which collects user data for model training. This requires cost-effective techniques to preserve the privacy and the security of model training within the resource-constrained environment. Traditional learning methods rely on the trust among the system for privacy and security. However, with the increase of the learning scale, maintaining every edge device’s trustworthiness could be expensive. To cost-effectively establish trust in a trustless environment, this paper proposes democratic learning (DemL), which makes the first step to explore hardware/software co-design for blockchain-secured decentralized on-device learning. By utilizing blockchain’s decentralization and tamper-proofing, our design secures AI learning in a trustless environment. To tackle the extra overhead introduced by blockchain , we propose PoMC (an algorithm and architecture co-design) as a novel blockchain consensus mechanism , which first exploits cross-domain reuse (AI learning and blockchain consensus) in AI learning architecture. Evaluation results show our DemL can protect AI learning from privacy leakage and model pollution, and demonstrated that privacy and security come with trivial hardware overhead and power consumption (2%). We believe that our work will open the door of synergizing blockchain and on-device learning for security and privacy. Mingcong Song, Tao Li 0006, Zhibin Yu 0001, Yuting Dai, Xiaoguang Liu 0001, Gang Wang 0001 |
J. Syst. Archit. | 4 |
| 2021 | GML: Efficiently Auto-Tuning Flink's Configurations Via Guided Machine LearningabstractThe increasingly popular fused batch-streaming big data framework, Apache Flink, has many performance-critical as well as untamed configuration parameters. However, how to tune them for optimal performance has not yet been explored. Machine learning (ML) has been chosen to tune the configurations for other big data frameworks (e.g., Apache Spark), showing significant performance improvements. However, it needs a long time to collect a large amount of training data by nature. In this article, we propose a guided machine learning (GML) approach to tune the configurations of Flink with significantly shorter time for collecting training data compared to traditional ML approaches. GML innovates two techniques. First, it leverages generative adversarial networks (GANs) to generate a part of training data, reducing the time needed for training data collection. Second, GML guides a ML algorithm to select configurations that the corresponding performance is higher than the average performance of random configurations. We evaluate GML on a lab cluster with 4 servers and a real production cluster in an internet company. The results show that GML significantly outperforms the state-of-the-art, DAC (Datasize-Aware-Configuration) (Z. Yu et al. 2018) for tuning the configurations of Spark, with 2.4× of reduced data collection time but with 30 percent reduced 99th percentile latency. When GML is used in the internet company, it reduces the latency by up to 57.8× compared to the configurations made by the company. Yijin Guo, Huasong Shan, Shixin Huang, Kai Hwang 0001, Jianping Fan 0002, Zhibin Yu 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2020 | BBS: Micro-Architecture Benchmarking Blockchain Systems through Machine Learning and Fuzzy SetabstractDue to the decentralization, irreversibility, and traceability, blockchain has attracted significant attention and has been deployed in many critical industries such as banking and logistics. However, the micro-architecture characteristics of blockchain programs still remain unclear. What's worse, the large number of micro-architecture events make understanding the characteristics extremely difficult. We even lack a systematic approach to identify the important events to focus on. In this paper, we propose a novel benchmarking methodology dubbed BBS to characterize blockchain programs at micro-architecture level. The key is to leverage fuzzy set theory to identify important micro-architecture events after the significance of them is quantified by a machine learning based approach. The important events for single programs are employed to characterize the programs while the common important events for multiple programs form an importance vector which is used to measure the similarity between benchmarks. We leverage BBS to characterize seven and six benchmarks from Blockbench and Caliper, respectively. The results show that BBS can reveal interesting findings. Moreover, by leveraging the importance characterization results, we improve that the transaction throughput of Smallbank from Fabric by 70% while reduce the transaction latency by 55%. In addition, we find that three of seven and two of six benchmarks from Blockbench and Caliper are redundant, respectively. Chao Chen 0022, Zihao Su, Weiguang Chen, Tao Li 0006, Zhibin Yu 0001 |
HPCA | 6 |
| 2020 | Thread-Level Locking for SIMT ArchitecturesabstractAs more emerging applications are moving to GPUs, thread-level synchronization has become a requirement. However, GPUs only provide warp-level and thread-block-level rather than thread-level synchronization. Moreover, it is highly possible to cause live-locks by using CPU synchronization mechanisms to implement thread-level synchronization for GPUs. In this article, we first propose a software-based thread-level synchronization mechanism called lock stealing for GPUs to avoid live-locks. We then describe how to implement our lock stealing algorithm in mutual exclusive locks and readers-writer locks with high performance. Finally, by putting it all together, we develop a thread-level locking library (TLLL) for commercial GPUs. To evaluate TLLL and show its general applicability, we use it to implement six widely used programs. We compare TLLL against the state-of-the-art ad-hoc GPU synchronization, GPU software transactional memory (STM), and CPU hardware transactional memory (HTM), respectively. The results show that, compared with the ad-hoc GPU synchronization for Delaunay mesh refinement (DMR), TLLL improves the performance by 22 percent on average on a GTX970 GPU, and shows up to 11 percent of performance improvement on a Volta V100 GPU. Moreover, it significantly reduces the required memory size. Such low memory consumption enables DMR to successfully run on the GTX970 GPU with the 10-million mesh size, and the V100 GPU with the 40-million mesh size, with which the ad-hoc synchronization can not run successfully. In addition, TLLL outperforms the GPU STM by 65 percent, and the CPU HTM (running on a Xeon E5-2620 v4 CPU with 16 hardware threads) by 43 percent on average. Lan Gao 0004, Rui Wang 0014, Zhongzhi Luan, Zhibin Yu 0001, Depei Qian 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | COPA: Highly Cost-Effective Power Back-Up for Green DatacentersabstractTraditional datacenters employ costly diesel generators (DG) and uninterrupted power supplies (UPS) to back up power. However, some or even all racks of a green datacenter can still be powered by renewable energy during grid power outages. This makes the utilization of the DGs and UPSs in green datacenters significantly lower than in traditional datacenters. In this paper, we propose a highly cost-effective power back-up (COPA) approach for green datacenters by leveraging the availability characteristics of renewable energy as well as grid power outages. COPA contributes three new techniques. The first technique, called least UPS capacity planning, determines the least rated power capability and runtime of the UPSs to guarantee the normal operations of a green datacenter during grid power outages. The second technique, named cooperative UPS/renewable power supply, employs UPS and renewable energy at the same time to supply power to each rack when grid power fails. The last one, dubbed renewable-energy-aware dynamic power management, controls the power consumption dynamically based on the available capacity of renewable energy and UPS. We build an experimental cluster consisting of 10 servers, and use four representative benchmarks as well as verified data about the availability characteristics of solar and wind energy to evaluate COPA. The results show that COPA reduces 47 percent and 70 percent of the power back-up cost for a solar energy powered datacenter and a wind energy powered datacenter, respectively. Moreover, COPA guarantees the application's Service Level Agreement (SLA) for at least 20 minutes (over 79 percent outages) and 56 minutes on average while enabling the back-up power to last for at least 2 hours and for 3 hours on average, which cannot be achieved by other under-provisioning power back-up approaches. Junmin Wu, Lieven Eeckhout, Amer Qouneh, Tao Li 0006, Zhibin Yu 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2019 | Adaptive memory-side last-level GPU cachingabstractEmerging GPU applications exhibit increasingly high computation demands which has led GPU manufacturers to build GPUs with an increasingly large number of streaming multiprocessors (SMs). Providing data to the SMs at high bandwidth puts significant pressure on the memory hierarchy and the Network-on-Chip (NoC). Current GPUs typically partition the memory-side last-level cache (LLC) in equally-sized slices that are shared by all SMs. Although a shared LLC typically results in a lower miss rate, we find that for workloads with high degrees of data sharing across SMs, a private LLC leads to a significant performance advantage because of increased bandwidth to replicated cache lines across different LLC slices. Xia Zhao 0004, Almutaz Adileh, Zhibin Yu 0001, Zhiying Wang 0003, Aamer Jaleel, Lieven Eeckhout |
ISCA | 3 |
| 2019 | Green-Up: under-provisioning power backup infrastructure for green datacenters
Zhibin Yu 0001 |
CCF Trans. High Perform. Comput. | 2 |
| 2019 | MiC: Multi-level Characterization and Optimization of GPGPU KernelsabstractGraphics processing units (GPUs) 1 have enjoyed increasing popularity in recent years, which benefits from, for example, general-purpose GPU (GPGPU) for parallel programs and new computing paradigms, such as the Internet of Things (IoT). GPUs hold great potential in providing effective solutions for big data analytics while the demands for processing large quantities of data in real time are also increasing. However, the pervasive presence of GPUs on mobile devices presents great challenges for GPGPU, mainly because GPGPU integrates a large amount of processor arrays and concurrent executing threads (up to hundreds of thousands). In particular, the root causes of performance loss in a GPGPU program can not be revealed in detail by current approaches. In this article, we propose MiC (Multi-level Characterization), a framework that comprehensively characterizes GPGPU kernels at the instruction, Basic Block (BBL), and thread levels. Specifically, we devise Instruction Vectors (IV) and Basic Blocks Vectors (BBV), a Thread Similarity Matrix (TSM), and a Divergence Flow Statistics Graph (DFSG) to profile information in each level. We use MiC to provide insights into GPGPU kernels through the characterizations of 34 kernels from popular GPGPU benchmark suites such as Compute Unified Device Architecture (CUDA) Software Development Kit (SDK), Rodinia, and Parboil. In comparison with Central Processing Unit (CPU) workloads, we conclude the key findings as follows: (1) There are comparable Instruction-Level Parallelism (ILP); (2) The BBL count is significantly smaller than CPU workloads—only 22.8 on average; (3) The dynamic instruction count per thread varies from dozens to tens of thousands and it is extremely small compared to CPU benchmarks; (4) The Pareto principle (also called 90/10 rule) does not apply to GPGPU kernels while it pervasively exists in CPU programs; (5) The loop patterns are dramatically different from those in CPU workloads; (6) The branch ratio is lower than that of CPU programs but higher than pure GPU workloads. In addition, we have also shown how TSM and DFSG are used to characterize the branch divergence in a visual way, to enable the analysis of thread behavior in GPGPU programs. In addition, we show an optimization case for a GPGPU kernel from the bottleneck identified through its characterization result, which improves 16.8% performance. Qixiao Liu, Zhibin Yu 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2018 | Datasize-Aware High Dimensional Configurations Auto-Tuning of In-Memory Cluster ComputingabstractIn-Memory cluster Computing (IMC) frameworks (e.g., Spark) have become increasingly important because they typically achieve more than 10× speedups over the traditional On-Disk cluster Computing (ODC) frameworks for iterative and interactive applications. Like ODC, IMC frameworks typically run the same given programs repeatedly on a given cluster with similar input dataset size each time. It is challenging to build performance model for IMC program because: 1) the performance of IMC programs is more sensitive to the size of input dataset, which is known to be difficult to be incorporated into a performance model due to its complex effects on performance; 2) the number of performance-critical configuration parameters in IMC is much larger than ODC (more than 40 vs. around 10), the high dimensionality requires more sophisticated models to achieve high accuracy. To address this challenge, we propose DAC, a datasize-aware auto-tuning approach to efficiently identify the high dimensional configuration for a given IMC program to achieve optimal performance on a given cluster. DAC is a significant advance over the state-of-the-art because it can take the size of input dataset and 41 configuration parameters as the parameters of the performance model for a given IMC program, --- unprecedented in previous work. It is made possible by two key techniques: 1) Hierarchical Modeling (HM), which combines a number of individual sub-models in a hierarchical manner; 2) Genetic Algorithm (GA) is employed to search the optimal configuration. To evaluate DAC, we use six typical Spark programs, each with five different input dataset sizes. The evaluation results show that DAC improves the performance of six typical Spark programs, each with five different input dataset sizes compared to default configurations by a factor of 30.4x on average and up to 89x. We also report that the geometric mean speedups of DAC over configurations by default, expert, and RFHOC are 15.4x, 2.3x, and 1.5x, respectively. Zhibin Yu 0001, Zhendong Bei, Xuehai Qian |
ASPLOS | 1 |
| 2018 | The Elasticity and Plasticity in Semi-Containerized Co-locating Cloud Workload: a View from Alibaba TraceabstractCloud computing with large-scale datacenters provides great convenience and cost-efficiency for end users. However, the resource utilization of cloud datacenters is very low, which wastes a huge amount of infrastructure investment and energy to operate. To improve resource utilization, cloud providers usually co-locate workloads of different types on shared resources. However, resource sharing makes the quality of service (QoS) unguaranteed. In fact, improving resource utilization (IRU) and guaranteeing QoS at the same time in cloud has been a dilemma which we name an IRU-QoS curse. To tackle this issue, characterizing the workloads from real production cloud computing platforms is extremely important. Qixiao Liu, Zhibin Yu 0001 |
SoCC | 2 |
| 2018 | CounterMiner: Mining Big Performance Data from Hardware CountersabstractModern processors typically provide a small number of hardware performance counters to capture a large number of microarchitecture events. These counters can easily generate a huge amount (e.g., GB or TB per day) of data, which we call big performance data in cloud computing platforms with more than thousands of servers and millions of complex workloads running in a "24/7/365" manner. The big performance data provides a precious foundation for root cause analysis of performance bottlenecks, architecture and compiler optimization, and many more. However, it is challenging to extract value from the big performance data due to: 1) the many unperceivable errors (e.g., outliers and missing values); and 2) the difficulty of obtaining insights, e.g., relating events to performance. In this paper, we propose CounterMiner, a rigorous methodology that enables the measurement and understanding of big performance data by using data mining and machine learning techniques. It includes three novel components: 1) using data cleaning to improve data quality by replacing outliers and filling in missing values; 2) iteratively quantifying, ranking, and pruning events based on their importance with respect to performance; 3) quantifying interaction intensity between two events by residual variance. We use sixteen benchmarks (eight from CloudSuite and eight from the Spark version of HiBench) to evaluate CounterMiner. The experimental results show that CounterMiner reduces the average error from 28.3% to 7.7% when multiplexing 10 events on 4 hardware counters. We also conduct a real-world case study, showing that identifying important configuration parameters of Spark programs by event importance is much faster than directly ranking the importance of these parameters. Yirong Lv, Qingyi Luo, Jing Wang 0055, Zhibin Yu 0001, Xuehai Qian |
MICRO | 5 |
| 2018 | Configuring in-memory cluster computing using random forest
Zhendong Bei, Zhibin Yu 0001, Ni Luo, Chuntao Jiang, Cheng-Zhong Xu 0001, Shengzhong Feng |
Future Gener. Comput. Syst. | 2 |
| 2018 | QIG: Quantifying the Importance and Interaction of GPGPU Architecture ParametersabstractGraphic processing units (GPUs) are widely used for general-purpose computing-so-called GPGPU computing. GPUs feature a large number of architecture parameters, resulting in a huge design space. To quickly explore this design space and identify the optimum architecture for a group of widely used computing kernels, it is critical to know how important each parameter is and how strongly these parameters interact with each other. This paper proposes an ensemble-learning-based approach, called quantifying the importance and interaction of Gpgpu architecture parameters (QIG), to quantify the importance of architecture parameters and their interactions with respect to performance. QIG employs a stochastic gradient boosted regression tree to construct performance models using performance data from a random set of GPU architectures. Leveraging these models, QIG observes the impact of each architecture parameter on performance, and calculates its importance and interaction intensity with other parameters. Using 25 widely used GPGPU kernels, we demonstrate that QIG accurately ranks the importance and interaction of GPU architecture parameters while the previously proposed Plackett-Burman design does not. Moreover, we show that QIG leads to a substantially more accurate performance model compared to prior work, including Starchart and approaches using artificial neural networks and supported vector machines: average error of 4.2% for QIG versus 23+% for prior work. Finally, QIG reveals a number of interesting insights for GPU architectures running GPGPU workloads. Zhibin Yu 0001, Jing Wang 0055, Lieven Eeckhout, Cheng-Zhong Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | MIA: Metric Importance Analysis for Big Data Workload CharacterizationabstractData analytics is at the foundation of both high-quality products and services in modern economies and societies. Big data workloads run on complex large-scale computing clusters, which implies significant challenges for deeply understanding and characterizing overall system performance. In general, performance is affected by many factors at multiple layers in the system stack, hence it is challenging to identify the key metrics when understanding big data workload performance. In this paper, we propose a novel workload characterization methodology using ensemble learning, called Metric Importance Analysis (MIA), to quantify the respective importance of workload metrics. By focusing on the most important metrics, MIA reduces the complexity of the analysis without losing information. Moreover, we develop the MIA-based Kiviat Plot (MKP) and Benchmark Similarity Matrix (BSM) which provide more insightful information than the traditional linkage clustering based dendrogram to visualize program behavior (dis)similarity. To demonstrate the applicability of MIA, we use it to characterize three big data benchmark suites: HiBench, CloudRank-D and SZTS. The results show that MIA is able to characterize complex big data workloads in a simple, intuitive manner, and reveal interesting insights. Moreover, through a case study, we demonstrate that tuning the configuration parameters related to the important metrics found by MIA results in higher performance improvements than through tuning the parameters related to the less important ones. Zhibin Yu 0001, Lieven Eeckhout, Zhendong Bei, Avi Mendelson, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | POSTER: BACM: Barrier-Aware Cache Management for Irregular Memory-Intensive GPGPU WorkloadsabstractGeneral-purpose workloads running on modern graphics processing units (GPGPUs) rely on hardware-based barriers to synchronize warps within a thread block (TB). However, imbalance may exist before reaching a barrier if a GPGPU workload contains irregular memory accesses, i.e., some warps may be critical while others may not. Ideally, cache space should be reserved for the critical warps. Unfortunately, current cache management policies are unaware of the existence of barriers and critical warps, which significantly limits the performance of irregular memory-intensive GPGPU workloads.In this work, we propose Barrier-Aware Cache Management (BACM), which is built on top of two underlying policies: a greedy policy and a friendly policy. The greedy policy does not allow non-critical warps to allocate cache lines in the L1 data cache; only critical warps can. The friendly policy allows non-critical warps to allocate cache lines but only over invalid or lower-priority cache lines. Based on the L1 data cache hit rate of non-critical warps, BACM dynamically chooses between the greedy and friendly policies. By doing so, BACM reserves more cache space to accelerate critical warps, thereby improving overall performance. Experimental results show that BACM achieves an average performance improvement of 24% and 20% compared to the GTO and BAWS policies, respectively. BACM's hardware cost is limited to 96 bytes per streaming multiprocessor. Xia Zhao 0004, Zhibin Yu 0001, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Lieven Eeckhout |
PACT | 3 |
| 2017 | BACM: Barrier-Aware Cache Management for Irregular Memory-Intensive GPGPU WorkloadsabstractGeneral-purpose workloads running on modern graphics processing units rely on hardware-based barriers to synchronize warps within a thread block (TB). However, imbalance may exist before reaching a barrier if a GPGPU workload contains irregular memory accesses, i.e., some warps may be critical while others may not. Ideally, cache space should be reserved for the critical warps. Unfortunately, current cache management policies are unaware of the existence of barriers and critical warps, which significantly limits the performance of irregular memory-intensive GPGPU workloads. In this paper, we propose Barrier-Aware Cache Management (BACM) which is built on top of two underlying policies: a greedy policy and a friendly policy. The greedy policy does not allow non-critical warps to allocate cache lines in the L1 data cache; only critical warps can. The friendly policy allows non-critical warps to allocate cache lines but only over invalid or lower-priority cache lines. BACM dynamically chooses between the greedy and friendly policies based on the L1 data cache hit rate for the non-critical warps. By doing so, BACM reserves more cache space to accelerate critical warps, thereby improving overall performance. Experimental results show that BACM achieves an average performance improvement of 24% and 20% compared to the GTO and BAWS policies, respectively. BACM's hardware cost is limited to 96 bytes per streaming multiprocessor. Xia Zhao 0004, Zhibin Yu 0001, Zhenlin Wang 0003, Xiaolin Wang 0001, Yingwei Luo, Lieven Eeckhout |
ICCD | 3 |
| 2016 | Thread Similarity Matrix: Visualizing Branch Divergence in GPGPU ProgramsabstractGraphics processing units (GPUs) have recently evolved into popular accelerators for general-purpose parallel programs -- so-called GPGPU computing. Although programming models such as CUDA and OpenCL significantly improve GPGPU programmability, optimizing GPGPU programs is still far from trivial. Branch divergence is one of the root causes reducing GPGPU performance. Existing approaches are able to calculate the branch divergence rate but are unable to reveal how the branches diverge in a GPGPU program. In this paper, we propose the Thread Similarity Matrix (TSM) to visualize how branches diverge and in turn help find optimization opportunities. TSM contains an element for each pair of threads, representing the difference in code being executed by the pair of threads. The darker the element, the more similar the threads are, the lighter, the more dissimilar. TSM therefore allows GPGPU programmers to easily understand an application's branch divergence behavior and pinpoint performance anomalies. We present a case study to demonstrate how TSM can help optimize GPGPU programs: we improve the performance of a highly-optimized GPGPU kernel by 35% by reorganizing its thread organization to reduce its branch divergence rate. Zhibin Yu 0001, Lieven Eeckhout, Cheng-Zhong Xu 0001 |
ICPP | 1 |
| 2016 | Barrier-Aware Warp Scheduling for Throughput ProcessorsabstractParallel GPGPU applications rely on barrier synchronization to align thread block activity. Few prior work has studied and characterized barrier synchronization within a thread block and its impact on performance. In this paper, we find that barriers cause substantial stall cycles in barrier-intensive GPGPU applications although GPGPUs employ lightweight hardware-support barriers. To help investigate the reasons, we define the execution between two adjacent barriers of a thread block as a warp-phase. We find that the execution progress within a warp-phase varies dramatically across warps, which we call warp-phase-divergence. While warp-phase-divergence may result from execution time disparity among warps due to differences in application code or input, and/or shared resource contention, we also pinpoint that warp-phase-divergence may result from warp scheduling. Zhibin Yu 0001, Lieven Eeckhout, Vijay Janapa Reddi, Yingwei Luo, Xiaolin Wang 0001, Zhenlin Wang 0003, Cheng-Zhong Xu 0001 |
ICS | 2 |
| 2016 | QIM: Quantifying Hyperparameter Importance for Deep Learning
Dan Jia, Rui Wang 0014, Cheng-Zhong Xu 0001, Zhibin Yu 0001 |
NPC | 4 |
| 2016 | Two-Level Hybrid Sampled Simulation of Multithreaded ApplicationsabstractSampled microarchitectural simulation of single-threaded applications is mature technology for over a decade now. Sampling multithreaded applications, on the other hand, is much more complicated. Not until very recently have researchers proposed solutions for sampled simulation of multithreaded applications. Time-Based Sampling (TBS) samples multithreaded application execution based on time—not instructions as is typically done for single-threaded applications—yielding estimates for a multithreaded application’s execution time. In this article, we revisit and analyze previously proposed TBS approaches (periodic and cantor fractal based sampling), and we obtain a number of novel and surprising insights, such as (i) accurately estimating fast-forwarding IPC , that is, performance in-between sampling units, is more important than accurately estimating sample IPC , that is, performance within the sampling units; (ii) fast-forwarding IPC estimation accuracy is determined by both the sampling unit distribution and how to use the sampling units to predict fast-forwarding IPC; and (iii) cantor sampling is more accurate at small sampling unit sizes, whereas periodic is more accurate at large sampling unit sizes. These insights lead to the development of Two-level Hybrid Sampling (THS) , a novel sampling methodology for multithreaded applications that combines periodic sampling’s accuracy at large time scales (i.e., uniformly selecting coarse-grain sampling units across the entire program execution) with cantor sampling’s accuracy at small time scales (i.e., the ability to accurately predict fast-forwarding IPC in-between small sampling units). The clustered occurrence of small sampling units under cantor sampling also enables shortened warmup and thus enhanced simulation speed. Overall, THS achieves an average absolute execution time prediction error of 4% while yielding an average simulation speedup of 40 × compared to detailed simulation, which is both more accurate and faster than the current state-of-the-art. Case studies illustrate THS’ ability to accurately predict relative performance differences across the design space. Chuntao Jiang, Zhibin Yu 0001, Lieven Eeckhout, Hai Jin 0001, Xiaofei Liao, Cheng-Zhong Xu 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2016 | ShenZhen transportation system (SZTS): a novel big data benchmark suite
Zhibin Yu 0001, Lieven Eeckhout, Zhengdong Bei, Fan Zhang 0019, Cheng-Zhong Xu 0001 |
J. Supercomput. | 2 |
| 2016 | RFHOC: A Random-Forest Approach to Auto-Tuning Hadoop's ConfigurationabstractHadoop is a widely-used implementation framework of the MapReduce programming model for large-scale data processing. Hadoop performance however is significantly affected by the settings of the Hadoop configuration parameters. Unfortunately, manually tuning these parameters is very time-consuming, if at all practical. This paper proposes an approach, called RFHOC, to automatically tune the Hadoop configuration parameters for optimized performance for a given application running on a given cluster. RFHOC constructs two ensembles of performance models using a random-forest approach for the map and reduce stage respectively. Leveraging these models, RFHOC employs a genetic algorithm to automatically search the Hadoop configuration space. The evaluation of RFHOC using five typical Hadoop programs, each with five different input data sets, shows that it achieves a performance speedup by a factor of 2.11$\times$on average and up to 7.4$\times$over the recently proposed cost-based optimization (CBO) approach. In addition, RFHOC's performance benefit increases with input data set size. Zhendong Bei, Zhibin Yu 0001, Cheng-Zhong Xu 0001, Lieven Eeckhout, Shengzhong Feng |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Shorter On-Line Warmup for Sampled Simulation of Multi-threaded ApplicationsabstractWarm up is a crucial issue in sampled micro architectural simulation to avoid performance bias by constructing accurate states for micro-architectural structures before each sampling unit. Not until very recently have researchers proposed Time-Based Sampling (TBS) for the sampled simulation of multi-threaded applications. However, warm up in TBS is challenging and complicated, because (i) full functional warm up in TBS causes very high overhead, limiting overall simulation speed, (ii) traditional adaptive functional warm up for sampling single-threaded applications cannot be readily applied to TBS, and (iii) check pointing is inflexible (even invalid) due to the huge storage requirements and the variations across different runs for multi-threaded applications. In this work, we propose Shorter On-Line (SOL) warm up, which employs a two-stage strategy, using 'prime' warm up in the first stage, and an extended 'No-State-Loss (NSL)' method in the second stage. SOL is a single-pass, on-line warm up technique that addresses the warm up challenges posed in TBS in parallel simulators. SOL is highly accurate and efficient, providing a good trade-off between simulation accuracy and speed, and is easily deployed to different TBS techniques. For the PARSEC benchmarks on a simulated 8-core system, two state-of-the-art TBS techniques with SOL warm up provide a 7.2× and 37× simulation speedup over detailed simulation, respectively, compared to 3.1× and 4.5× under full warm up. SOL sacrifices only 0.3% in absolute execution time prediction accuracy on average. Chuntao Jiang, Zhibin Yu 0001, Hai Jin 0001, Xiaofei Liao, Lieven Eeckhout, Yonggang Zeng, Cheng-Zhong Xu 0001 |
ICPP | 2 |
| 2015 | SZTS: A Novel Big Data Transportation System Benchmark SuiteabstractData analytics is at the core of the supply chain for both products and services in modern economies and societies. Big data workloads however, are placing unprecedented demands on computing technologies, calling for a deep understanding and characterization of these emerging workloads. In this paper, we propose Shen Zhen Transportation System (SZTS), a novel big data Hadoop benchmark suite comprised of real-life transportation analysis applications with real-life input data sets from Shenzhen in China. SZTS uniquely focuses on a specific and real-life application domain whereas other existing Hadoop benchmark suites, such as Hi Bench and Cloud Rank-D, consist of generic algorithms with synthetic inputs. We perform a cross-layer workload characterization at both the job and micro architecture level, revealing unique characteristics of SZTS compared to existing Hadoop benchmarks as well as general-purpose multi-core PARSEC benchmarks. We also study the sensitivity of workload behavior with respect to input data size, and propose a methodology for identifying representative input data sets. Zhibin Yu 0001, Lieven Eeckhout, Zhengdong Bei, Fan Zhang 0019, Cheng-Zhong Xu 0001 |
ICPP | 2 |
| 2015 | GPGPU-MiniBench: Accelerating GPGPU Micro-Architecture SimulationabstractGraphics processing units (GPU), due to their massive computational power with up to thousands of concurrent threads and general-purpose GPU (GPGPU) programming models such as CUDA and OpenCL, have opened up new opportunities for speeding up general-purpose parallel applications. Unfortunately, pre-silicon architectural simulation of modern-day GPGPU architectures and workloads is extremely time-consuming. This paper addresses the GPGPU simulation challenge by proposing a framework, called GPGPU-MiniBench, for generating miniature, yet representative GPGPU workloads. GPGPU-MiniBench first summarizes the inherent execution behavior of existing GPGPU workloads in a profile. The central component in the profile is the Divergence Flow Statistics Graph (DFSG), which characterizes the dynamic control flow behavior including loops and branches of a GPGPU kernel. GPGPU-MiniBench generates a synthetic miniature GPGPU kernel that exhibits similar execution characteristics as the original workload, yet its execution time is much shorter thereby dramatically speeding up architectural simulation. Our experimental results show that GPGPU-MiniBench can speed up GPGPU architectural simulation by a factor of 49× on average and up to 589×, with an average IPC error of 4.7 percent across a broad set of GPGPU benchmarks from the CUDA SDK, Rodinia and Parboil benchmark suites. We also demonstrate the usefulness of GPGPU-MiniBench for driving GPU architecture exploration. Zhibin Yu 0001, Lieven Eeckhout, Nilanjan Goswami, Tao Li 0006, Lizy Kurian John, Hai Jin 0001, Cheng-Zhong Xu 0001, Junmin Wu |
IEEE Trans. Computers | 1 |
| 2013 | A characterization of big data benchmarksabstractRecently, big data has been evolved into a buzzword from academia to industry all over the world. Benchmarks are important tools for evaluating an IT system. However, benchmarking big data systems is much more challenging than ever before. First, big data systems are still in their infant stage and consequently they are not well understood. Second, big data systems are more complicated compared to previous systems such as a single node computing platform. While some researchers started to design benchmarks for big data systems, they do not consider the redundancy between their benchmarks. Moreover, they use artificial input data sets rather than real world data for their benchmarks. It is therefore unclear whether these benchmarks can be used to precisely evaluate the performance of big data systems. In this paper, we first analyze the redundancy among benchmarks from ICTBench, HiBench and typical workloads from real world applications: spatio-temporal data analysis for Shenzhen transportation system. Subsequently, we present an initial idea of a big data benchmark suite for spatio-temporal data. There are three findings in this work: (1) redundancy exists in these pioneering benchmark suites and some of them can be removed safely. (2) The workload behavior of trajectory data analysis applications is dramatically affected by their input data sets. (3) The benchmarks created for academic research cannot represent the cases of real world applications. Zhibin Yu 0001, Zhendong Bei, Juanjuan Zhao 0001, Fan Zhang 0019, Yubin Zou, Ye Li 0002, Cheng-Zhong Xu 0001 |
IEEE BigData | 2 |
| 2013 | Application-Aware Workload Consolidation to Minimize Both Energy Consumption and Network Load in Cloud EnvironmentsabstractIn this paper we tackle the problem of virtual machine (VM) placement onto physical servers to jointly optimize two objective functions. The first objective is to minimize the total energy spent within a cloud due to the servers that are commissioned to satisfy the computational demands of VMs. The second objective is to minimize the total network overhead incurred due to: (a) communicational dependencies between VMs, and (b) the VM migrations performed for the transition from an old assignment scheme to a new one. We study different methodologies for solving the aforementioned problem. The first approach is based on VM packing algorithms that optimize the above objective functions separately, reaching a single solution. The other approach is to tackle simultaneously the two optimization targets and define a set of non-dominating solutions. Performance evaluation using simulation experiments reveals interesting trade-offs between energy consumption and network load. Nikos Tziritas, Cheng-Zhong Xu 0001, Thanasis Loukopoulos, Samee Ullah Khan, Zhibin Yu 0001 |
ICPP | 5 |
| 2013 | Accelerating GPGPU architecture simulationabstractRecently, graphics processing units (GPUs) have opened up new opportunities for speeding up general-purpose parallel applications due to their massive computational power and up to hundreds of thousands of threads enabled by programming models such as CUDA. However, due to the serial nature of existing micro-architecture simulators, these massively parallel architectures and workloads need to be simulated sequentially. As a result, simulating GPGPU architectures with typical benchmarks and input data sets is extremely time-consuming. This paper addresses the GPGPU architecture simulation challenge by generating miniature, yet representative GPGPU kernels. We first summarize the static characteristics of an existing GPGPU kernel in a profile, and analyze its dynamic behavior using the novel concept of the divergence flow statistics graph (DFSG). We subsequently use a GPGPU kernel synthesizing framework to generate a miniature proxy of the original kernel, which can reduce simulation time significantly. The key idea is to reduce the number of simulated instructions by decreasing per-thread iteration counts of loops. Our experimental results show that our approach can accelerate GPGPU architecture simulation by a factor of 88X on average and up to 589X with an average IPC relative error of 5.6%. Zhibin Yu 0001, Lieven Eeckhout, Nilanjan Goswami, Tao Li 0006, Lizy Kurian John, Hai Jin 0001, Cheng-Zhong Xu 0001 |
SIGMETRICS | 1 |
| 2013 | PCantorSim: Accelerating parallel architecture simulation through fractal-based samplingabstractComputer architects rely heavily on microarchitecture simulation to evaluate design alternatives. Unfortunately, cycle-accurate simulation is extremely slow, being at least 4 to 6 orders of magnitude slower than real hardware. This longstanding problem is further exacerbated in the multi-/many-core era, because single-threaded simulation performance has not improved much, while the design space has expanded substantially. Parallel simulation is a promising approach, yet does not completely solve the simulation challenge. Furthermore, existing sampling techniques, which are widely used for single-threaded applications, do not readily apply to multithreaded applications as thread interaction and synchronization must now be taken into account. This work presents PCantorSim , a novel Cantor set (a classic fractal)--based sampling scheme to accelerate parallel simulation of multithreaded applications. Through the use of the proposed methodology, only less than 5% of an application's execution time is simulated in detail. We have implemented our approach in Sniper (a parallel multicore simulator) and evaluated it by running the PARSEC benchmarks on a simulated 8-core system. The results show that PCantorSim increases simulation speed over detailed parallel simulation by a factor of 20×, on average, with an average absolute execution time prediction error of 5.3%. Chuntao Jiang, Zhibin Yu 0001, Hai Jin 0001, Cheng-Zhong Xu 0001, Lieven Eeckhout, Wim Heirman, Trevor E. Carlson, Xiaofei Liao |
ACM Trans. Archit. Code Optim. | 2 |
| 2012 | FractalMRC: Online Cache Miss Rate Curve Prediction on Commodity SystemsabstractShared caches in chip multi-processors (CMPs) have important benefits such as accelerating inter-core communication, yet the inherent cache contention among multiple processes on such architectures can significantly degrade performance. To address this problem, cache partitioning has been studied based on the prediction of the cache miss rate curve (MRC) of the concurrently running programs. On-line MRC prediction, however, either requires special hardware support or incurs a high overhead when conducted purely in software. This paper presents a new MRC prediction scheme based on a fractal model and hence called Fractal MRC. It uses the easily available features in performance monitoring units of modern Intel processors and predicts the MRC of a running program with low overhead and high accuracy. No changes to applications and hardware are required. The prediction is validated against the measured results for 26 applications from SPEC CPU2006 benchmark suite. The highest prediction accuracy is 99.3%, the accuracy of 12 of the applications is over 80%, and the average is 76%. The cost of prediction is 2% slowdown on average. The new, efficient and accurate MRC prediction has enabled a dynamic technique to partition cache between pairs of applications at run time to match or exceed the best performance attainable with static cache partitioning. Lulu He, Zhibin Yu 0001, Hai Jin 0001 |
IPDPS | 2 |
| 2010 | System-level max power (SYMPO): a systematic approach for escalating system-level power consumption using synthetic benchmarksabstractTo effectively design a computer system for the worst case power consumption scenario, system architects often use hand-crafted maximum power consuming benchmarks at the assembly language level. These stressmarks, also called power viruses, are very tedious to generate and require significant domain knowledge. In this paper, we propose SYMPO, an automatic SYstem level Max POwer virus generation framework, which maximizes the power consumption of the CPU and the memory system using genetic algorithm and an abstract workload generation framework. For a set of three ISAs, we show the efficacy of the power viruses generated using SYMPO by comparing the power consumption with that of MPrime torture test, which is widely used by industry to test system stability. Our results show that the usage of SYMPO results in the generation of power viruses that consume 14-41% more power compared to MPrime on SPARC ISA. The genetic algorithm achieved this result in about 70 to 90 generations in 11 to 15 hours when using a full system simulator. We also show that the power viruses generated in the Alpha ISA consume 9-24% more power compared to the previous approach of stressmark generation. We measure and provide the power consumption of these benchmarks on hardware by instrumenting a quad-core AMD Phenom II X4 system. The SYMPO power virus consumes more power compared to various industry grade power viruses on x86 hardware. We also provide a microarchitecture independent characterization of various industry standard power viruses. Karthik Ganesan 0006, Jungho Jo, William Lloyd Bircher, Dimitris Kaseridis, Zhibin Yu 0001, Lizy Kurian John |
PACT | 5 |
| 2010 | CantorSim: Simplifying Acceleration of Micro-architecture SimulationsabstractMany techniques have been developed to accelerate micro-architecture simulation since it is becoming increasingly urgent as the complexity of workloads and simulated processors increases. However, most of popular techniques need profiles or trial simulations to determine parameters before real simulations. When the number of dynamic instructions of workloads such as SPEC CPU2006 is huge, the profiles or trial simulations often take a long time. What's worse, any changes of benchmarks need to repeat the above work. This paper proposes a novel approach based on fractals to solve this problem. It employs the generation procedure of trisection Cantor Set, a classic fractal, to select instructions simulated in detail. We implement a simulator named CantorSim using the trisection Cantor fractal approach. The simulator is evaluated with benchmark programs from SPEC CPU2000 and SPEC CPU2006. We show that our approach is micro-architecture independent as the same parameter value for a single program yields similar results on processors with different configurations. The unified analytical model used to determine sampling parameters for different benchmarks indicates the proposed approach can save a lot of time for preparing parameters. On the other hand, for SPEC CPU2000 and SPEC CPU2006, CantorSim is slightly faster than SMARTSim. The validated average CPI relative errors of the SPEC CPU2000 and SPEC CPU2006 are 3.2% and 2.2% respectively. Therefore, our proposed approach can simplify acceleration of simulations significantly without sacrificing speed and accuracy. Zhibin Yu 0001, Hai Jin 0001, Jian Chen 0030, Lizy Kurian John |
MASCOTS | 1 |
| 2009 | TSS: Applying two-stage sampling in micro-architecture simulationsabstractAccelerating micro-architecture simulation is becoming increasingly urgent as the complexity of workload and simulated processor increases. This paper presents a novel two-stage sampling (TSS) scheme to accelerate the sampling-based simulation. It firstly selects some large samples from a dynamic instruction stream as candidates of detail simulation and then samples some small groups from each selected first stage sample to do detail simulation. Since the distribution of standard deviation of cycle per instruction (CPI) is insensitive to microarchitecture, TSS could be used to speedup design space exploration by splitting the sampling process into two stages, which is able to remove redundant instruction samples from detail simulation when the program is in stable program phase (standard deviation of CPI is near zero). It also adopts systematic sampling to accelerate the functional warm-up in sampling simulation. Experimental results show that, by combining these two techniques, TSS achieves an average and maximum speedup of 1.3 and 2.29 over SMARTS, with the average CPI relative error is less than 3%. TSS could significantly accelerate the time consuming iterative early design evaluation process. Zhibin Yu 0001, Hai Jin 0001, Jian Chen 0030, Lizy Kurian John |
MASCOTS | 1 |