Ruijin Zhou

dblp:129/2868 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Energy-efficient computing · 65% Memory systems · 15% Cloud and datacenter computing · 15%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing
datacenter power management
0.322013
Enabling datacenter servers to scale out economically and sustainably · MICRO 2013
Enabling distributed generation powered sustainable high-performance data center · HPCA 2013
Memory systems › non-volatile memory
phase change memory
0.212013
Exploring high-performance and energy proportional interface for phase change memory systems · HPCA 2013
Energy-efficient computing
power management
0.212013
Enabling distributed generation powered sustainable high-performance data center · HPCA 2013
Energy-efficient computing › datacenter power management
power oversubscription
0.212013
Enabling datacenter servers to scale out economically and sustainably · MICRO 2013
Energy-efficient computing › datacenter power management
power provisioning
0.012013
Enabling datacenter servers to scale out economically and sustainably · MICRO 2013

Methods — techniques the papers use, named apart from their topics

workload-aware optimization · 0.2rank partitioning · 0.2load following · 0.2channel partitioning · 0.2
YearPublicationVenuePosition
2022 GE-IDS: an intrusion detection system based on grayscale and entropy
Dan Liao, Ruijin Zhou, Hui Li 0067
Peer-to-Peer Netw. Appl.2
2017 Performance Analysis of Large Scale Distributed Systems by Ranking Dominant Features
abstract
Large scale distributed systems generate so many metrics/features to monitor and analyze their performances. Most of the analysis to detect performance regressions and root causes of the regressions is done by manually looking at graphs of the metric values. This approach is error prone and doesn't scale. There are some recent works to automate root cause analysis. However these schemes either rely on specific network metrics or aggregate values such as median. Such approach, which doesn't use various physical and virtual device specific metrics, is ineffective in detecting anomalies and their root causes for large scale distributed systems/clusters such as VMware Virtual Storage Area Network (vSAN) and vSphere.
Debessay Fesehaye, Lenin Singaravelu, Amitabha Banerjee, Ruijin Zhou, Xiaobo Huang, Chien-Chia Chen, Rajesh Somasundaran
BDCAT4
2017 Group Clustering Using Inter-Group Dissimilarities
abstract
Various systems have natural groupings. For instance in large scale distributed system, we can have groups of virtual and/or physical devices. A system can also have groups of time series datasets collected at different time intervals. Such groups are usually characterized by multidimensional metrics (features) set. Clustering such groups using their multidimensional datasets has various applications, such as identifying different performance levels for anomaly detection and load balancing. Traditional algorithms focus on clustering a single time series dataset and not on such groups with multidimensional metrics datasets. In this paper, we present the design, implementation and analysis of two sets of group clustering algorithms. The first set is called one-to-many as it generates clusters of groups by comparing each group against all other groups. The second set of algorithms is called pairwise as it generates the clusters of groups using pairwise group dissimilarity matrix. Both sets of algorithms first generate group dissimilarity weights using metric ranking algorithms. We implemented the group clustering algorithms by extending a well known machine learning package and using a front-end visualizer.We validated the clustering algorithms using real world datasets on the VMware vSAN product. Experimental results show that 7 out of the 8 proposed algorithms can generate expected clusters in at least 4 out of the 6 detailed experiments. In 5 out of the 6 experiments, 3 out of the 8 proposed algorithms can generate the expected clusters. One of the pairwise algorithms can generate the expected clusters in all 6 of the 6 experiments.
Debessay Fesehaye, Lenin Singaravelu, Chien-Chia Chen, Xiaobo Huang, Amitabha Banerjee, Ruijin Zhou, Rajesh Somasundaran
ICDCS6
2015 Towards Lightweight and Swift Storage Resource Management in Big Data Cloud Era
abstract
Workload IO behavior in modern data centers is fluctuating and unpredictable due to the rapidly adopted, public cloud environment. Nevertheless, existing storage resource management systems, such as VMware SDRS, are incapable of performing real time policy-based storage management due to the high cost of migrating large size virtual disks. Hence, the traditional storage management schemes become ineffective due to the lack of quick response to the frequent IO bursts and the inaccurate storage latency prediction in the light of a highly fluctuating environment. To address the aforementioned issues, we propose LightSRM, which can work properly in a time-variant cloud environment. To mitigate the storage migration cost, we leverage copy-on-write/read snapshots to redirect the IO requests without moving the virtual disk. To support snapshots in storage management, we also build a performance model specifically for snapshots. We employ exponentially weighted moving average with adjustable sliding window to provide quick and accurate performance prediction. Furthermore, we propose a hybrid management scheme, which can dynamically choose either snapshot or migration for fastest performance tuning. We build our prototype in a QEMU/KVM based virtualized environment. Our empirical evaluation results show that snapshot can redirect IO requests in a faster manner than migration can do when the virtual disk size is large. Besides, snapshot method has less disk performance impact on the applications. By employing hybrid snapshot/migration method, LightSRM yields less overall latency, better load balance, and less IO traffic overhead.
Ruijin Zhou, Huixiang Chen 0001, Tao Li 0006
ICS1
2014 An end-to-end analysis of file system features on sparse virtual disks
abstract
Software Defined Data Center (SDDC) is now an emerging area drawing considerable attention in enterprise computing. Software-Defined Storage (SDS), as a key element to enable the SDDC concept, is considered one of the most disruptive storage technologies in modern times. SDS introduces a variety of novel features and functionalities thereby changing the traditional view of the storage stack. In VMware's ESXi virtualization platform, several sparse virtual disk formats have been implemented to support critical features for SDS such as virtual machine (VM) snapshots, Fault-Tolerance (FT), thin provisioning and linked clones. Each virtual disk format supports unique features that may incur complex interactions with other layers of the storage stack such as guest file systems and storage devices.
Ruijin Zhou, Sankaran Sivathanu, Jinpyo Kim, Bing Tsai, Tao Li 0006
ICS1
2014 On Characterization of Performance and Energy Efficiency in Heterogeneous HPC Cloud Data Centers
abstract
The relocation of high performance computing systems (HPC) to the cloud poses new challenges for data center architects and IT managers. These challenges are due to heterogeneity injected into data centers by cutting-edge virtualization technologies and hardware accelerators used to support emerging cloud applications and services. Although hardware accelerators like General Purpose Graphics Processing Units (GPGPUs) and virtualization technologies have been well studied and evaluated individually, a detailed analysis of their combined architectures and collective behavior from the data center point of view is lacking. Using real platforms and high performance computing workloads, we study the power performance tradeoffs due to various granularities of heterogeneity across hardware and software layers and expose hidden opportunities for optimizing overall data center efficiency. Our approach is to evaluate server power and performance from a data center point of view as opposed to evaluating hardware accelerators and virtualization technologies themselves. Our results show that performance on cloud is affected by virtualization overhead and fraction of serial code. Moreover, GPU workloads achieve 25% and 30% savings in power and energy consumption when executed on low power platforms, and only 50% of our GPU workloads are more energy efficient than their corresponding CPU implementations. The results also show that it is much more power efficient to collocate GPU virtual machines with non-GPU virtual machines.
Amer Qouneh, Nilanjan Goswami, Ruijin Zhou, Tao Li 0006
MASCOTS3
2014 Towards Automated Provisioning and Emergency Handling in Renewable Energy Powered Datacenters
Chao Li 0009, Rui Wang 0014, Yang Hu 0001, Ruijin Zhou, Ming Liu 0006, Longjun Liu, Jingling Yuan, Tao Li 0006, Depei Qian 0001
J. Comput. Sci. Technol.4
2013 Enabling distributed generation powered sustainable high-performance data center
abstract
The necessity for capping carbon emission has significantly restricted the potential of modern data centers. For this matter, both industry and academia are proactively seeking opportunities on cross-layer power management schemes that could open a door for sustainable high-performance computing platform. In this paper we investigate an emerging trend in the IT industry: using promising onsite distributed generation (DG) techniques to provide premium clean energy to the computing load. We develop data center power demand shaping (PDS), a novel technique that allows data centers to utilize onsite green energy efficiently. In contrast to prior design, PDS takes advantage of a so-far unexplored power supply feature, i.e., the load following capabilities of DG systems to avoid the high performance penalty issue incurred during supply tracking. In addition, PDS features two adaptive power management schemes: DGR Boost and UPS Boost. These two workload-aware optimization methods leverage mature computer tuning knobs to achieve attractive data center performance improvement. Using real-world data center traces and industry data of distributed generation systems, we show that our technique can come within 1.2% performance of an ideal oracle, which is roughly a 37% improvement over existing supply tracking based design. Our design could save over 100 metric tons of carbon emissions annually for a 10MW data center.
Chao Li 0009, Ruijin Zhou, Tao Li 0006
HPCA2
2013 Exploring high-performance and energy proportional interface for phase change memory systems
abstract
Phase change memory is emerging as a promising candidate for building up future energy efficient memory systems. To achieve high-performance and energy proportional design, phase change memory devices need to be reorganized so that (1) the relatively long latency of phase change memory devices should be hidden; (2) unnecessary power waste of phase change memory need to be preserved. Previous studies show that conventional memory ranks could be broken down into multiple smaller ranks for increased concurrency and lower power consumption. Nevertheless, the conventional electrical bus is incapable of supporting a large number of memory chips due to its insufficient load capacity and signal traversing speed. In this paper, we propose a phase change memory system design that leverages the state-of-art photonic links to overcome this issue. Moreover, thanks to the flexibility of photonic links, it is possible to amortize the small-rank penalty (e.g. the rank-to-rank switch overhead) by partitioning the channels either statically or dynamically. Our experimental results show that photonically interconnected phase change memory can increase the system performance (IPC) by up to 19% while saving 35% memory system power.
Zhong-Qi Li, Ruijin Zhou, Tao Li 0006
HPCA2
2013 Enabling datacenter servers to scale out economically and sustainably
abstract
As cloud applications proliferate and data-processing demands increase, server resources must grow to unleash the performance of emerging workloads that scale well with large number of compute nodes. Nevertheless, power has become a crucial bottleneck that restricts horizontal scaling (scale out) of server systems, especially in datacenters that employ power over-subscription. When a datacenter hits the maximum capacity of its power provisioning equipment, the owner has to either build another facility or upgrade existing utility power infrastructure -- both approaches add huge capital expenditure, require significant construction lead time, and can further increase the owner's carbon footprint.
Chao Li 0009, Yang Hu 0001, Ruijin Zhou, Ming Liu 0006, Longjun Liu, Jingling Yuan, Tao Li 0006
MICRO3
2013 Leveraging phase change memory to achieve efficient virtual machine execution
abstract
Virtualization technology is being widely adopted by servers and data centers in the cloud computing era to improve resource utilization and energy efficiency. Nevertheless, the heterogeneous memory demands from multiple virtual machines (VM) make it more challenging to design efficient memory systems. Even worse, mission critical VM management activities (e.g. checkpointing) could incur significant runtime overhead due to intensive IO operations. In this paper, we propose to leverage the adaptable and non-volatile features of the emerging phase change memory (PCM) to achieve efficient virtual machine execution. Towards this end, we exploit VM-aware PCM management mechanisms, which 1) smartly tune SLC/MLC page allocation within a single VM and across different VMs and 2) keep critical checkpointing pages in PCM to reduce I/O traffic. Experimental results show that our single VM design (IntraVM) improves performance by 10% and 20% compared to pure SLC- and MLC- based systems. Further incorporating VM-aware resource management schemes (IntraVM+InterVM) increases system performance by 15%. In addition, our design saves 46% of checkpoint/restore duration and reduces 50% of overall IO penalty to the system.
Ruijin Zhou, Tao Li 0006
VEE1
2013 Optimizing virtual machine live storage migration in heterogeneous storage environment
abstract
Virtual machine (VM) live storage migration techniques significantly increase the mobility and manageability of virtual machines in the era of cloud computing. On the other hand, as solid state drives (SSDs) become increasingly popular in data centers, VM live storage migration will inevitably encounter heterogeneous storage environments. Nevertheless, conventional migration mechanisms do not consider the speed discrepancy and SSD's wear-out issue, which not only causes significant performance degradation but also shortens SSD's lifetime. This paper, for the first time, addresses the efficiency of VM live storage migration in heterogeneous storage environments from a multi-dimensional perspective, i.e., user experience, device wearing, and manageability. We derive a flexible metric (migration cost), which captures various design preference. Based on that, we propose and prototype three new storage migration strategies, namely: 1) Low Redundancy (LR), which generates the least amount of redundant writes; 2) Source-based Low Redundancy (SLR), which keeps the balance between IO performance and write redundancy; and 3) Asynchronous IO Mirroring, which seeks the highest IO performance. The evaluation of our prototyped system shows that our techniques outperform existing live storage migration by a significant margin. Furthermore, by adaptively mixing our proposed schemes, the cost of massive VM live storage migration can be even lower than that of only using the best of individual mechanism.
Ruijin Zhou, Fang Liu 0002, Chao Li 0009, Tao Li 0006
VEE1