Shuang Chen 0002

dblp:07/6433-2 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Cloud and datacenter computing · 58% Memory systems · 21% Energy-efficient computing · 14%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 100%

Topics — the 22 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
resource management
1.322024
UFO: The Ultimate QoS-Aware Core Management for Virtualized and Oversubscribed Public Clouds · NSDI 2024
PIMCloud: QoS-Aware Resource Management of Latency-Critical Applications in Clouds with Processing-in-Memory · HPCA 2022
Cloud and datacenter computing
virtualization
0.812024
UFO: The Ultimate QoS-Aware Core Management for Virtualized and Oversubscribed Public Clouds · NSDI 2024
Energy-efficient computing › power management
dynamic voltage and frequency scaling
0.612022
ReTail: Opting for Learning Simplicity to Enable QoS-Aware Power Management in the Cloud · HPCA 2022
Energy-efficient computing
power management
0.612022
ReTail: Opting for Learning Simplicity to Enable QoS-Aware Power Management in the Cloud · HPCA 2022
Memory systems
processing-in-memory
0.612022
PIMCloud: QoS-Aware Resource Management of Latency-Critical Applications in Clouds with Processing-in-Memory · HPCA 2022
Cloud and datacenter computing › resource management
qos-aware resource management
0.612022
PIMCloud: QoS-Aware Resource Management of Latency-Critical Applications in Clouds with Processing-in-Memory · HPCA 2022
Cloud and datacenter computing
latency-critical applications
0.532022
PIMCloud: QoS-Aware Resource Management of Latency-Critical Applications in Clouds with Processing-in-Memory · HPCA 2022
ReTail: Opting for Learning Simplicity to Enable QoS-Aware Power Management in the Cloud · HPCA 2022
PARTIES: QoS-Aware Resource Partitioning for Multiple Interactive Services · ASPLOS 2019
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.412019
PARTIES: QoS-Aware Resource Partitioning for Multiple Interactive Services · ASPLOS 2019
Cloud and datacenter computing
multi-tenancy
0.412019
PARTIES: QoS-Aware Resource Partitioning for Multiple Interactive Services · ASPLOS 2019
Cloud and datacenter computing › quality of service
qos-aware scheduling
0.412019
PARTIES: QoS-Aware Resource Partitioning for Multiple Interactive Services · ASPLOS 2019
Cloud and datacenter computing › resource allocation › resource allocation policy
resource partitioning
0.412019
PARTIES: QoS-Aware Resource Partitioning for Multiple Interactive Services · ASPLOS 2019
Memory systems
cache management
0.312017
SWAP: Effective Fine-Grain Management of Shared Last-Level Caches with Minimum Hardware Support · HPCA 2017
Memory systems › cache management
cache partitioning
0.312017
SWAP: Effective Fine-Grain Management of Shared Last-Level Caches with Minimum Hardware Support · HPCA 2017
Memory systems › memory hierarchy › cache hierarchy
last-level cache
0.312017
SWAP: Effective Fine-Grain Management of Shared Last-Level Caches with Minimum Hardware Support · HPCA 2017
Query processing and optimization
sorting
0.212016
A Study of Sorting Algorithms on Approximate Memory · SIGMOD Conference 2016
Storage systems › energy-efficient storage
approximate storage
0.212016
A Study of Sorting Algorithms on Approximate Memory · SIGMOD Conference 2016
Memory systems
non-volatile memory
0.212016
A Study of Sorting Algorithms on Approximate Memory · SIGMOD Conference 2016
Operating systems › resource management › memory management
page coloring
0.112017
SWAP: Effective Fine-Grain Management of Shared Last-Level Caches with Minimum Hardware Support · HPCA 2017
Processor architecture and microarchitecture
chip multiprocessor
0.112017
SWAP: Effective Fine-Grain Management of Shared Last-Level Caches with Minimum Hardware Support · HPCA 2017
Processor architecture and microarchitecture
many-core architecture
0.112017
SWAP: Effective Fine-Grain Management of Shared Last-Level Caches with Minimum Hardware Support · HPCA 2017
Emerging computing paradigms
approximate computing
0.112016
A Study of Sorting Algorithms on Approximate Memory · SIGMOD Conference 2016
High-performance computing › lossy compression
precision-resolution trade-off
0.112016
A Study of Sorting Algorithms on Approximate Memory · SIGMOD Conference 2016

Methods — techniques the papers use, named apart from their topics

resource scheduling · 0.6performance evaluation · 0.6latency prediction · 0.6feature selection · 0.6colocation · 0.6hybrid storage · 0.5approx-refine execution · 0.5hardware and software resource partitioning · 0.4
YearPublicationVenuePosition
2024 UFO: The Ultimate QoS-Aware Core Management for Virtualized and Oversubscribed Public Clouds
Yajuan Peng, Shuang Chen 0002, Yi Zhao 0018
NSDI2
2022 ReTail: Opting for Learning Simplicity to Enable QoS-Aware Power Management in the Cloud
abstract
Many cloud services have Quality-of-Service (QoS) requirements; most requests have to to complete within a given latency constraint. Recently, researchers have begun to investigate whether it is possible to meet QoS while attempting to save power on a per-request basis. Existing work shows that one can indeed hand-tune a request latency predictor offline for a particular cloud application, and consult it at runtime to modulate CPU voltage and frequency, resulting in substantial power savings.In this paper, we propose ReTail, an automated and general solution for request-level power management of latency-critical services with QoS constraints. We present a systematic process to select the features of any given application that best correlate with its request latency. ReTail uses these features to predict latency, and adjust CPU’s power consumption. ReTail’s predictor is trained fully at runtime. We show that unlike previous findings, simple techniques perform better than complex machine learning models, when using the right input features. For a web search engine, ReTail outperforms prior mechanisms based on complex hand-tuned predictors for that application domain. Furthermore, ReTail’s systematic approach also yields superior power savings across a diverse set of cloud applications.
Shuang Chen 0002, Angela Jin, Christina Delimitrou, José F. Martínez
HPCA1
2022 PIMCloud: QoS-Aware Resource Management of Latency-Critical Applications in Clouds with Processing-in-Memory
abstract
The slowdown of Moore’s Law, combined with advances in 3D stacking of logic and memory, have pushed architects to revisit the concept of processing-in-memory (PIM) to overcome the memory wall bottleneck. This PIM renaissance finds itself in a very different computing landscape from the one twenty years ago, as more and more computation shifts to the cloud. Most PIM architecture papers still focus on best-effort applications, while PIM’s impact on latency-critical cloud applications is not well understood.This paper explores how datacenters can exploit PIM architectures in the context of latency-critical applications. We adopt a general-purpose cloud server with HBM-based, 3D-stacked logic+memory modules, and study the impact of PIM on six diverse interactive cloud applications. We reveal the previously neglected opportunity that PIM presents to these services, and show the importance of properly managing PIM-related resources to meet the QoS targets of interactive services and maximize resource efficiency. Then, we present PIMCloud, a QoS-aware resource manager designed for cloud systems with PIM allowing colocation of multiple latency-critical and best-effort applications. We show that PIMCloud efficiently manages PIM resources: it (1) improves effective machine utilization by up to 70% and 85% (average 24% and 33%) under 2-app and 3-app mixes, compared to the best state-of-the-art manager; (2) helps latency-critical applications meet QoS; and (3) adapts to varying load patterns.
Shuang Chen 0002, Christina Delimitrou, José F. Martínez
HPCA1
2019 PARTIES: QoS-Aware Resource Partitioning for Multiple Interactive Services
abstract
Multi-tenancy in modern datacenters is currently limited to a single latency-critical, interactive service, running alongside one or more low-priority, best-effort jobs. This limits the efficiency gains from multi-tenancy, especially as an increasing number of cloud applications are shifting from batch jobs to services with strict latency requirements. We present PARTIES, a QoS-aware resource manager that enables an arbitrary number of interactive, latency-critical services to share a physical node without QoS violations. PARTIES leverages a set of hardware and software resource partitioning mechanisms to adjust allocations dynamically at runtime, in a way that meets the QoS requirements of each co-scheduled workload, and maximizes throughput for the machine. We evaluate PARTIES on state-of-the-art server platforms across a set of diverse interactive services. Our results show that PARTIES improves throughput under QoS by 61% on average, compared to existing resource managers, and that the rate of improvement increases with the number of co-scheduled applications per physical host.
Shuang Chen 0002, Christina Delimitrou, José F. Martínez
ASPLOS1
2017 SWAP: Effective Fine-Grain Management of Shared Last-Level Caches with Minimum Hardware Support
abstract
Performance isolation is an important goal in server-class environments. Partitioning the last-level cache of a chip multiprocessor (CMP) across co-running applications has proven useful in this regard. Two popular approaches are (a) hardware support for way partitioning, or (b) operating system support for set partitioning through page coloring. Unfortunately, neither approach by itself is scalable beyond a handful of cores without incurring in significant performance overheads. We propose SWAP, a scalable and fine-grained cache management technique that seamlessly combines set and way partitioning. By cooperatively managing cache ways and sets, SWAP (“Set and WAy Partitioning”) can successfully provide hundreds of fine-grained cache partitions for the manycore era. SWAP requires no additional hardware beyond way partitioning. In fact, SWAP can be readily implemented in existing commercial servers whose processors do support hardware way partitioning. In this paper, we prototype SWAP on a 48-core Cavium ThunderX platform running Linux, and we show average speedups over no cache partitioning that are twice as large as those attained with ThunderX's hardware way partitioning alone.
Shuang Chen 0002, Jeff Setter, José F. Martínez
HPCA2
2017 Bank Stealing for a Compact and Efficient Register File Architecture in GPGPU
abstract
Modern general-purpose graphic processing units (GPGPUs) have emerged as pervasive alternatives for parallel high-performance computing. The extreme multithreading in modern GPGPUs demands a large register file (RF), which is typically organized into multiple banks to support the massive parallelism. Although a heavily banked structure benefits RF throughput, its associated area and energy costs with diminishing performance gains greatly limit the future RF scaling. In this paper, we propose an improved RF design with bank stealing techniques, which enable a high RF throughput with compact area. By deeply investigating the GPGPU microarchitecture, we find that the state-of-the-art RF designs' is far from optimal due to the deficiency in bank utilization, which is the intrinsic limitation to a high RF throughput and a compact RF area. We investigate the causes for bank conflicts and identify that most conflicts can be eliminated by leveraging the fact that the highly banked RF oftentimes experiences underutilization. This is especially true in GPGPUs, where multiple ready warps are available at the scheduling stage with their operands to be wisely coordinated. In this paper, we propose two lightweight bank stealing techniques that can opportunistically fill the idle banks and register entries for better operand service. Using the proposed architecture, the average GPGPU performance can be improved under a smaller energy budget with significant area saving, which makes it promising for sustainable RF scaling.
Naifeng Jing, Shunning Jiang, Shuang Chen 0002, Jingjie Zhang, Li Jiang 0002, Chao Li 0009, Xiaoyao Liang
IEEE Trans. Very Large Scale Integr. Syst.3
2016 A Study of Sorting Algorithms on Approximate Memory
abstract
Hardware evolution has been one of the driving factors for the redesign of database systems. Recently, approximate storage emerges in the area of computer architecture. It trades off precision for better performance and/or energy consumption. Previous studies have demonstrated the benefits of approximate storage for applications that are tolerant to imprecision such as image processing. However, it is still an open question whether and how approximate storage can be used for applications that do not expose such intrinsic tolerance. In this paper, we study one of the most basic operations in database--sorting on a hybrid storage system with both precise storage and approximate storage. Particularly, we start with a study of three common sorting algorithms on approximate storage. Experimental results show that a 95% sorted sequence can be obtained with up to 40% reduction in total write latencies. Thus, we propose an approx-refine execution mechanism to improve the performance of sorting algorithms on the hybrid storage system to produce precise results. Our optimization gains the performance benefits by offloading the sorting operation to approximate storage, followed by an efficient refinement to resolve the unsortedness on the output of the approximate storage. Our experiments show that our approx-refine can reduce the total memory access time by up to 11%. These studies shed light on the potential of approximate hardware for improving the performance of applications that require precise results.
Shuang Chen 0002, Shunning Jiang, Bingsheng He, Xueyan Tang
SIGMOD Conference1
2015 Bank stealing for conflict mitigation in GPGPU Register File
abstract
Modern General Purpose Graphic Processing Unit (GPGPU) demands a large Register File (RF), which is typically organized into multiple banks to support the massive parallelism. Although heavy banking benefits RF throughput, its associated area and energy costs with diminishing performance gains greatly limit future RF s-caling. In this paper, we propose an improved RF design with a bank stealing technique, which enables a high RF throughput with compact area. By deeply investigating the GPGPU microarchitecture, we identify the deficiency in the state-of-the-art RF designs as the bank conflict problem, while the majority of conflicts can be eliminated leveraging the fact that the highly-banked RF oftentimes experiences under-utilization. This is especially true in GPGPU where multiple ready warps are available at the scheduling stage with their operands to be wisely coordinated. Our lightweight bank stealing technique can opportunistically fill the idle banks for better operand service, and the average GPGPU performance can be improved under smaller energy budget with significant area saving, which makes it promising for sustainable RF scaling.
Naifeng Jing, Shuang Chen 0002, Shunning Jiang, Li Jiang 0002, Chao Li 0009, Xiaoyao Liang
ISLPED2