Sui Chen

dblp:151/4519 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
0since 2021 · last 2020
0000-0001-7251-7760ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 6 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 34% GPUs and heterogeneous computing · 26% Parallel and multicore computing · 19%
Databases, data mining, and information retrieval
1 paper
Transaction processing and concurrency control · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
non-volatile memory
0.822020
Architectural Support for NVRAM Persistence in GPUs · IEEE Trans. Parallel Distributed Syst. 2020
Efficient GPU NVRAM Persistence with Helper Warps · DAC 2019
Storage systems › transaction support
durable transactions
0.412020
Architectural Support for NVRAM Persistence in GPUs · IEEE Trans. Parallel Distributed Syst. 2020
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy
0.412020
Architectural Support for NVRAM Persistence in GPUs · IEEE Trans. Parallel Distributed Syst. 2020
GPUs and heterogeneous computing
GPU memory management
0.412019
Efficient GPU NVRAM Persistence with Helper Warps · DAC 2019
Memory systems
cache coherence
0.312017
Accelerating GPU Hardware Transactional Memory with Snapshot Isolation · ISCA 2017
Parallel and multicore computing › transactional memory
hardware transactional memory
0.312017
Accelerating GPU Hardware Transactional Memory with Snapshot Isolation · ISCA 2017
Memory systems
snapshot isolation
0.312017
Accelerating GPU Hardware Transactional Memory with Snapshot Isolation · ISCA 2017
Distributed systems › distributed coordination
conflict resolution
0.212016
Efficient GPU hardware transactional memory through early conflict resolution · HPCA 2016
GPUs and heterogeneous computing
GPU transactional memory
0.212016
Efficient GPU hardware transactional memory through early conflict resolution · HPCA 2016
Parallel and multicore computing
synchronization
0.212016
Efficient GPU hardware transactional memory through early conflict resolution · HPCA 2016
Parallel and multicore computing
transactional memory
0.212016
Efficient GPU hardware transactional memory through early conflict resolution · HPCA 2016
Transaction processing and concurrency control
concurrency control
0.112017
Accelerating GPU Hardware Transactional Memory with Snapshot Isolation · ISCA 2017
Transaction processing and concurrency control › isolation levels
snapshot isolation
0.112017
Accelerating GPU Hardware Transactional Memory with Snapshot Isolation · ISCA 2017

Methods — techniques the papers use, named apart from their topics

helper warps · 0.8write skew elimination · 0.6snapshot isolation · 0.6transactional memory · 0.4logging · 0.4asynchronous persistence · 0.4pause-and-go execution · 0.2early-abort conflict resolution · 0.2
YearPublicationVenuePosition
2020 Architectural Support for NVRAM Persistence in GPUs
abstract
Non-volatile Random Access Memories (NVRAM) have emerged in recent years to bridge the performance gap between the main memory and external storage devices, such as Solid State Drives (SSD). In addition to higher storage density, NVRAM provides byte-addressability, higher bandwidth, near-DRAM latency, and easier access compared to block devices such as traditional SSDs. This enables new programming paradigms taking advantage of durability and larger memory footprint. With the range and size of GPU workloads expanding, NVRAM will present itself as a promising addition to GPU's memory hierarchy. To utilize the non-volatility of NVRAMs, programs should allow durable stores, maintaining consistency through a power loss event. This is usually done through a logging mechanism that works in tandem with a transaction execution layer which can consist of a transactional memory or a locking mechanism. Together, this results in a transaction processing system that preserves the ACID properties. GPUs are designed with high throughput in mind, leveraging high degrees of parallelism. Transactional memory proposals enable fine-grained transactions at the GPU thread-level. However, with lower write bandwidths compared to that of DRAMs, using NVRAM as-is may yield sub-optimal overall system performance when threads experience long latency. To address this problem, we propose using Helper Warps to move persistence out of the critical path of transaction execution, alleviating the impact of latencies. Our mechanism achieves a speedup of 4.4 and 1.5 under bandwidth limits of 1.6 GB/s and 12 GB/s and is projected to maintain speed advantage even when NVRAM bandwidth gets as high as hundreds of GB/s in certain cases. Due to the speedup, our proposed method also results in reduction in overall energy consumption.
Sui Chen, Lei Liu 0037, Lu Peng 0001
IEEE Trans. Parallel Distributed Syst.1
2019 Efficient GPU NVRAM Persistence with Helper Warps
abstract
Non-volatile Random-Access Memories (NVRAM) have emerged in recent years to bridge the performance gap between the main memory and external storage devices. To utilize the non-volatility of NVRAMs, programs should allow durable stores, meaning consistency must be maintained during a power loss event. GPUs are designed with high throughput, leveraging high degrees of parallelism. However, with lower NVRAM write bandwidths compared to that of DRAMs, using NVRAM as is may yield suboptimal overall system performance. To address this problem, we propose using Helper Warps to move persistence out of the critical path of transaction execution, alleviating the impact of latencies. Our mechanism achieves a speedup of 4.4 and 1.5 under bandwidth limits of 1.6 GB/s and 12 GB/s and is projected to maintain speed advantage even when NVRAM bandwidth gets as high as hundreds of GB/s in certain cases.
Sui Chen, Faen Zhang, Lei Liu 0037, Lu Peng 0001
DAC1
2017 Accelerating GPU Hardware Transactional Memory with Snapshot Isolation
abstract
Snapshot Isolation (SI) is an established model in the database community, which permits write-read conflicts to pass and aborts transactions only on write-write conflicts. With the Write Skew anomaly correctly eliminated, SI can reduce the occurrence of aborts, save the work done by transactions, and greatly benefit long transactions involving complex data structures.
Sui Chen, Lu Peng 0001, Samuel Irving
ISCA1
2017 Soft error resilience of Big Data kernels through algorithmic approaches
Travis LeCompte, Walker Legrand, Sui Chen, Lu Peng 0001
J. Supercomput.3
2016 Efficient GPU hardware transactional memory through early conflict resolution
abstract
It has been proposed that Transactional Memory be added to Graphics Processing Units (GPUs) in recent years. One proposed hardware design, Warp TM, can scale to 1000s of concurrent transactions. As a programming method that can atomicize an arbitrary number of memory access locations and greatly reduce the efforts to program parallel applications, transactional memory handles the complexity of inter-thread synchronization. However, when thousands of transactions run concurrently on a GPU, conflicts and resource contentions arise, causing performance loss. In this paper, we identify and analyze the cause of conflicts and contentions and propose two enhancements that try to resolve conflicts early: (1) Early-Abort global conflict resolution that allows conflicts to be detected before they reach the Commit Units so that contention in the Commit Units is reduced and (2) Pause-and-Go execution scheme that reduces the chance of conflict and the performance penalty of re-executing long transactions. These two enhancements are enabled by a single hardware modification. Our evaluation shows the combination of the two enhancements greatly improves overall execution speed while reducing energy consumption.
Sui Chen, Lu Peng 0001
HPCA1
2016 Soft error resilience in Big Data kernels through modular analysis
Sui Chen, Greg Bronevetsky, Lu Peng 0001, Bin Li 0008, Xin Fu 0001
J. Supercomput.1
2015 NBTI alleviation on FinFET-made GPUs by utilizing device heterogeneity
Ying Zhang 0016, Sui Chen, Lu Peng 0001, Shaoming Chen
Integr.2
2015 A framework for evaluating comprehensive fault resilience mechanisms in numerical programs
Sui Chen, Greg Bronevetsky, Bin Li 0008, Marc Casas, Lu Peng 0001
J. Supercomput.1