Shuo Wei

dblp:58/9488 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
2since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-authorSystems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 72% GPUs and heterogeneous computing · 21% Parallel and multicore computing · 6%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › non-volatile memory
persistent memory
0.712023
PMLiteDB: Streamlining Access Paths for High-Performance Persistent Memory Document Database Systems · IEEE Trans. Computers 2023
Memory systems › cache management › cache insertion policy
selective caching
0.712023
PMLiteDB: Streamlining Access Paths for High-Performance Persistent Memory Document Database Systems · IEEE Trans. Computers 2023
GPUs and heterogeneous computing
GPU computing
0.512021
Feluca: A Two-Stage Graph Coloring Algorithm With Color-Centric Paradigm on GPU · IEEE Trans. Parallel Distributed Syst. 2021
Graph algorithms and graph theory
graph coloring
0.512021
Feluca: A Two-Stage Graph Coloring Algorithm With Color-Centric Paradigm on GPU · IEEE Trans. Parallel Distributed Syst. 2021
Graph algorithms and graph theory › graph coloring
parallel graph coloring
0.512021
Feluca: A Two-Stage Graph Coloring Algorithm With Color-Centric Paradigm on GPU · IEEE Trans. Parallel Distributed Syst. 2021
Memory systems
data movement
0.212023
PMLiteDB: Streamlining Access Paths for High-Performance Persistent Memory Document Database Systems · IEEE Trans. Computers 2023
Memory systems
DRAM
0.212023
PMLiteDB: Streamlining Access Paths for High-Performance Persistent Memory Document Database Systems · IEEE Trans. Computers 2023
Parallel and multicore computing › parallel algorithms
parallel algorithm design
0.112021
Feluca: A Two-Stage Graph Coloring Algorithm With Color-Centric Paradigm on GPU · IEEE Trans. Parallel Distributed Syst. 2021

Methods — techniques the papers use, named apart from their topics

sequential spread · 1.0recursion-based coloring · 1.0color-centric paradigm · 1.0
YearPublicationVenuePosition
2023 PMLiteDB: Streamlining Access Paths for High-Performance Persistent Memory Document Database Systems
abstract
The advent of byte-addressable persistent memory opens an important opportunity for document databases to read and write durable data fetching them into DRAM. Reaping the benefit of persistent memory is not straightforward, as existing document databases are tailored for disk storage. They assume that the disk and DRAM data movement dominates the performance. However, this paper points out that data indexing becomes the performance bottleneck when porting document databases to persistent memory. The paper proposes PMLiteDB, the first persistent memory document database with streamlined access paths. PMLiteDB introduces two techniques,direct readingandselective caching.Direct readingstreamlines the translation from document IDs to the address of documents whenever possible by swizzling the IDs intopersistent memory references. It guarantees to use only up-to-datepersistent memory referenceswhen document movements invalidate associated references.Selective cachingreduces data movements between DRAM and persistent memory by selectively caching only frequently accessed persistent memory data pages with a DRAM buffer. For other pages, the database loads data on them directly without caching. Compared to the design that adopts persistent memory as a fast disk without exploiting the byte-addressability, PMLiteDB achieves 2.33× on average and up to 6.18× speedup.
Hai Jin 0001, Shuo Wei, Yan Sha, Chencheng Ye 0001, Haikun Liu, Xiaofei Liao
IEEE Trans. Computers2
2021 Feluca: A Two-Stage Graph Coloring Algorithm With Color-Centric Paradigm on GPU
abstract
There are great challenges in performing graph coloring on GPU in general. First, the long-tail problem exists in the recursion algorithm because the conflict (i.e., different threads assign the adjacent nodes to the same color) becomes more likely to occur as the number of iterations increases. Second, it is hard to parallelize the sequential spread algorithm because the color allocation depends on the adjoining iteration. Third, the atomic operation is widely used on GPU to maintain the color list, which can greatly reduce the efficiency of GPU threads. In this article, we propose a two-stage high-performance graph coloring algorithm, called Feluca, aiming to address the above challenges. Feluca combines the recursion-based method with the sequential spread-based method. In the first stage, Feluca uses a recursive routine to color a majority of vertices in the graph. Then, it switches to the sequential spread method to color the remaining vertices in order to avoid the conflicts of the recursive algorithm. Moreover, the following techniques are proposed to further improve the graph coloring performance. i) A new method is proposed to eliminate the cycles in the graph; ii) a top-down scheme is developed to avoid the atomic operation originally required for color selection; and iii) a novel color-centric coloring paradigm is designed to improve the degree of parallelism for the sequential spread part. All these newly developed techniques, together with further GPU-specific optimizations such as coalesced memory access, comprise an efficient parallel graph coloring solution in Feluca. We have conducted extensive experiments on NVIDIA GPU. The results show that Feluca can achieve 1.19 - 8.39× speedup over the state-of-the-art algorithms.
Zhigao Zheng 0001, Xuanhua Shi, Ligang He, Hai Jin 0001, Shuo Wei, Hulin Dai
IEEE Trans. Parallel Distributed Syst.5
2020 RSAN: Residual Subtraction and Attention Network for Single Image Super-Resolution
abstract
The single-image super-resolution (SISR) aims to recover a potential high-resolution image from its low-resolution version. Recently, deep learning-based methods have played a significant role in super-resolution field due to its effectiveness and efficiency. However, most of the SISR methods neglect the importance among the feature map channels. Moreover, they can not eliminate the redundant noises, making the output image be blurred. In this paper, we propose the residual subtraction and attention network (RSAN) for powerful feature expression and channels importance learning. More specifically, RSAN firstly implements one redundance removal module to learn noise information in the feature map and subtract noise through residual learning. Then it introduces the channel attention module to amplify high-frequency information and suppress the weight of effectless channels. Experimental results on extensive public benchmarks demonstrate our RSAN achieves significant improvement over the previous SISR methods in terms of both quantitative metrics and visual quality.
Shuo Wei, Xin Sun 0003, Junyu Dong
ICPR1
2011 Automatic image segmentation based on PCNN with adaptive threshold time constant
Shuo Wei, Qu Hong, Mengshu Hou
Neurocomputing1