Xueliang Du

dblp:165/4515 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
1since 2021 · last 2023
0009-0000-6368-0558ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Processor architecture and microarchitecture · 61% Energy-efficient computing · 39%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing › low-power design
low-power processor design
0.212016
MaPU: A novel mathematical computing architecture · HPCA 2016
Processor architecture and microarchitecture
SIMD
0.212016
MaPU: A novel mathematical computing architecture · HPCA 2016
Processor architecture and microarchitecture › SIMD
SIMD datapath
0.212016
MaPU: A novel mathematical computing architecture · HPCA 2016

Methods — techniques the papers use, named apart from their topics

state-machine-based program model · 0.2multi-granularity parallel memory · 0.2
YearPublicationVenuePosition
2023 Optimizing Memory Allocation for Multi-Subgraph Mapping on Spatial Accelerators
abstract
Spatial accelerators enable the pervasive use of energy-efficient solutions for computation-intensive applications. In the mapping of spatial accelerators, a large kernel is usually partitioned into multiple subgraphs for resource constraints, leading to more memory accesses and access conflicts. To minimize the access conflicts, existing works either neglect the interference of multiple subgraphs or pay little attention to data's life cycle along the execution order. To this end, this paper proposes an optimized memory allocation approach for multi-subgraph mapping on spatial accelerators by constructing an optimization problem using Integer Linear Programming (ILP). The experimental results demonstrate that our work can find conflict-free solutions for most kernels and achieve 1.15× speedup, as compared to the state-of-the-art approach.
Decai Pan, Dajiang Liu, Xueliang Du
SYSTOR5
2020 Baidu Kunlun An AI processor for diversified workloads
abstract
This article consists only of a collection of slides from the author's conference presentation.
Jian Ouyang, Mijung Noh, Yin Ma, Canghai Gu, SoonGon Kim, Ki-il Hong, Wang-Keun Bae, Zhibiao Zhao, Xiaozhang Gong, Jiaxin Shi, Hefei Zhu, Xueliang Du
Hot Chips Symposium16
2016 MaPU: A novel mathematical computing architecture
abstract
As the feature size of the semiconductor process is scaling down to 10nm and below, it is possible to assemble systems with high performance processors that can theoretically provide computational power of up to tens of PLOPS. However, the power consumption of these systems is also rocketing up to tens of millions watts, and the actual performance is only around 60% of the theoretical performance. Today, power efficiency and sustained performance have become the main foci of processor designers. Traditional computing architecture such as superscalar and GPGPU are proven to be power inefficient, and there is a big gap between the actual and peak performance. In this paper, we present the MaPU architecture, a novel architecture which is suitable for data-intensive computing with great power efficiency and sustained computation throughput. To achieve this goal, MaPU attempts to optimize the application from a system perspective, including the hardware, algorithm and corresponding program model. It uses an innovative multi-granularity parallel memory system with intrinsic shuffle ability, cascading pipelines with wide SIMD data paths and a state-machine-based program model. When executing typical signal processing algorithms, a single MaPU core implemented with a 40nm process exhibits a sustained performance of 134 GLOPS while consuming only 2.8 W in power, which increases the actual power efficiency by an order of magnitude comparable with the traditional CPU and GPGPU.
Xueliang Du, Leizu Yin, Weili Ren, Shaolin Xie, Zhonghua Pu, Guangxin Ding, Mengchen Zhu, Lipeng Yang, Ruoshan Guo, Yongyong Yang, Wenqin Sun, Fabiao Zhou, NuoZhou Xiao
HPCA2
2015 Design of a Distributed Compressor for Astronomy SSD
abstract
SSD (solid state device) has shown a great potential in astronomy data storage. Data compression is an essential task to obtain higher storage density and bandwidth. This paper proposes a distributed compressor customized for FPGA-based astronomy SSD. Our data-driven compressor cope with astronomy data in the unit of byte, two compression algorithms, run length and length-limited huffman are utilized, a distributed length-limited huffman encoder for SSD is further developed to reduce the latency. Experimental results indicate that our proposed compressor achieves a 1GB/s bandwidth with less than 2500 LUTs utilized while the compression ratio is only 10% lower than Gzip level9.
Xi Jin 0002, Xueliang Du
FCCM4