EDBT 2026 Demo / reviewers in the wild / expert
Xueliang Du
dblp:165/4515
· DBLP profile ↗
4ranked-venue papers
0as first author
1since 2021 · last 2023
0009-0000-6368-0558ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Processor architecture and microarchitecture · 61% Energy-efficient computing · 39% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy-efficient computing › low-power design
low-power processor design |
0.2 | 1 | 2016 | MaPU: A novel mathematical computing architecture · HPCA 2016 |
Processor architecture and microarchitecture
SIMD |
0.2 | 1 | 2016 | MaPU: A novel mathematical computing architecture · HPCA 2016 |
Processor architecture and microarchitecture › SIMD
SIMD datapath |
0.2 | 1 | 2016 | MaPU: A novel mathematical computing architecture · HPCA 2016 |
Methods — techniques the papers use, named apart from their topics
state-machine-based program model · 0.2multi-granularity parallel memory · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Optimizing Memory Allocation for Multi-Subgraph Mapping on Spatial AcceleratorsabstractSpatial accelerators enable the pervasive use of energy-efficient solutions for computation-intensive applications. In the mapping of spatial accelerators, a large kernel is usually partitioned into multiple subgraphs for resource constraints, leading to more memory accesses and access conflicts. To minimize the access conflicts, existing works either neglect the interference of multiple subgraphs or pay little attention to data's life cycle along the execution order. To this end, this paper proposes an optimized memory allocation approach for multi-subgraph mapping on spatial accelerators by constructing an optimization problem using Integer Linear Programming (ILP). The experimental results demonstrate that our work can find conflict-free solutions for most kernels and achieve 1.15× speedup, as compared to the state-of-the-art approach. Decai Pan, Dajiang Liu, Xueliang Du |
SYSTOR | 5 |
| 2020 | Baidu Kunlun An AI processor for diversified workloadsabstractThis article consists only of a collection of slides from the author's conference presentation. Jian Ouyang, Mijung Noh, Yin Ma, Canghai Gu, SoonGon Kim, Ki-il Hong, Wang-Keun Bae, Zhibiao Zhao, Xiaozhang Gong, Jiaxin Shi, Hefei Zhu, Xueliang Du |
Hot Chips Symposium | 16 |
| 2016 | MaPU: A novel mathematical computing architectureabstractAs the feature size of the semiconductor process is scaling down to 10nm and below, it is possible to assemble systems with high performance processors that can theoretically provide computational power of up to tens of PLOPS. However, the power consumption of these systems is also rocketing up to tens of millions watts, and the actual performance is only around 60% of the theoretical performance. Today, power efficiency and sustained performance have become the main foci of processor designers. Traditional computing architecture such as superscalar and GPGPU are proven to be power inefficient, and there is a big gap between the actual and peak performance. In this paper, we present the MaPU architecture, a novel architecture which is suitable for data-intensive computing with great power efficiency and sustained computation throughput. To achieve this goal, MaPU attempts to optimize the application from a system perspective, including the hardware, algorithm and corresponding program model. It uses an innovative multi-granularity parallel memory system with intrinsic shuffle ability, cascading pipelines with wide SIMD data paths and a state-machine-based program model. When executing typical signal processing algorithms, a single MaPU core implemented with a 40nm process exhibits a sustained performance of 134 GLOPS while consuming only 2.8 W in power, which increases the actual power efficiency by an order of magnitude comparable with the traditional CPU and GPGPU. Xueliang Du, Leizu Yin, Weili Ren, Shaolin Xie, Zhonghua Pu, Guangxin Ding, Mengchen Zhu, Lipeng Yang, Ruoshan Guo, Yongyong Yang, Wenqin Sun, Fabiao Zhou, NuoZhou Xiao |
HPCA | 2 |
| 2015 | Design of a Distributed Compressor for Astronomy SSDabstractSSD (solid state device) has shown a great potential in astronomy data storage. Data compression is an essential task to obtain higher storage density and bandwidth. This paper proposes a distributed compressor customized for FPGA-based astronomy SSD. Our data-driven compressor cope with astronomy data in the unit of byte, two compression algorithms, run length and length-limited huffman are utilized, a distributed length-limited huffman encoder for SSD is further developed to reduce the latency. Experimental results indicate that our proposed compressor achieves a 1GB/s bandwidth with less than 2500 LUTs utilized while the compression ratio is only 10% lower than Gzip level9. Xi Jin 0002, Xueliang Du |
FCCM | 4 |