EDBT 2026 Demo / reviewers in the wild / expert
Shaolin Xie
dblp:178/3220
· DBLP profile ↗
8ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Processor architecture and microarchitecture · 90% Energy-efficient computing · 10% | |
| Databases, data mining, and information retrieval
1 paper |
Query processing and optimization · 56% Machine learning and data management · 44% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning and data management
learned database components |
0.9 | 1 | 2025 | LIMAO: A Framework for Lifelong Modular Learned Query Optimization · Proc. VLDB Endow. 2025 |
Query processing and optimization › query optimization
learned query optimization |
0.9 | 1 | 2025 | LIMAO: A Framework for Lifelong Modular Learned Query Optimization · Proc. VLDB Endow. 2025 |
Processor architecture and microarchitecture
instruction set architecture |
0.8 | 1 | 2024 | Scalable, Programmable and Dense: The HammerBlade Open-Source RISC-V Manycore · ISCA 2024 |
Processor architecture and microarchitecture
many-core architecture |
0.8 | 1 | 2024 | Scalable, Programmable and Dense: The HammerBlade Open-Source RISC-V Manycore · ISCA 2024 |
Processor architecture and microarchitecture › instruction set architecture
RISC-V |
0.8 | 1 | 2024 | Scalable, Programmable and Dense: The HammerBlade Open-Source RISC-V Manycore · ISCA 2024 |
Energy-efficient computing › low-power design
low-power processor design |
0.2 | 1 | 2016 | MaPU: A novel mathematical computing architecture · HPCA 2016 |
Processor architecture and microarchitecture
SIMD |
0.2 | 1 | 2016 | MaPU: A novel mathematical computing architecture · HPCA 2016 |
Processor architecture and microarchitecture › SIMD
SIMD datapath |
0.2 | 1 | 2016 | MaPU: A novel mathematical computing architecture · HPCA 2016 |
Methods — techniques the papers use, named apart from their topics
modular lifelong learning · 0.9attention-based neural network · 0.9state-machine-based program model · 0.2multi-granularity parallel memory · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | QuIT your B+-tree for the Quick Insertion Tree
Aneesh Raman, Konstantinos Karatsenidis, Shaolin Xie, Matthaios Olma, Subhadeep Sarkar 0001, Manos Athanassoulis |
EDBT | 3 |
| 2025 | LIMAO: A Framework for Lifelong Modular Learned Query OptimizationabstractQuery optimizers are crucial for the performance of database systems. Recently, many learned query optimizers (LQOs) have demonstrated significant performance improvements over traditional optimizers. However, most of them operate under a limited assumption: a static query environment. This limitation prevents them from effectively handling complex, dynamic query environments in real-world scenarios. Extensive retraining can lead to the well-known catastrophic forgetting problem which reduces the LQO generalizability over time. In this paper, we address this limitation and introduce LIMAO (Lifelong Modular Learned Query Optimizer), a framework for lifelong learning of plan cost prediction that can be seamlessly integrated into existing LQOs. LIMAO leverages a modular lifelong learning technique, an attention-based neural network composition architecture, and an efficient training paradigm designed to retain prior knowledge while continuously adapting to new environments. We implement LIMAO in two LQOs, showing that our approach is agnostic to underlying engines. Experimental results show that LIMAO significantly enhances the performance of LQOs, achieving up to a 40% improvement in query execution time and reducing the variance of execution time by up to 60% under dynamic workloads. By leveraging a precise and self-consistent design, LIMAO effectively mitigates catastrophic forgetting, ensuring stable and reliable plan quality over time. Compared to Postgres, LIMAO achieves up to a 4× speedup on selected benchmarks, highlighting its practical advantages in real-world query optimization. Qihan Zhang, Shaolin Xie, Ibrahim Sabek |
Proc. VLDB Endow. | 2 |
| 2024 | Scalable, Programmable and Dense: The HammerBlade Open-Source RISC-V ManycoreabstractExisting tiled manycore architectures propose to convert abundant silicon resources into general-purpose parallel processors with unmatched computational density and programmability. However, as we approach 100 K cores in one chip, conventional manycore architectures struggle to navigate three key axes: scalability, programmability, and density. Many manycores sacrifice programmability for density; or scalability for programmability. In this paper, we explore HammerBlade, which simultaneously achieves scalability, programmability and density. HammerBlade is a fully open-source RISC-V manycore architecture, which has been silicon-validated with a 2048-core ASIC implementation using a 14/16nm process. We evaluate the system using a suite of parallel benchmarks that captures a broad spectrum of computation and communication patterns. Dai Cheol Jung, Max Ruttenberg, Paul Gao 0001, Scott Davidson 0004, Daniel Ruelas-Petrisko, Kangli Li, Aditya K. Kamath, Shaolin Xie, Peitian Pan, Zhongyuan Zhao 0004, Zichao Yue, Bandhav Veluri, Sripathi Muralitharan, Adrian Sampson, Andrew Lumsdaine, Zhiru Zhang, Christopher Batten, Mark Oskin, Dustin Richmond, Michael B. Taylor |
ISCA | 9 |
| 2018 | FBNA: A Fully Binarized Neural Network AcceleratorabstractIn recent researches, binarized neural network (BNN) has been proposed to address the massive computations and large memory footprint problem of the convolutional neural network (CNN). Several works have designed specific BNN accelerators and showed very promising results. Nevertheless, only part of the neural network is binarized in their architecture and the benefits of binary operations were not fully exploited. In this work, we propose the first fully binarized convolutional neural network accelerator (FBNA) architecture, in which all convolutional operations are binarized and unified, even including the first layer and padding. The fully unified architecture provides more resource, parallelism and scalability optimization opportunities. Compared with the state-of-the-art BNN accelerator, our evaluation results show 3.1x performance, 5.4x resource efficiency and 4.9x power efficiency on CIFAR-10. Ruizhi Chen, Pin Li, Shaolin Xie |
FPL | 5 |
| 2018 | Low Latency Spiking ConvNets with Restricted Output Training and False Spike InhibitionabstractDeep convolutional neural networks (ConvNets) have achieved the state-of-the-art performance on many real-world applications. However, significant computation and storage demands are required by ConvNets. Spiking neural networks (SNNs), with sparsely activated neurons and event-driven computations, show great potential to take advantage of the ultra- low power spike-based hardware architectures. Yet, training SNN with similar accuracy as ConvNets is difficult. Recent researchers have demonstrated the work of converting ConvNets to SNNs (CNN-SNN conversion) with similar accuracy. However, the energy-efficiency of the converted SNNs is impaired by the increased classification latency. In this paper, we focus on optimizing the classification latency of the converted SNNs. First, we propose a restricted output training method to normalize the converted weights dynamically in the CNN-SNN training phase. Second, false spikes are identified and the false spike inhibition theory is derived to speedup the convergence of the classification process. Third, we propose a temporal max pooling method to approximate the max pooling operation in ConvNets without accuracy loss. The evaluation shows that the converted SNNs converge in about 30 time-steps and achieve the best classification accuracy of 94% on CIFAR -10 dataset. Ruizhi Chen, Shaolin Xie, Pin Li |
IJCNN | 4 |
| 2018 | Fast and Efficient Deep Sparse Multi-Strength Spiking Neural Networks with Dynamic PruningabstractDeep convolutional neural networks (CNNs) have shown state-of-the-art accuracy for various computer vision and speech tasks. However, CNNs are computation-intensive and energy-inefficient which are difficult to be deployed in real-time systems. Event-driven Spiking Neural Networks (SNNs) are extremely power efficient, which provides an alternative for ultra-low power applications. But effective training methods for SNN are still lacking. Due to its spatio-temporal feature of SNN, conventional training method for CNN can not be employed in SNN. To address this problem, some researchers proposed to convert the corresponding weights of trained CNNs into the synapse weights of SNNs (CNNs-SNNs). Nevertheless, limited by the [0, 1] constraints on the SNN neuron outputs, the accuracy of the converted SNNs is impaired. Besides, as the SNN network becomes deeper, the convergence speed of SNN inference are unacceptably slow. In this work, we proposed an innovative deep multi-strength SNN (M-SNN) structure which relaxes the restriction of the neuron output spike strength while the event-driven feature for low-power implementations is maintained. Using this architecture, large scale SNN can be converted from CNN with comparable accuracy and fast inference speed. The evaluation results show 3.7 × convergence speedup. Moreover, with multi-strength spike, aggressive pruning strategies can be applied to reduce the computational operations by almost 85% while maintaining the same accuracy. Ruizhi Chen, Shaolin Xie, Pin Li |
IJCNN | 3 |
| 2018 | Parallel Polar Encoding in 5G CommunicationabstractBecause of its theoretical capacity-achieving property, polar code has become the coding scheme of the control channel in the 5G communication standard. Although its encoding complexity is low, the data dependency in polar code makes it difficult to parallelize. This paper proposes a parallel polar encoding method for 5G communication and evaluates its performance with extended digital signal processor (DSP) instructions. Compared with the existing field-programmable gate array (FPGA) implementation, the performance improved by $300 \times$ with negligible area and power overhead. The extended instructions are based on our in-house DSP architecture, but the parallel scheme is applicable to other single instruction multiple data (SIMD) architectures. Shaolin Xie |
ISCC | 2 |
| 2016 | MaPU: A novel mathematical computing architectureabstractAs the feature size of the semiconductor process is scaling down to 10nm and below, it is possible to assemble systems with high performance processors that can theoretically provide computational power of up to tens of PLOPS. However, the power consumption of these systems is also rocketing up to tens of millions watts, and the actual performance is only around 60% of the theoretical performance. Today, power efficiency and sustained performance have become the main foci of processor designers. Traditional computing architecture such as superscalar and GPGPU are proven to be power inefficient, and there is a big gap between the actual and peak performance. In this paper, we present the MaPU architecture, a novel architecture which is suitable for data-intensive computing with great power efficiency and sustained computation throughput. To achieve this goal, MaPU attempts to optimize the application from a system perspective, including the hardware, algorithm and corresponding program model. It uses an innovative multi-granularity parallel memory system with intrinsic shuffle ability, cascading pipelines with wide SIMD data paths and a state-machine-based program model. When executing typical signal processing algorithms, a single MaPU core implemented with a 40nm process exhibits a sustained performance of 134 GLOPS while consuming only 2.8 W in power, which increases the actual power efficiency by an order of magnitude comparable with the traditional CPU and GPGPU. Xueliang Du, Leizu Yin, Weili Ren, Shaolin Xie, Zhonghua Pu, Guangxin Ding, Mengchen Zhu, Lipeng Yang, Ruoshan Guo, Yongyong Yang, Wenqin Sun, Fabiao Zhou, NuoZhou Xiao |
HPCA | 9 |