EDBT 2026 Demo / reviewers in the wild / expert
Jiesong Liu
dblp:337/2891
· DBLP profile ↗
11ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0002-8311-020XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Drop-In Solution for On-the-Fly Adaptation of Speculative Decoding in Large Language ModelsabstractLarge Language Models (LLMs) are cuttingedge generative AI models built on transformer architecture, which tend to be highly memoryintensive when performing real-time inference.Various strategies have been developed to enhance the end-to-end inference speed for LLMs, one of which is speculative decoding.This technique involves running a smaller LLM (draft model) for inference over a defined window size, denoted as γ, while simultaneously being validated by the larger LLM (target model).Choosing the optimal γ value and the draft model is essential for unlocking the potential of speculative decoding.But it is difficult to do due to the complicated influence from various factors, including the nature of the task, the hardware in use, and the combination of the large and small models.This paper introduces on-the-fly adaption of speculative decoding, a solution that dynamically adapts the choices to maximize the efficiency of speculative decoding for LLM inferences.As a drop-in solution, it needs no offline benchmarking or training.Experiments show that the solution can lead to 3.55-16.48%speed improvement over the standard speculative decoding, and 1.2-3.4×over the default LLMs. Jiesong Liu, Brian Park, Xipeng Shen |
ACL (1) | 1 |
| 2025 | Generalizing Reuse Patterns for Efficient DNN on MicrocontrollersabstractDeep Neural Networks (DNNs) face challenges in deployment on resource-constrained devices due to their high computational demands. Leveraging redundancy in input data and activation maps for computation reuse is an effective way to accelerate DNN inference, especially for microcontrollers where the computing power is very limited. This work points out an important limitation in current reuse-based DNN optimizations, the narrow definition of reuse patterns in data. It proposes the concept of generalized reuse and uncovers the relations between generalized reuse patterns and row/column reorder of a matrix view of the input or activation map of a DNN. It revolutionizes the conventional view of explorable reuse patterns, drastically expanding the reuse space. It further develops two novel analytical models for analyzing the impacts of reuse patterns on the accuracy and latency of DNNs, enabling efficient selection of appropriate reuse patterns. Experiments show that generalized reuse consistently brings significant benefits, regardless of the differences among DNNs or microcontroller hardware. It delivers 1.03-2.2x speedups or 1-8% accuracy improvement over conventional reuse. Jiesong Liu, Bin Ren 0002, Xipeng Shen |
ASPLOS (2) | 1 |
| 2025 | Fourier Token Merging: Understanding and Capitalizing Frequency Domain for Efficient Image GenerationabstractImage generation requires intensive computations and faces challenges due to long latency.
Exploiting redundancy in the input images and intermediate representations throughout the neural network pipeline is an effective way to accelerate image generation.
Token merging (ToMe) exploits similarities among input tokens by clustering them and merges similar tokens into one, thus significantly reducing the number of tokens that are fed into the transformer block.
This work introduces Fourier Token Merging, a new method for understanding and capitalizing frequency domain for efficient image generation.
By introducing frequency token merging, we find that transforming the token into the frequency domain representation for clustering can better exert the ability of clustering based on the underlying redundancy after de-correlation.
Through analytical and empirical studies, we demonstrate the benefits of using Fourier clustering over the original time domain clustering.
We experimented fourier token merging on the stable diffusion model, and the results show up to 25\% reduction in latency without impairing image quality.
The code is available at https://github.com/Fred1031/Fourier-Token-Merging. Jiesong Liu, Xipeng Shen |
NeurIPS | 1 |
| 2025 | A Systematic Study on Early Stopping Metrics in HPO and the Implications of UncertaintyabstractThe development of hyperparameter optimization (HPO) algorithms is an important topic within both the machine learning and data management domains. While numerous strategies employing early stopping mechanisms have been proposed to bolster HPO efficiency, there remains a notable deficiency in understanding how the selection of early stopping metrics influences the reliability of early stopping decisions and, by extension, the broader HPO outcomes. This paper undertakes a systematic exploration of the impact of metric selection on the effectiveness of early stopping-based HPO. Specifically, we introduce a set of metrics that incorporate uncertainty and highlight their practical significance in enhancing the reliability of early stopping decisions. Our empirical experiments on HPO and NAS benchmarks show that using training loss as an early stopping metric in the early training stages improves HPO outcomes by up to 24.76% compared to the more widely accepted validation loss. Furthermore, integrating uncertainty into the metric yields an additional improvement of up to 4% under budget constraints, translating into meaningful resource savings and scalability benefits in large-scale HPO scenarios. These findings demonstrate the critical role of metric selection while shedding light on the potential implications of integrating uncertainty as a metric. This research provides empirical insights that serve as a compass for the selection and formulation of metrics, thereby contributing to a more profound comprehension of mechanisms underpinning early stopping-based HPO. Jiawei Guan, Feng Zhang 0007, Jiesong Liu, Xiaoyong Du 0001, Xipeng Shen |
Proc. VLDB Endow. | 3 |
| 2024 | UQ-Guided Hyperparameter Optimization for Iterative LearnersabstractHyperparameter Optimization (HPO) plays a pivotal role in unleashing the potential of iterative machine learning models. This paper addresses a crucial aspect that has largely been overlooked in HPO: the impact of uncertainty in ML model training. The paper introduces the concept of uncertainty-aware HPO and presents a novel approach called the UQ-guided scheme for quantifying uncertainty. This scheme offers a principled and versatile method to empower HPO techniques in handling model uncertainty during their exploration of the candidate space.
By constructing a probabilistic model and implementing probability-driven candidate selection and budget allocation, this approach enhances the quality of the resulting model hyperparameters. It achieves a notable performance improvement of over 50\% in terms of accuracy regret and exploration time. Jiesong Liu, Feng Zhang 0007, Jiawei Guan, Xipeng Shen |
NeurIPS | 1 |
| 2024 | Enabling Efficient Deep Learning on MCU With Transient Redundancy EliminationabstractDeploying deep neural networks (DNNs) with satisfactory performance in resource-constrained environments is challenging. This is especially true of microcontrollers due to their tight space and computational capabilities. However, there is a growing demand for DNNs on microcontrollers, as executing large DNNs on microcontrollers is critical to reducing energy consumption, increasing performance efficiency, and eliminating privacy concerns. This paper presents a novel and systematic data redundancy elimination method to implement efficient DNNs on microcontrollers through innovations in computation and space optimization. By making the optimization itself a trainable component in the target neural networks, this method maximizes performance benefits while keeping the DNN accuracy stable. Experiments are performed on two microcontroller boards with three popular DNNs, namely CifarNet, ZfNet and SqueezeNet. Experiments show that this solution eliminates more than 96% of computations in DNNs and makes them fit well on microcontrollers, yielding 3.4-5$\times$speedup with little loss of accuracy. Jiesong Liu, Feng Zhang 0007, Jiawei Guan, Hsin-Hsuan Sung, Xiaoguang Guo, Saiqin Long, Xiaoyong Du 0001, Xipeng Shen |
IEEE Trans. Computers | 1 |
| 2024 | G-Learned Index: Enabling Efficient Learned Index on GPUabstractAI and GPU technologies have been widely applied to solve big data problems. The total data volume worldwide reaches 200 zettabytes in 2022. How to efficiently index the required content among massive data becomes serious. Recently, a promising learned index has been proposed to address this challenge: It has extremely high efficiency while retaining marginal space overhead. However, we notice that previous learned indexes have mainly focused on CPU architecture, while ignoring the advantages of GPU. Because traditional indexes like B-Tree, LSM, and bitmap have greatly benefited from GPU acceleration, a combination of a learned index and GPU has great potentials to reach tremendous speedups. In this paper, we propose a GPU-based learned index, called G-Learned Index, to significantly improve the performance of learned index structures. The primary challenges in developing G-Learned Index lie in the use of thousands of GPU cores including minimization of synchronization and branch divergence, data structure design for parallel operations, and usage of memory bandwidth including limited memory transactions and multi-memory hierarchy. To overcome these challenges, a series of novel technologies are developed, including efficient thread organization, succinct data structures, and heterogeneous memory hierarchy utilization. Compared to the state-of-the-art learned index, the proposed G-Learned Index achieves an average of 174× speedup (and 107× of its parallel version). Meanwhile, we attain 2× less query time over the state-of-the-art GPU B-Tree. Our further exploration of range queries shows that G-Learned Index is 17× faster than CPU multi-dimensional learned index. We have made G-Learned Index available athttps://anonymous.4open.science/r/G-Learned-Index-8D89. Jiesong Liu, Feng Zhang 0007, Lv Lu, Chang Qi, Xiaoguang Guo, Dong Deng 0001, Guoliang Li 0001, Huanchen Zhang, Jidong Zhai, Hechen Zhang, Yuxing Chen 0003, Anqun Pan, Xiaoyong Du 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | Space-Efficient TREC for Enabling Deep Learning on MicrocontrollersabstractDeploying deep neural networks (DNNs) for a resource-constrained environment and achieving satisfactory performance is challenging. It is especially so on microcontrollers for their stringent space and computing power. This paper focuses on new ways to make TREC, an optimization recently proposed to enable computation reuse in DNNs, space and time efficient on Microcontrollers. The solution maximizes the performance benefits while keeping the DNN accuracy stable. Experiments show that the solution eliminates over 96% computations in DNNs and makes them fit well into microcontrollers, producing 3.4-5× speedups with only marginal accuracy loss. Jiesong Liu, Feng Zhang 0007, Jiawei Guan, Hsin-Hsuan Sung, Xiaoguang Guo, Xiaoyong Du 0001, Xipeng Shen |
ASPLOS (3) | 1 |
| 2022 | TREC: Transient Redundancy Elimination-based ConvolutionabstractThe intensive computations in convolutional neural networks (CNNs) pose challenges for resource-constrained devices; eliminating redundant computations from convolution is essential. This paper gives a principled method to detect and avoid transient redundancy, a type of redundancy existing in input data or activation maps and hence changing across inferences. By introducing a new form of convolution (TREC), this new method makes transient redundancy detection and avoidance an inherent part of the CNN architecture, and the determination of the best configurations for redundancy elimination part of CNN backward propagation. We provide a rigorous proof of the robustness and convergence of TREC-equipped CNNs. TREC removes over 96% computations and achieves 3.51x average speedups on microcontrollers with minimal (about 0.7%) accuracy loss. Jiawei Guan, Feng Zhang 0007, Jiesong Liu, Hsin-Hsuan Sung, Xiaoyong Du 0001, Xipeng Shen |
NeurIPS | 3 |
| 2022 | Approximating Probabilistic Group Steiner Trees in GraphsabstractConsider an edge-weighted graph, and a number of properties of interests (PoIs). Each vertex has a probability of exhibiting each PoI. The joint probability that a set of vertices exhibits a PoI is the probability that this set contains at least one vertex that exhibits this PoI. The probabilistic group Steiner tree problem is to find a tree such that (i) for each PoI, the joint probability that the set of vertices in this tree exhibits this PoI is no smaller than a threshold value, e.g. , 0.97; and (ii) the total weight of edges in this tree is the minimum. Solving this problem is useful for mining various graphs with uncertain vertex properties, but is NP-hard. The existing work focuses on certain cases, and cannot perform this task. To meet this challenge, we propose 3 approximation algorithms for solving the above problem. Let |Γ| be the number of PoIs, and ξ be an upper bound of the number of vertices for satisfying the threshold value of exhibiting each PoI. Algorithms 1 and 2 have tight approximation guarantees proportional to |Γ| and ξ, and exponential time complexities with respect to ξ and |Γ|, respectively. In comparison, Algorithm 3 has a looser approximation guarantee proportional to, and a polynomial time complexity with respect to, both |Γ| and ξ. Experiments on real and large datasets show that the proposed algorithms considerably outperform the state-of-the-art related work for finding probabilistic group Steiner trees in various cases. Yahui Sun 0001, Jiesong Liu, Xiaokui Xiao, Rong-Hua Li 0001, Zhewei Wei |
Proc. VLDB Endow. | 3 |
| 2022 | Exploring Query Processing on CPU-GPU Integrated Edge DeviceabstractHuge amounts of data have been generated on edge devices every day, which requires efficient data analytics and management. However, due to the limited computing capacity of these edge devices, query processing at the edge faces tremendous pressure. Fortunately, in recent years, hardware vendors have integrated heterogeneous coprocessors, such as GPUs, into the edge device, which can provide much more computing power. Furthermore, the CPU-GPU integrated edge device has shown significant benefits in a variety of situations. Therefore, the exploration of query processing on such CPU-GPU integrated edge devices becomes an urgent need. In this article, we develop a fine-grained query processing engine, called FineQuery, which can perform efficient query processing on CPU-GPU integrated edge devices. Particularly, FineQuery can take advantage of both architectural features of edge devices and query characteristics by performing fine-grained workload scheduling between the CPU and the GPU. Experiments show that on TPC-H workloads, FineQuery reduces 42.81% latency and improves 2.39× bandwidth utilization on average compared to the implementation of using only GPU or CPU. Furthermore, query processing at the edge can bring significant performance-per-cost benefits and energy efficiency. On average, FineQuery at the edge brings 21× performance-per-cost ratio and 4× energy efficiency compared with processing the data on a discrete GPU platform. Jiesong Liu, Feng Zhang 0007, Hourun Li, Dalin Wang, Weitao Wan, Xiaokun Fang, Jidong Zhai, Xiaoyong Du 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |