Zikun Li

dblp:273/0029 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
abstract
Modern large language model (LLM) applications exhibit diverse service-level objectives (SLOs), from low-latency requirements in interactive coding assistants to more relaxed constraints in data wrangling tasks. Existing LLM serving systems, which rely on uniform batching and scheduling strategies, often fail to meet these heterogeneous SLOs concurrently. We present AdaServe, the first LLM serving system designed to support efficient multi-SLO serving through SLO-customized speculative decoding. AdaServe formulates multi-SLO serving as a constrained optimization problem and introduces a hardware-aware algorithm that constructs a speculation tree tailored to each request's latency target. It features a speculate-select-verify pipeline that enables fine-grained control over decoding speed while maximizing system throughput. AdaServe further adapts to workload variation by dynamically adjusting speculation parameters. Evaluations across diverse workloads show that AdaServe reduces SLO violations by up to 4.3X and improves goodput by up to 1.9X compared to the best-performing baselines, highlighting its effectiveness in multi-SLO serving.
Zikun Li, Zhuofu Chen, Remi Delacourt, Gabriele Oliaro, Qinghan Chen, Shuhuai Lin, April Yang, Zhihao Zhang 0001, Zhuoming Chen, Yi-Hsiang Lai, Xinhao Cheng, Xupeng Miao
EuroSys1
2025 TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
abstract
Large language models (LLMs) have driven significant advancements across diverse NLP tasks, with long-context models gaining prominence for handling extended inputs. However, the expanding key-value (KV) cache size required by Transformer architectures intensifies the memory constraints, particularly during the decoding phase, creating a significant bottleneck. Existing sparse attention mechanisms designed to address this bottleneck have two limitations: (1) they often fail to reliably identify the most relevant tokens for attention, and (2) they overlook the spatial coherence of token selection across consecutive Transformer layers, which can lead to performance degradation and substantial overhead in token selection. This paper introduces TidalDecode, a simple yet effective algorithm and system for fast and accurate LLM decoding through position persistent sparse attention. TidalDecode leverages the spatial coherence of tokens selected by existing sparse attention methods and introduces a few token selection layers that perform full attention to identify the tokens with the highest attention scores, while all other layers perform sparse attention with the pre-selected tokens. This design enables TidalDecode to substantially reduce the overhead of token selection for sparse attention without sacrificing the quality of the generated results. Evaluation on a diverse set of LLMs and tasks shows that TidalDecode closely matches the generative performance of full attention methods while reducing the LLM decoding latency by up to $2.1\times$.
Lijie Yang 0003, Zhihao Zhang 0001, Zhuofu Chen, Zikun Li
ICLR4
2025 WirMAE: Learning Well-Logging Interval Representations via Masked Autoencoders for Gas Hydrate Reservoir Characterization
abstract
Reservoir characterization (identification and parameter estimation) is critical for gas hydrate exploration and development. While machine learning techniques excel at capturing complex relationships in well-logging reservoir characterization, existing research mainly focuses on end-to-end supervised learning approaches relying on costly labeled data and lacking multi-task learning capabilities. In this paper, we introduce a self-supervised learning (SSL) framework to learn general representations of large-scale unlabeled well-logging intervals from sandy, silty, and clayey hydrate reservoirs via masked autoencoders (WirMAE). By incorporating channel-based attention mechanisms and a masked reconstruction pre-training strategy for variable tokens, WirMAE effectively extracts intrinsic multivariate correlations within logging data, enabling the generation of various missing log curves. The model is validated on data from globally distributed hydrate with diverse accumulation patterns. Compared to supervised deep learning methods and classical machine learning models, Fine-turned WirMAE achieves superior reservoir identification accuracy (average F1-score: 0.864) using only 1% labeled data in complex geological settings. Combined with domain expertise, WirMAE yields precise estimation of key reservoir parameters such as hydrate saturation and permeability, outperforming conventional petrophysical methods. Additionally, embedding visualizations and attention analyses reveal the inner workings of the model and its consistency with expert-driven geological interpretations. Our findings highlight the potential of SSL for advancing more accurate and transparent intelligent reservoir characterization using well-log data, indicating that the application of WirMAE could be extended to broader hydrocarbon reservoirs in the future.
Zikun Li, Liangxiao Jiang, Jiaxin Sun, Kyungbook Lee, Fulong Ning
IEEE Trans. Geosci. Remote. Sens.1
2024 Quarl: A Learning-Based Quantum Circuit Optimizer
abstract
Optimizing quantum circuits is challenging due to the very large search space of functionally equivalent circuits and the necessity of applying transformations that temporarily decrease performance to achieve a final performance improvement. This paper presents Quarl, a learning-based quantum circuit optimizer. Applying reinforcement learning (RL) to quantum circuit optimization raises two main challenges: the large and varying action space and the non-uniform state representation. Quarl addresses these issues with a novel neural architecture and RL-training procedure. Our neural architecture decomposes the action space into two parts and leverages graph neural networks in its state representation, both of which are guided by the intuition that optimization decisions can be mostly guided by local reasoning while allowing global circuit-wide reasoning. Our evaluation shows that Quarl significantly outperforms existing circuit optimizers on almost all benchmark circuits. Surprisingly, Quarl can learn to perform rotation merging—a complex, non-local circuit optimization implemented as a separate pass in existing optimizers.
Zikun Li, Jinjun Peng, Yixuan Mei, Sina Lin, Yi Wu 0013, Oded Padon
Proc. ACM Program. Lang.1
2023 Missing Sonic Logs Generation for Gas Hydrate-Bearing Sediments via Hybrid Networks Combining Deep Learning With Rock Physics Modeling
abstract
Logging-while-drilling (LWD) sonic data is critical for marine gas hydrate reservoir evaluation and production prediction. However, acquiring complete acoustic logs, particularly shear-wave, poses significant challenges and incurs high costs. To tackle this issue, we develop a two-branch hybrid framework for predicting LWD sonic logs of hydrate-bearing sediments from existing logging data. One branch based on a rock physics model is utilized to generate background (no-hydrate/no-gas) elastic wave velocity profiles, while the other deep learning branch (DLB) compensates for the residuals between actual observations and the outputs of the first branch. The state-of-the-art Transformer encoder block is employed in the DLB to extract potentially intricate patterns within logging sequences. Such a scientific knowledge-guided network architecture with additional physics-based feature construction provides an explainable process that improves the physical consistency of predictions. Our method is tested with two publicly available datasets from the Cascadia continental margin. The hybrid model greatly enhances the predictive accuracy of the physical process model (with a minimum mean absolute percentage error of 0.73% and 4.33% for P- and S-wave velocities, respectively) and demonstrates outstanding generalization performance compared to pure data-driven approaches. The well-trained model offers impressive extrapolation beyond observed conditions for unmeasured high hydrate saturation (>40%) intervals in the region.
Zikun Li, Jialong Xia, Gang Lei 0003, Kyungbook Lee, Fulong Ning
IEEE Trans. Geosci. Remote. Sens.1
2023 BurstSketch: Finding Bursts in Data Streams
abstract
Burstis a common pattern in data streams which is characterized by a sudden increase in terms of arrival rate followed by a sudden decrease. Burst detection has attracted extensive attention from the research community. To detect bursts accurately in real time, we propose a novel sketch, namely BurstSketch, which consists of two stages. Stage 1 uses the technique Running Track to select potential burst items efficiently. Stage 2 monitors the potential burst items and captures the key features of burst pattern by a technique called Snapshotting. We further propose an optimization, namely Dynamic Buckets, which can improve the accuracy of BurstSketch. We provide theoretical error bounds for Stage 1, Stage 2 and the optimized version. Experimental results show that, compared with the strawman solution, Burstsketch achieves 2.00 to 11.63 times higher F1 score, and 1.56 times higher throughput. We also integrate BurstSketch into Apache Flink, and show that using BurstSketch can be faster than simply using the built-in APIs provided by Apache Flink.
Ruijie Miao, Jiarui Guo, Zikun Li, Tong Yang 0003, Bin Cui 0001
IEEE Trans. Knowl. Data Eng.4
2022 Quartz: superoptimization of Quantum circuits
abstract
Existing quantum compilers optimize quantum circuits by applying circuit transformations designed by experts. This approach requires significant manual effort to design and implement circuit transformations for different quantum devices, which use different gate sets, and can miss optimizations that are hard to find manually. We propose Quartz, a quantum circuit superoptimizer that automatically generates and verifies circuit transformations for arbitrary quantum gate sets. For a given gate set, Quartz generates candidate circuit transformations by systematically exploring small circuits and verifies the discovered transformations using an automated theorem prover. To optimize a quantum circuit, Quartz uses a cost-based backtracking search that applies the verified transformations to the circuit. Our evaluation on three popular gate sets shows that Quartz can effectively generate and verify transformations for different gate sets. The generated transformations cover manually designed transformations used by existing optimizers and also include new transformations. Quartz is therefore able to optimize a broad range of circuits for diverse gate sets, outperforming or matching the performance of hand-tuned circuit optimizers.
Mingkuan Xu, Zikun Li, Oded Padon, Sina Lin, Jessica Pointing, Auguste Hirth, Henry Ma, Jens Palsberg, Alex Aiken, Umut A. Acar
PLDI2
2021 BurstSketch: Finding Bursts in Data Streams
abstract
Burst is a common pattern in data streams which is characterized by a sudden increase in terms of arrival rate followed by a sudden decrease. Burst detection has attracted extensive attention from the research community. In this paper, we propose a novel sketch, namely BurstSketch, to detect bursts accurately in real time. BurstSketch first uses the technique Running Track to select potential burst items efficiently, and then monitors the potential burst items and capture the key features of burst pattern by a technique called Snapshotting. Experimental results show that our sketch achieves a 1.75 times higher recall rate than the strawman solution.
Shen Yan 0004, Zikun Li, Decheng Tan, Tong Yang 0003, Bin Cui 0001
SIGMOD Conference3
2020 WavingSketch: An Unbiased and Generic Sketch for Finding Top-k Items in Data Streams
abstract
Finding top-k items in data streams is a fundamental problem in data mining. Existing algorithms that can achieve unbiased estimation suffer from poor accuracy. In this paper, we propose a new sketch, WavingSketch, which is much more accurate than existing unbiased algorithms. WavingSketch is generic, and we show how it can be applied to four applications: finding top-k frequent items, finding top-k heavy changes, finding top-k persistent items, and finding top-k Super-Spreaders. We theoretically prove that WavingSketch can provide unbiased estimation, and then give an error bound of our algorithm. Our experimental results show that, compared with the state-of-the-art, WavingSketch has 4.50 times higher insertion speed and up to 9 x 106 times (2 x 104 times in average) lower error rate in finding frequent items when memory size is tight. For other applications, WavingSketch can also achieve up to 286 times lower error rate. All related codes are open-sourced and available at Github anonymously.
Jizhou Li, Zikun Li, Shiqi Jiang 0004, Tong Yang 0003, Bin Cui 0001, Yafei Dai, Gong Zhang 0001
KDD2