Wei Li 0322

dblp:64/6025-322 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0005-7862-9989ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RefineDedup: efficient deduplication for mobile systems via application-wise learning
Wei Li 0322, Xianzhang Chen, Xingjie Zhou, Duo Liu 0002, Yujuan Tan, Ao Ren, Kan Zhong, Lei Qiao 0002
Sci. China Inf. Sci.1
2025 CoSF: A Co-Optimization Framework for Operator Splitting and Fusion
Wei Li 0322, Ao Ren, Qingqiu Lan, Haining Fang, Zhenyu Wang 0002, Yujuan Tan, Kan Zhong, Duo Liu 0002
Euro-Par (1)1
2025 CAST: An Efficient Framework for Schedules Performance Prediction Based on Compact ASTs
abstract
With the advances of deep learning, efficient model inference is crucial. Deep learning compilers optimize inference by decomposing models into subgraphs and searching schedules for them, whose evaluation relies on accurate cost models. Existing methods suffer from high transformation overheads or limited prediction accuracy caused by insufficient structural representation of subgraphs and schedules. To address these limitations, we propose CAST, a framework that predicts schedule performance based on Abstract Syntax Trees (ASTs). CAST proposes AST classification based on structural similarity and class-specific cost models. Experiments show CAST achieves significantly reduced prediction errors and up to$13 \times$higher efficiency than prior methods.
Qingqiu Lan, Ao Ren, Zhenyu Wang 0002, Wei Li 0322, Hongbin Zhu, Yujuan Tan, Duo Liu 0002, Kan Zhong, Chaoxia Qin
ICCD4
2025 FASP: A Fast and Accurate Framework for Schedule Performance Evaluation
abstract
With the widespread application of deep neural networks, improving inference efficiency has become increasingly critical. To speed up the inference, deep learning compilers search for high-performance schedules for the DNN tensor programs. During the process, cost models have been extensively studied to evaluate the performance of the schedules, such that high-performance ones can be efficiently obtained. However, existing methods suffer from either high overhead or low accuracy of performance evaluation, both of which limit the efficiency of the final schedule. To address these issues, we propose FASP, a fast and accurate framework for schedule performance evaluation, based on Abstract Syntax Trees (ASTs). First, we propose a redundancy-aware ASTs reduction method to generate compact ASTs for more accurate feature extraction. Second, we propose a feature extraction method based on compact ASTs, which extracts features by accounting for computation nodes, loop nodes, and their structural relationships. Third, we propose a composition-similarity-driven ASTs classification method and a class-specific cost model architecture for more accurate performance evaluation. FASP overcomes the limitations of prior methods by significantly reducing evaluation errors. Experiments show its excellent performance in both single-model and cross-model evaluation, with errors ranging from 6 % to$\mathbf{1 3 \%}$. Moreover, FASP can obtain high-performance schedules with$13 \times$lower latency.
Qingqiu Lan, Ao Ren, Zhenyu Wang 0002, Wei Li 0322, Hongbin Zhu, Yujuan Tan, Duo Liu 0002, Kan Zhong, Chaoxia Qin
ICPADS4
2024 FinerDedup: Sifting Fingerprints for Efficient Data Deduplication on Mobile Devices
abstract
Data deduplication is promised to extend the lifetime and capacity of storage on mobile devices. However, existing data deduplication works show high memory consumption and indexing costs for maintaining a fingerprint for each data block, especially when the duplicate ratio of data blocks on mobile systems is about 10% to 30%. In this paper, we propose a novel approach called FinerDedup to optimize the memory costs and retrieval efficiency of data deduplication. FinerDedup drastically reduces the number of fingerprints by screening out the duplicate data blocks via random forest and Bloom filter. We implement FinerDedup on real mobile devices with Android 10 and evaluate it with real workloads. Extensive experimental results show that FinerDedup can reduce 85% of fingerprints and 20% of I/O latency over the widely-used DmDedup.
Xianzhang Chen, Xingjie Zhou, Wei Li 0322, Duo Liu 0002, Yujuan Tan, Ao Ren
DAC3