VLDB 2026 Research / reviewers in the wild / expert
Xiaoyang Lu
dblp:80/8488
· DBLP profile ↗
20ranked-venue papers
8as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 7 first-author · 12 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | I/O Analysis is All You Need: An I/O Analysis for Long-Sequence AttentionabstractAs GPUs and other accelerators become increasingly popular, optimizing I/O operations between on-chip and off-chip memory is increasingly critical. I/O analysis, however, is complex, requiring a deep understanding of application dataflow and memory hierarchy. Developing a practical I/O analysis methodology remains a timely challenge. Self-attention is employed extensively in transformer models, but its quadratic memory complexity poses significant challenges to modern memory systems. In this study, we explore how to use I/O analysis to develop optimal solutions for accelerating exact long-sequence self-attention. We first introduce a novel I/O analysis for tall-and-skinny matrix-matrix multiplication, which captures the dominant data movement behavior of long-sequence self-attention. Guided by systematic I/O analysis, we develop AttenIO, an I/O-driven accelerator for exact long-sequence self-attention with three key optimizations: (1) an analytically derived I/O-optimal tiling and scheduling to minimize I/O operations, (2) fine-grained three-level communication-computation overlapping to hide I/O stalls, and (3) parallel execution patterns for efficient softmax. Our evaluation shows that AttenIO achieves a 1.6×-8.8× speedup over the state-of-the-art solutions. Although AttenIO is designed for self-attention, it also highlights the broader potential of I/O analysis as a principled foundation for guiding high-performance I/O optimizations. Xiaoyang Lu, Boyu Long, Xiaoming Chen 0003, Yinhe Han 0001, Xian-He Sun |
ASPLOS (2) | 1 |
| 2026 | Zion: A Comprehensive, Adaptive, and Lightweight Hardware PrefetcherabstractAs the gap between processor and memory performance widens, optimizing data access performance becomes increasingly critical. Hardware prefetching is a widely used technique to hide long-latency off-chip memory accesses, but state-of-the-art prefetchers struggle with diverse and dynamic access patterns. Their limited adaptability leads to excessive storage overhead and reduced effectiveness under memory-intensive workloads. We propose Zion, a comprehensive, adaptive, and lightweight hardware prefetcher for memory-intensive workloads. At its core, Zion uses Independent Temporal-Spatial Modules (ITSM) for broad pattern coverage and runtime adaptability to diverse memory access patterns. Moreover, Zion leverages runtime feedback to dynamically guide prefetching decisions and maintain efficiency under memory pressure. Extensive multi-core evaluations show that Zion consistently outperforms state-of-the-art prefetchers, achieving up to 43.2% performance improvement on SPEC and 43.0% on self-attention workloads, while maintaining low overhead and broad effectiveness. Vadim Biryukov, Xiaoyang Lu, Zirui Liu 0001, Kaixiong Zhou, Xian-He Sun |
DATE | 2 |
| 2026 | I/O-Aware PIM Acceleration for Long-Sequence LLM Inference with Hybrid Sparse Attention
Xiaoyang Lu, Lihan Hu, Hongrui Huang, Peng Jiang 0004, Xian-He Sun |
IPDPS | 1 |
| 2026 | QCP: A Practical Separation Logic-Based C Program Verification Tool
Xiwei Wu, Yueyang Feng, Xiaoyang Lu, Tianchuan Lin, Shushu Wu, Lihan Xie, Chengxi Yang, Hongyi Zhong, Juanru Li, Naijun Zhan, Zhenjiang Hu 0002, Qinxiang Cao |
TASE | 3 |
| 2026 | Bayesian Tensor Interpolative Decomposition for IoT-Based Healthcare Data Multiway AnalyticsabstractThe IoT-based Healthcare Industry 5.0 offers intelligent solutions that utilize numerous sensors, generating massive amounts of data. These healthcare data with composite properties can be represented by high-order tensors, which enables the application of tensor multi-way analytics to uncover the latent features in data. However, existing tensor decomposition techniques often face challenges of limited interpretability and weak robustness, which constrain their application in the healthcare industry. To mitigate these challenges, this paper introduces a novel Bayesian Tensor Interpolative Decomposition (BTID) model, which is an extension of the concept of matrix interpolative decomposition and is designed to efficiently and interpretably analyze multi-way tensors. The proposed model performs low-rank approximation using a subset of the original data and employs the Bayesian learning approach to infer the weight matrices, thereby ensuring theoretical interpretability. Additionally, a hierarchical prior is incorporated within this model to address the complexities of the data and enhance model robustness. Finally, a clustering-driven strategy is developed to construct the skeleton matrix, ensuring its physical interpretability while facilitating the CP rank estimation. The BTID model has been comprehensively evaluated on synthetic and healthcare datasets, where the experimental outcomes confirm its superior accuracy, robustness, and interpretability compared with traditional tensor decomposition methods. Zecan Yang, Huaimin Wang 0002, Laurence T. Yang, Honglu Zhao, Songhe Yuan, Xiaoyang Lu |
IEEE Internet Things J. | 6 |
| 2025 | Concurrency-Aware Cache Miss Cost Prediction with Perceptron Learning
Xiaoyang Lu, Xiaoming Chen 0003, Yinhe Han 0001, Xian-He Sun |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | COSMOS: RL-Enhanced Locality-Aware Counter Cache Optimization for Secure MemoryabstractSecure memory systems employing AES-CTR encryption face significant performance challenges due to high counter (CTR) cache miss rates, especially in applications with irregular memory access patterns.These high miss rates increase memory traffic and latency, as each CTR cache miss triggers additional DRAM accesses.To address these bottlenecks and adapt to diverse access patterns, we propose COSMOS (Counter Optimized Secure Memory Operation Scheme), a novel solution leveraging reinforcement learning to reduce long memory access latency.COSMOS integrates two RL-based specialized predictors: one for data location prediction and another for CTR locality prediction, each with a well-defined state space, action space, and reward function.The RL-based data location predictor determines whether data reside on-chip or offchip after an L1 cache miss, enabling early CTR access for off-chip predictions with minimal changes to the existing cache hierarchy.The RL-based CTR locality predictor identifies CTRs with high locality, supporting a locality-centric CTR cache (LCR-CTR) to improve cache efficiency and reduce miss rates.COSMOS improves performance over MorphCtr by 25% in for irregular memory access applications, with minimal hardware overhead. Xiaoyang Lu, Yuezhi Che, Ziang Tian, Dazhao Cheng, Xian-He Sun, Michael T. Niemier, Xiaobo Sharon Hu |
MICRO | 2 |
| 2025 | VEP: A Two-stage Verification Toolchain for Full eBPF Programmability
Xiwei Wu, Yueyang Feng, Tianyi Huang, Xiaoyang Lu, Shengkai Lin, Lihan Xie, Shizhen Zhao, Qinxiang Cao |
NSDI | 4 |
| 2025 | ProMiner: Enhancing Locality, Parallelism, and Offloading for Graph Mining on Processing-in-Memory SystemsabstractGraph mining, critical for discovering specific patterns within complex structures, is becoming increasingly important in our data-driven world. Due to their memory-bound nature, graph mining applications encounter significant limitations with conventional processor-centric systems, like central processing units (CPUs) and graphics processing units (GPUs), stemming from the costly data movement between memory and processing units. Memory-centric computing systems, such as processing-in-memory (PIM) where computation occurs directly within or near memory modules, have the potential to accelerate graph mining. However, accelerating graph mining applications with PIM presents three primary challenges: (1) the difficulty in utilizing locality, (2) the challenge of exploring parallelism, and (3) the complexity of workload offloading between PIM and CPU. Addressing these intricate challenges, we introduce ProMiner, a novel framework that integrates three key techniques through cohesive software and hardware co-design. First, we propose a partitioning method tailored for graph mining to enhance data locality. Second, we design a coarse-fine parallelism optimization scheme to explore parallelism across different levels of memory. Third, we introduce a concurrency-aware mechanism for performance estimation, aimed at identifying the optimal computing engine for workload offloading to maximize performance. Our experimental results demonstrate that ProMiner significantly advances the state-of-the-art in graph mining, achieving 48.8% and 29.9% execution time reduction over NDMiner and DIM- Mining, respectively. Xiaoyang Lu, Xiaoming Chen 0003, Xingqi Zou, Yinhe Han 0001, Xian-He Sun |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | ACES: Accelerating Sparse Matrix Multiplication with Adaptive Execution Flow and Concurrency-Aware Cache OptimizationsabstractSparse matrix-matrix multiplication (SpMM) is a critical computational kernel in numerous scientific and machine learning applications. SpMM involves massive irregular memory accesses and poses great challenges to conventional cache-based computer architectures. Recently dedicated SpMM accelerators have been proposed to enhance SpMM performance. However, current SpMM accelerators still face challenges in adapting to varied sparse patterns, fully exploiting inherent parallelism, and optimizing cache performance. To address these issues, we introduce ACES, a novel SpMM accelerator in this study. First, ACES features an adaptive execution flow that dynamically adjusts to diverse sparse patterns. The adaptive execution flow balances parallel computing efficiency and data reuse. Second, ACES incorporates locality-concurrency co-optimizations within the global cache. ACES utilizes a concurrency-aware cache management policy, which considers data locality and concurrency for optimal replacement decisions. Additionally, the integration of a non-blocking buffer with the global cache enhances concurrency and reduces computational stalls. Third, the hardware architecture of ACES is designed to integrate all innovations. The architecture ensures efficient support across the adaptive execution flow, advanced cache optimizations, and fine-grained parallel processing. Our performance evaluation demonstrates that ACES significantly outperforms existing solutions, providing a 2.1× speedup and marking a substantial advancement in SpMM acceleration. Xiaoyang Lu, Boyu Long, Xiaoming Chen 0003, Yinhe Han 0001, Xian-He Sun |
ASPLOS (3) | 1 |
| 2024 | CHROME: Concurrency-Aware Holistic Cache Management Framework with Online Reinforcement LearningabstractCache management is a critical aspect of computer architecture, encompassing techniques such as cache replacement, bypassing, and prefetching. Existing research has often focused on individual techniques, overlooking the potential benefits of joint optimization. Moreover, many of these approaches rely on static and intuition-driven policies, limiting their performance under complex and dynamic workloads. To address these challenges, this paper introduces CHROME, a novel concurrencyaware cache management framework. CHROME takes a holistic approach by seamlessly integrating intelligent cache replacement and bypassing with pattern-based prefetching. By leveraging online reinforcement learning, CHROME dynamically adapts cache decisions based on multiple program features and applies a reward for each decision that considers the accuracy of the action and the system-level feedback information. Our performance evaluation demonstrates that CHROME outperforms current state-of-the-art schemes, exhibiting significant improvements in cache management. Notably, CHROME achieves a remarkable performance boost of up to 13.7% over the traditional LRU method in multi-core systems with only modest overhead. Xiaoyang Lu, Hamed Najafi, Jason Liu 0001, Xian-He Sun |
HPCA | 1 |
| 2024 | AceMiner: Accelerating Graph Pattern Matching using PIM with Optimized Cache SystemabstractGraph pattern matching (GPM), a critical algorithm for discovering specific patterns within complex structures, is becoming increasingly important in the data-driven world. GPM applications are memory-bound and can be accelerated by memory-centric computing systems, such as processing-in-memory (PIM). However, there are three primary challenges when it comes to accelerating GPM applications with PIM: (1) difficulty in utilizing locality, (2) heavy data movement, and (3) heavy comparison overhead due to pruning. To address these challenges, we propose AceMiner, a framework to accelerate GPM applications with a software and hardware co-design per-spective using PIM. In AceMiner, we embed hybridCache, a novel in-DRAM cache system with lower access latency and optimized replacement policy, to leverage the potential locality and reduce data movement in PIM. Additionally, we introduce a comparison unit to address the huge pruning overhead. Experimental results show that AceMiner outperforms the state-of-the-art, achieving speedups of 40.2% and 13.3% over NDMiner and DIMMining respectively, with less energy consumption and design overhead. Xiaoyang Lu, Xiaoming Chen 0003, Xingqi Zou, Yinhe Han 0001, Xian-He Sun |
ICCD | 2 |
| 2024 | Short-term photovoltaic power forecasting using hybrid contrastive learning and temporal convolutional network under future meteorological information absenceabstractAbstract Photovoltaic (PV) power generation is widely utilized to satisfy the increasing energy demand due to its cleanness and inexhaustibility. Accurate PV power forecasting can improve the penetration of PV power in the grid. However, it is pretty challenging to predict PV power in short‐term under precious future meteorological information absence conditions. To address this problem, this study proposes the hybrid Contrastive Learning and Temporal Convolutional Network (CL‐TCN), and this forecasting approach consists of two parts, including model training and adaptive processes of forecasting models. In the model training stage, this forecasting method firstly trains 18 TCN models for 18 time points from 9:00 a.m. to 17:30 p.m. These TCN models are trained by only using historical PV power data samples, and each model is used to predict the next half‐hour power output. The adaptive process of models means that, in a practical forecasting stage, PV power samples from historical data are firstly evaluated and scored by a CL based data scoring mechanism to search for the most similar data samples to current measured samples. Then these similar samples are further applied to training a single above‐mentioned well‐trained TCN model to improve its performance in forecasting the next half‐hour PV power. The experimental results tested at the time resolution of 30 min demonstrate that the proposed approach has superior performance in forecasting accuracy not only in smooth PV power samples but also in fluctuating PV power samples. Moreover, the proposed CL based data scoring mechanism can filter useless data samples effectively accelerating the forecasting process. Xiaoyang Lu, Yandang Chen, Qibin Li, Pingping Yu |
Comput. Intell. | 1 |
| 2023 | CARE: A Concurrency-Aware Enhanced Lightweight Cache Management FrameworkabstractImproving cache performance is a lasting research topic. While utilizing data locality to enhance cache performance becomes more and more difficult, data access concurrency provides a new opportunity for cache performance optimization. In this work, we propose a novel concurrency-aware cache management framework that outperforms state-of-the-art locality-only cache management schemes. First, we investigate the merit of data access concurrency and pinpoint that reducing the miss rate may not necessarily lead to better overall performance. Next, we introduce the pure miss contribution (PMC) metric, a lightweight and versatile concurrency-aware indicator, to accurately measure the cost of each outstanding miss access by considering data concurrency. Then, we present CARE, a dynamic adjustable, concurrency-aware, low-overhead cache management framework with the help of the PMC metric. We evaluate CARE with extensive experiments across different application domains and show significant performance gains with the consideration of data concurrency. In a 4-core system, CARE improves IPC by 10.3% over LRU replacement. In 8 and 16-core systems where more concurrent data accesses exist, CARE outperforms LRU by 13.0% and 17.1%, respectively. Xiaoyang Lu, Rujia Wang, Xian-He Sun |
HPCA | 1 |
| 2023 | DeepSim: A Transformer Based Model For Fast Simulation And Exploring Computer System Design SpaceabstractNo abstract available. Hamed Najafi, Xiaoyang Lu |
SIGSIM-PADS | 2 |
| 2023 | ABUSDet: A Novel 2.5D deep learning model for automated breast ultrasound tumor detection
Xudong Song, Xiaoyang Lu, Gengfa Fang, Xiangjian He, Xiaochen Fan, Le Cai, Wenjing Jia |
Appl. Intell. | 2 |
| 2023 | The Memory-Bounded Speedup Model and Its Impacts in Computing
Xian-He Sun, Xiaoyang Lu |
J. Comput. Sci. Technol. | 2 |
| 2021 | Premier: A Concurrency-Aware Pseudo-Partitioning Framework for Shared Last-Level CacheabstractAs the number of on-chip cores and application demands increase, efficient management of shared cache resources becomes imperative. Cache partitioning techniques have been studied for decades to reduce interference between applications in a shared cache and provide performance and fairness guarantees. However, there are few studies on how concurrent memory accesses affect the effectiveness of partitioning. When concurrent memory requests exist, cache miss does not reflect concurrency overlapping well. In this work, we first introduce pure misses per kilo instructions (PMPKI), a metric that quantifies the cache efficiency considering concurrent access activities. Then we propose Premier, a dynamically adaptive concurrency-aware cache pseudo-partitioning framework. Premier provides insertion and promotion policies based on PMPKI curves to achieve the benefits of cache partitioning. Finally, our evaluation of various workloads shows that Premier outperforms state-of-the-art cache partitioning schemes in terms of performance and fairness. In an 8-core system, Premier achieves 15.45% higher system performance and 10.91% better fairness than the UCP scheme. Xiaoyang Lu, Rujia Wang, Xian-He Sun |
ICCD | 1 |
| 2021 | CoPIM: A Concurrency-aware PIM Workload Offloading Architecture for Graph ApplicationsabstractProcessing-in-Memory (PIM) is considered a promising solution to improve the performance of graph-computing applications by minimizing the data movement between the host and memory. Which workload to offload and how to offload it to PIM logic determine whether the PIM architecture is well utilized. Offloading too much or too little workload from the host processor to the PIM side could hurt overall performance. On the other hand, the offloading granularity needs to be representative without losing generality. In this paper, we present CoPIM, a novel PIM workload offloading architecture that can dynamically determine which portion of the graph workload can benefit more from PIM-side computation. CoPIM focuses on the loop code blocks of graph applications and evaluates the necessity of offloading based on a concurrent memory access model. We also provide detailed architectural designs to support the offloading. In this way, CoPIM reduces the size of offloading instructions and also improves the overall performance with less energy consumption. The experimental results show that compared with other state-of-the-art PIM workload offloading frameworks, CoPIM achieves a speedup by the geometric mean of 19.5% and 11.4% than PEI and GraphPIM, respectively. On the other hand, CoPIM also reduces the un-core energy consumption by 6.8% and 6.5% on average over PEI and GraphPIM, respectively. Mingzhe Zhang 0005, Rujia Wang, Xiaoming Chen 0003, Xingqi Zou, Xiaoyang Lu, Yinhe Han 0001, Xian-He Sun |
ISLPED | 6 |
| 2020 | APAC: An Accurate and Adaptive Prefetch Framework with Concurrent Memory Access AnalysisabstractPrefetching techniques have been studied for decades. However, there are few studies on how concurrent memory accesses may affect prefetching effectiveness. When there are multiple concurrent memory requests, we can classify them into sub-classes by analyzing the overlapping relationship. In this work, we first propose pure prefetch coverage (PPC), a novel prefetching metric that can identify an accurate prefetch coverage under the concurrent memory access model. Then we propose APAC, an adaptive prefetch framework with PPC metric that can capture the dynamics of applications and adjust the prefetching aggressiveness. Our experimental results show that the PPC metric has a higher IPC correlation compared to the conventional prefetch coverage (PC) metric. For memory-intensive single-thread benchmarks, APAC provides an average performance improvement by 17.3% and 5.9% compared to the state-of-the-art adaptive prefetch framework FDP and NST. In a multi-core system, APAC outperforms FDP and NST by 8.5% and 5.0% IPC on average, respectively. Xiaoyang Lu, Rujia Wang, Xian-He Sun |
ICCD | 1 |