Zhuoxin Bai

dblp:305/3942 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0005-5135-5594ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HitKV: Activation Frequency Knows Which Tokens Are Important
abstract
The demand for long-context processing in large language models (LLMs) continues to escalate alongside rapid advancements in their capabilities. However, the intermediate attention keys and values (KV cache) employed to avoid re-computations, also grow linearly with sequence length, far exceeding the memory capacity of consumer-grade GPUs. Consequently, many studies have proposed KV cache compression methods that evict unimportant tokens based on variant attention scoring strategies. These methods typically retain the KV pairs of the top-k scoring tokens under a fixed memory budget. However, they still face several limitations. First, they disregard the activation frequency of tokens, specifically the count of times tokens achieve top-k scores in the attention distribution of following tokens. The methods based on variant attention scores may incorrectly evict some high-activation-frequency yet low final-scoring tokens. Second, the activation frequency exhibits different distribution patterns across layers and tasks. Neglecting these differences negatively impacts model performance and task adaptability. Our analysis of the actual token activation frequency and its unique characteristics across layers and task types reveals potential opportunities to address these issues. In this paper, we propose HitKV, which employs hit rates to directly characterize token activation frequencies, enabling adaptive layer-aware and task-aware KV cache eviction under the uniform memory allocation strategies. Also, HitKV can be easily integrated into layer-specific memory allocation methods. Experimental results demonstrate that HitKV maintains model performance with preserving only 3% of the KV cache, achieves high-quality generation outputs in long-text generation tasks, and delivers 4× throughput improvement over baselines.
Sanle Zhao, Yujuan Tan, Jing Yu 0026, Zhuoxin Bai, Zongjie Wang, Ao Ren
AAAI4
2026 Latency Optimization in Hybrid Memory System for GNNs
abstract
Graph Neural Networks (GNNs) require high-capacity, low-latency memory systems to process large graphs. A hierarchical hybrid memory architecture combining high-capacity Non-Volatile Memory (NVM) and low-latency DRAM offers a promising solution. However, the inherent sparsity of graph data results in poor locality for GNN memory requests, leading to low DRAM cache hit rates and numerous misses, which significantly impairs the hybrid memory system’s performance. A critical issue is that DRAM misses in serial access mode incur substantial latency. While parallel access mode can mitigate this for misses, it introduces long-tail latency and wastes bandwidth for DRAM hits. In this paper, we focus on addressing these issues from two aspects: increasing the cache hit rate and decreasing the miss latency. We mainly propose two predictors: a future data access predictor that enables accurate prefetching to DRAM, thereby improving cache hit rates, and a data location predictor that determines whether data resides in DRAM or NVM, optimizing the choice between serial and parallel access modes to reduce miss latency. By integrating these predictors, we achieve efficient data access in both DRAM and NVM. Our experiments show a 49.5% reduction in memory delay and a 38.1% increase in memory bandwidth utilization compared to baseline.
Zhaoyang Zeng, Yujuan Tan, Wei Chen 0101, Zhuoxin Bai, Ao Ren, Duo Liu 0002, Xianzhang Chen
IEEE Trans. Computers5
2025 Cocache: An Accurate and Low-Overhead Dynamic Caching Method for GNNs
Zhaoyang Zeng, Yujuan Tan, Zhuoxin Bai, Kan Zhong, Duo Liu 0002, Ao Ren
Euro-Par (2)4
2025 DualSpar: A Dual-Granularity Memory Framework with Adaptive Sparsity for Efficient LLM Inference
abstract
The block-based inference engine, powered by noncontiguous key-value (KV) cache management, has emerged as a new paradigm for large language model (LLM) inference due to its efficient memory utilization. However, in large-batch, longcontext workloads, the substantial demand for KV cache remains a major bottleneck in inference performance. Existing research leverages the sparsity of the attention mechanisms by removing non-critical tokens to limit the KV cache size. However, we observe that in block-based inference engines, current sparse methods require defragmentation after token removal to maintain tensor continuity, incurring significant overhead. Additionally, we find a correlation between request length and sparsity potential, yet existing methods apply a uniform sparsity strategy at batch level without dynamic adjustment. To address these, we propose DualSpar, a novel sparse KV cache framework with dual-granularity memory and adaptive sparsity strategy. First, it binds token importance to KV cache block granularity, achieving low-overhead pre-consolidation. Second, it incorporates system load and request length into sparsity decision, fully exploiting the sparse potential of different requests while reducing the queuing latency in large-batch processing. Evaluations show that DualSpar achieves up to a 3.16× throughput improvement, a 3.75× faster time-to-first-token (TTFT), and an 87.2% reduction in defragmentation overhead while maintaining high accuracy.
Yujuan Tan, Zhuoxin Bai, Sanle Zhao, Yujiao Wang, Zongjie Wang, Ao Ren, Kan Zhong
ICCD3
2025 RobTrack: A Robust 3D Multi-object Tracking Method for Edge Devices
Mingrui Qiang, Ao Ren, Yujuan Tan, Jing Yu 0026, Zhuoxin Bai, Duo Liu 0002, Kan Zhong, Chaoxia Qin
ICIC (5)5
2025 GNNBoost: Accelerating sampling-based GNN training on large scale graph by optimizing data preparation
Yujuan Tan, Yan Gan, Zhaoyang Zeng, Zhuoxin Bai, Lei Qiao 0002, Duo Liu 0002, Kan Zhong, Ao Ren
J. Syst. Archit.4
2025 LAShards: Low-Overhead and Self-Adaptive MRC Construction for Non-Stack Algorithms
abstract
Shared cache systems have become increasingly crucial, especially in cloud services, where the Miss Ratio Curve (MRC) is a widely used tool for evaluating cache performance. The MRC depicts the relationship between the cache miss ratio and cache size, indicating how cache performance trends with varying cache sizes. Recent advancements have enabled efficient MRC construction for stack replacement policies. For non-stack policies, miniature simulation downsizes the actual cache size and data stream through spatially hashed sampling, providing a general method for MRC construction. However, this approach still faces significant challenges. Firstly, constructing an MRC requires numerous mini-caches to obtain miss ratios, consuming significant cache resources, leading to tremendous memory and computing overhead. Secondly, it cannot adapt to the dynamic I/O workloads, resulting in less precise MRC.To address these issues, we propose LAShards, a low-overhead and self-adaptive MRC construction method for non-stack replacement policies. The key idea behind LAShards is to exploit the locality and burstiness in access patterns. It can statically reduce memory usage and dynamically adapt to workloads. Compared to previous works, LAShards can save up to 20× of memory resources, and increase throughput by up to 10×.
Sanle Zhao, Yujuan Tan, Zhaoyang Zeng, Jing Yu 0026, Zhuoxin Bai, Ao Ren, Xianzhang Chen, Duo Liu 0002
IEEE Trans. Computers5
2024 BGS: Accelerate GNN training on multiple GPUs
Yujuan Tan, Zhuoxin Bai, Duo Liu 0002, Zhaoyang Zeng, Yan Gan, Ao Ren, Xianzhang Chen, Kan Zhong
J. Syst. Archit.2