VLDB 2026 Research / reviewers in the wild / expert
Limin Xiao 0001
dblp:31/5990-1
· DBLP profile ↗
45ranked-venue papers
0as first author
43since 2021 · last 2026
0000-0001-9438-9181ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 33 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-DesignabstractEfficiently training large-scale models (LMs) in GPU clusters involves two separate avenues: inter-job dynamic scheduling and intra-job adaptive parallelism (AP). However, existing dynamic schedulers struggle with large-model scheduling due to the mismatch between static parallelism (SP)-aware scheduling and AP-based execution, leading to cluster inefficiencies such as degraded throughput and prolonged job queuing. This paper presents Arena, a large-model training system that co-designs dynamic scheduling and adaptive parallelism to achieve high cluster efficiency. To reduce scheduling costs while improving decision quality, Arena designs low-cost, disaggregated profiling and AP-tailored, load-aware performance estimation, while unifying them by sharding the joint scheduling-parallelism optimization space via a grid abstraction. Building on this, Arena dynamically schedules profiled jobs in elasticity and heterogeneity dimensions, and executes them using efficient AP with pruned search space. Evaluated on heterogeneous testbeds and production workloads, Arena reduces job completion time by up to 49.3% and improves cluster throughput by up to 1.6X. Chunyu Xue, Weihao Cui, Quan Chen 0002, Chen Chen 0067, Han Zhao 0005, Shulai Zhang, Linmei Wang, Limin Xiao 0001, Weifeng Zhang 0003, Jing Yang 0017, Bingsheng He, Minyi Guo |
EuroSys | 9 |
| 2026 | RL-Paxos: Relieving the Leader's Burden with Efficient Task Offloading in Distributed Consensus
Jinquan Wang, Bing Wei 0002, Xiaojian Liao, Limin Xiao 0001 |
ICDE | 6 |
| 2026 | NasZip: Software and Hardware Co-Design to Accelerate Approximate Nearest Neighbor Search with DIMM-Based Near-Data Processing
Chen Nie, Limin Xiao 0001, Weifeng Zhang 0003, Zhezhi He |
ISCA | 7 |
| 2026 | EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
Yitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou, Siyuan Cao, Xujie Fan, Yuchen Xu 0003, Junkai Chen, Chenqi Zhao, Nengyuan Zhang, Shaoke Fang, Jiangyuan Chen, Yuanfeng Chen, Zhan Wang 0003, Yuchao Zhang 0004, Yang Liu 0038, Xiangrui Yang 0002, Xiaohe Hu, Limin Xiao 0001, Weifeng Zhang 0003, Yazhu Lan, Jianbo Dong, Binzhang Fu, Wenfei Wu |
SIGCOMM | 24 |
| 2026 | A two-stage data placement strategy for cloud-edge-device collaborative environment
Runnan Shen, Jinquan Wang, Zhisheng Huo, Limin Xiao 0001, Shengyang Tan, Yuntong Li, Xiangrong Xu 0002, Liang Wang 0020 |
Comput. Commun. | 4 |
| 2026 | MEIS: Optimizing deduplication system with efficient index structure
Runnan Shen, Jinquan Wang, Zhisheng Huo, Limin Xiao 0001, Jiantong Huo, Minyi Guo, Jing Shang 0001 |
J. Syst. Archit. | 4 |
| 2026 | Accelerating LLM Inference via Low-Bit Fine-Grained Quantization Algorithm and Bit-Level Accelerator Co-DesignabstractLarge language models (LLMs) have emerged as one of the most impactful and transformative paradigms in natural language processing. Despite their remarkable success, the intensive computational demands and substantial memory footprint impose a significant barrier to efficient LLM inference.In this paper, we present a comprehensive solution to improve LLM inference performance under ultra-low weight precision, meticulously optimized through algorithm and architecture co-design. To achieve this, we first propose a fine-grained intra-cluster bit allocation method that partitions the weights into small clusters and explicitly considers the distribution of outliers and salient points within each cluster. Then, an intra-cluster protection mechanism is proposed to selectively preserve important weights during quantization, where an extended integer format and group-wise scale factor search are further introduced to mitigate accuracy degradation caused by aggressive bit-width reduction. Furthermore, we develop a memory-aligned encoding scheme to facilitate efficient memory access while enabling flexible identification of mixed-precision representations. Finally, we design a lightweight bit-level accelerator for low-bit LLM inference, offering simplified hardware design and enhanced adaptability through parallel bit-level computation. Compared to existing state-of-the-art quantization algorithms, our algorithm achieves higher model accuracy under ultra-low weight precision. Meanwhile, the proposed bit-level accelerator delivers speedups of 1.59×, 1.38×, and 1.61×, along with energy efficiency improvements of 1.52×, 1.42×, and 1.22× over ANT, OliVe, and FineQ, respectively. Xilong Xie, Liang Wang 0020, Limin Xiao 0001, Tairan Zhang, Jinquan Wang, Yongyue Wang, Xiaojian Liao |
IEEE Trans. Computers | 3 |
| 2026 | Is Intelligence the Right Direction in New OS Scheduling for Multiple Resources in Cloud Environments?abstractMaking it intelligent is a promising way in System/OS design. This article proposes OSML+, a new ML-based resource scheduling mechanism for co-located cloud services. OSML+ intelligently schedules the cache and main memory bandwidth resources at the memory hierarchy and the computing core resources simultaneously. OSML+ uses a multi-model collaborative learning approach during its scheduling and thus can handle complicated cases, e.g., avoiding resource cliffs, sharing resources among applications, enabling different scheduling policies for applications with different priorities, and so on. OSML+ can converge faster using ML models than previous studies. Moreover, OSML+ can automatically learn on the fly and handle dynamically changing workloads accordingly. Using transfer learning technologies, we show our design can work well across various cloud servers, including the latest off-the-shelf large-scale servers. Our experimental results show that OSML+ supports higher loads and meets QoS targets with lower overheads than previous studies. Xinglei Dou, Lei Liu 0037, Limin Xiao 0001 |
ACM Trans. Storage | 3 |
| 2025 | CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited MemoryabstractLarge language models like GPT-4 are resource-intensive, but recent advancements suggest that smaller, specialized experts can outperform the monolithic models on specific tasks. The Collaboration-of-Experts (CoE) approach integrates multiple expert models, improving the accuracy of generated results and offering great potential for precision-critical applications, such as automatic circuit board quality inspection. However, deploying CoE serving systems presents challenges to memory capacity due to the large number of experts required, which can lead to significant performance overhead from frequent expert switching across different memory and storage tiers. Jiashun Suo, Xiaojian Liao, Limin Xiao 0001, Jinquan Wang, Xiao Su 0002, Zhisheng Huo |
ASPLOS (2) | 3 |
| 2025 | CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUsabstractWith the rapid development of DNN applications, multi-tenant execution, where multiple DNNs are co-located on a single SoC, is becoming a prevailing trend. Although many methods are proposed in prior works to improve multi-tenant performance, the impact of shared cache is not well studied. This paper proposes CaMDN, an architecture-scheduling co-design to enhance cache efficiency for multi-tenant DNNs on integrated NPUs. Specifically, a lightweight architecture is proposed to support model-exclusive, NPU-controlled regions inside shared cache to eliminate unexpected cache contention. Moreover, a cache scheduling method is proposed to improve shared cache utilization. In particular, it includes a cache-aware mapping method for adaptability to the varying available cache capacity and a dynamic allocation algorithm to adjust the usage among co-located DNNs at runtime. Compared to prior works, CaMDN reduces the memory access by 33.4% on average and achieves a model speedup of up to 2.56 × (1.88 × on average). Tianhao Cai, Liang Wang 0020, Limin Xiao 0001, Xiaojian Liao |
DAC | 3 |
| 2025 | FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMsabstractLarge language models (LLMs) have significantly advanced the natural language processing paradigm but impose substantial demands on memory and computational resources. Quantization is one of the most effective ways to reduce memory consumption of LLMs. However, advanced single-precision quantization methods experience significant accuracy degradation when quantizing to ultra-low bits. Existing mixed-precision quantization methods are quantized by groups with coarse granularity. Employing high precision for group data leads to substantial memory overhead, whereas low precision severely impacts model accuracy. To address this issue, we propose FineQ, software-hardware co-design for low-bit fine-grained mixed-precision quantization of LLMs. First, FineQ partitions the weights into finer-grained clusters and considers the distribution of outliers within these clusters, thus achieving a balance between model accuracy and memory overhead. Then, we propose an outlier protection mechanism within clusters that uses 3 bits to represent outliers and introduce an encoding scheme for index and data concatenation to enable aligned memory access. Finally, we introduce an accelerator utilizing temporal coding that effectively supports the quantization algorithm while simplifying the multipliers in the systolic array. FineQ achieves higher model accuracy compared to the SOTA mixed-precision quantization algorithm at a close average bit-width. Meanwhile, the accelerator achieves up to 1.79x energy efficiency and reduces the area of the systolic array by 61.2%. Xilong Xie, Liang Wang 0020, Limin Xiao 0001, Xiangrong Xu 0002 |
DATE | 3 |
| 2025 | Swift-Sim: A Modular and Hybrid GPU Architecture Simulation FrameworkabstractSimulation tools are critical for architects to quickly estimate the impact of aggressive new features of GPU architecture. Existing cycle-accurate GPU simulators are typically cumbersome and slow to run. We observe that it is time-consuming and unnecessary for cycle-accurate GPU simulators to perform detailed simulations for the entire GPU when exploring the design space of specific components. This paper proposes Swift-Sim, a modular and hybrid GPU simulation framework. With a highly modular design, our framework can choose appropriate modeling approaches for each component according to requirements. For components of interest to architects, we use cycle-accurate simulation to evaluate new GPU architectures. For other components, we use analytical modeling, which accelerates simulation speed with only minor and acceptable degradation in overall accuracy. Based on this simulation framework, we present two working examples of hybrid modeling that simulate the ALU pipeline and memory accesses using analytical models. We further implement two GPU performance simulators with different levels of simplification based on Swift-Sim and evaluate them using configurations from real GPUs. The results show that the two simulators achieve an 82.6x and 211.2x geometric mean speedup compared to Accel-Sim with insignificant accuracy degradation, Xiangrong Xu 0002, Yuanqiu Lv, Liang Wang 0020, Limin Xiao 0001, Runnan Shen, Jinquan Wang |
DATE | 4 |
| 2025 | PointISA: ISA-Extensions for Efficient Point Cloud Analytics via Architecture and Algorithm Co-Design
Liang Wang 0020, Limin Xiao 0001, Hao Zhang 0205, Xilong Xie, Jianfeng Zhu 0001, Shaojun Wei, Leibo Liu |
MICRO | 3 |
| 2025 | Amove: Accelerating LLMs through Mitigating Outliers and Salient Points via Fine-Grained Grouped Vectorized Data TypeabstractThe quantization of Large Language Models (LLMs) poses significant challenges due to the heterogeneous nature of feature point distributions in low-bit quantization scenarios, including salient points, normal outliers, and massive outliers.These challenges are particularly pronounced in supporting both weight-only and weight-activation quantization modes, as existing methods often focus on a single mode and fail to address the diverse feature characteristics holistically, resulting in suboptimal model accuracy and hardware efficiency trade-offs.To tackle these limitations, we introduce Amove, a novel codesign framework that synergistically integrates data type and hardware architecture design for efficient LLM quantization.Our approach is threefold: First, we conduct a comprehensive analysis of quantization granularity and propose a residual approximation mechanism that balances model accuracy and memory overhead under fine-grained quantization.Second, we design a flexible finegrained grouped vectorized data type, enabling seamless support for both weight-activation and low-bit weight-only quantization modes within a unified framework.Third, we implement the hardware architecture of Amove on both GPU tensor core and systolic arraybased architectures.The Amove-enhanced tensor core achieves an average speedup of 2.13× and a 1.70× reduction in energy consumption over the state-of-the-art OliVe design.Furthermore, an Amove-based accelerator achieves up to 2.67× speedup and 1.68× energy reduction over the state-of-the-art accelerator. Xilong Xie, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Xiangrong Xu 0002, Jinquan Wang, Xiaojian Liao |
MICRO | 3 |
| 2025 | Efficient Performance-Aware GPU Sharing with Compatibility and Isolation through Kernel Space Interception
Shulai Zhang, Quan Chen 0002, Han Zhao 0005, Weihao Cui, Limin Xiao 0001, Minyi Guo |
USENIX ATC | 8 |
| 2025 | ICCG: low-cost and efficient consistency with adaptive synchronization for metadata replication
Liang Wang 0020, Jing Shang 0001, Zhiwen Xiao, Limin Xiao 0001, Bing Wei 0002, Runnan Shen, Jinquan Wang |
Frontiers Comput. Sci. | 5 |
| 2025 | Exploiting intra-chip locality for multi-chip GPUs via two-level shared L1 cache
Xiangrong Xu 0002, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Yuanqiu Lv, Xilong Xie, Xiaojian Liao |
J. Syst. Archit. | 3 |
| 2025 | An Intelligent Scheduling Approach on Mobile OS for Optimizing UI Smoothness and PowerabstractMobile devices need to respond quickly to diverse user inputs. The existing approaches often heuristically raise the CPU/GPU frequency according to the empirical rules when facing burst inputs and various changes. Although doing so can be effective sometimes, the existing approaches still need improvements. For instance, raising processors’ frequency can lead to high power consumption when the frequency is over-provisioned or fail to meet user demands when the frequency is under-provisioned. To this end, we propose MobiRL, a reinforcement learning-based scheduler for intelligently adjusting the CPU/GPU frequency to satisfy user demands accurately on mobile systems. MobiRL monitors the mobile system status and autonomously learns to optimize UI smoothness and power consumption by conducting CPU/GPU frequency-adjusting actions. The experimental results on the latest delivered smartphones show that MobiRL outperforms the widely used commercial scheduler on real devices—reducing the frame drop rate by 4.1% and reducing power consumption by 42.8%, respectively. Moreover, compared with a study using Q-Learning for CPU frequency scheduling, MobiRL achieves up to a 2.5% lower frame drop rate and reduces power consumption by 32.6%, respectively. Our approach has been deployed in mobile phone products. Xinglei Dou, Lei Liu 0037, Limin Xiao 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | Hierarchical Hashing: A Dynamic Hashing Method With Low Write Amplification and High Performance for Non-Volatile MemoryabstractThe hashing method is widely used as the index structure, which can be stored in NVM to improve the application performance. However, existing hashing methods may cause high extra write amplification to NVM and bring high additional storage overhead on NVM while providing low request performance. To solve these problems, we have proposed a dynamic hashing method calledHierarchical Hashing, whose basic idea is to leverage a novel hash collision resolution mechanism that can dynamically expand the size of the hash table.Hierarchical Hashingcan incur no extra write amplification to NVM when resolving hash collisions. Additionally, it can directly address all cells when resizing the hash table, thereby avoiding the additional storage overhead caused by non-addressable linked lists. Furthermore, the request performance can be improved as all cells of the hash table are addressable when resizing to resolve hash collisions. The experimental results demonstrate thatHierarchical Hashingbrings no extra write amplification to NVM and achieves nearly 90% space utilization and high request performance while providing 99% memory utilization, compared with existing representative hashing methods. Jinquan Wang, Zhisheng Huo, Limin Xiao 0001, Jinqian Yang, Jiantong Huo, Minyi Guo |
IEEE Trans. Computers | 3 |
| 2024 | FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point CloudsabstractPoint cloud analytics has become a critical workload for embedded and mobile platforms across various applications. Farthest point sampling (FPS) is a fundamental and widely used kernel in point cloud processing. However, the heavy external memory access makes FPS a performance bottleneck for real-time point cloud processing. Although bucket-based farthest point sampling can significantly reduce unnecessary memory accesses during the point sampling stage, the KD-tree construction stage becomes the predominant contributor to execution time. In this paper, we present FuseFPS, an architecture and algorithm co-design for bucket-based farthest point sampling. We first propose a hardware-friendly sampling-driven KD-tree construction algorithm. The algorithm fuses the KD-tree construction stage into the point sampling stage, further reducing memory accesses. Then, we design an efficient accelerator for bucket-based point sampling. The accelerator can offload the entire bucket-based FPS kernel at a low hardware cost. Finally, we evaluate our approach on various point cloud datasets. The detailed experiments show that compared to the state-of-the-art accelerator QuickFPS, FuseFPS achieves about $4.3\times $ and about $6.1\times $ improvements on speed and power efficiency, respectively. Liang Wang 0020, Limin Xiao 0001, Hao Zhang 0205, Xilong Xie |
ASPDAC | 3 |
| 2024 | BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point CloudsabstractPoint cloud-based machine perception applications have achieved great success in various scenarios. In this work, we focus on point cloud k-Nearest Neighbor (kNN) search, an important kernel for point clouds. Existing kNN acceleration techniques have overlooked the operation-level optimization in the Euclidean distance computation operations, which suffer from low efficiency due to a number of unnecessary computations and various data precision requirements.We reconsider point cloud kNN search from a new bitserial computation perspective and propose BitNN, a bit-serial architecture for point cloud kNN search. BitNN supports adaptive precision processing and unnecessary computing reduction, significantly improving the performance and power efficiency of kNN search. To achieve that, we first propose a bit-serial computation method for kNN search, which derives a recursive expression to compute the Euclidean distance bit by bit. Then, the dimension-wise point cloud encoding method and point-wise data layout method are proposed to enable adaptive precision processing based on bit-serial computation. Furthermore, we present an early termination mechanism for bit-serial kNN search. By estimating the lower bound of distance based on a few bits, a number of unnecessary computations can be reduced. Finally, we design an efficient bit-serial accelerator for kNN search. The accelerator exploits the massive parallelism to improve computing efficiency.We evaluate BitNN with several widely used point cloud datasets. BitNN achieves up to $6.6 \times$ speedup and $3.6 \times$ power efficiency compared to a comparable sized architecture. Moreover, BitNN can be easily integrated into existing bit-parallel kNN accelerators. We enhance the state-of-the-art kNN accelerator, ParallelNN, with bit-serial computation techniques, achieving up to $4.4 \times$ speedup and $2.9 \times$ power efficiency Liang Wang 0020, Limin Xiao 0001, Hao Zhang 0205, Tianhao Cai, Xiangrong Xu 0002 |
ISCA | 3 |
| 2024 | Minimizing the cost of periodically replicated systems via model and quantitative analysis
Liang Wang 0020, Limin Xiao 0001, Shixuan Jiang, Jinquan Wang, Bing Wei 0002, Guangjun Qin |
Frontiers Comput. Sci. | 3 |
| 2024 | iSwap: A New Memory Page Swap Mechanism for Reducing Ineffective I/O Operations in Cloud EnvironmentsabstractThis article proposes iSwap , a new memory page swap mechanism that reduces the ineffective I/O swap operations and improves the QoS for applications with a high priority in cloud environments. iSwap works in the OS kernel. iSwap accurately learns the reuse patterns for memory pages and makes the swap decisions accordingly to avoid ineffective operations. In the cases where memory pressure is high, iSwap compresses pages that belong to the latency-critical (LC) applications (or high-priority applications) and keeps them in main memory, avoiding I/O operations for these LC applications to ensure QoS, and iSwap evicts low-priority applications’ pages out of main memory. iSwap has a low overhead and works well for cloud applications with large memory footprints. We evaluate iSwap on Intel x86 and ARM platforms. The experimental results show that iSwap can significantly reduce ineffective swap operations (8.0%–19.2%) and improve the QoS for LC applications (36.8%–91.3%) in cases where memory pressure is high, compared with the latest LRU-based approach widely used in modern OSes. Zhuohao Wang, Lei Liu 0037, Limin Xiao 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic ArrayabstractThe systolic accelerator is one of the premier architectural choices for DNN acceleration. However, the conventional systolic architecture suffers from low PE utilization due to the mismatch between the fixed array and diverse DNN workloads. Recent studies have proposed flexible systolic array architectures to adapt to DNN models. However, these designs support only coarse-grained reshaping or significantly increase hardware overhead. In this study, we propose ReDas, a flexible and lightweight systolic array that supports dynamic fine-grained reshaping and multiple dataflows. First, ReDas integrates lightweight and reconfigurable roundabout data paths, which achieve fine-grained reshaping using only short connections between adjacent PEs. Second, we redesign the PE microarchitecture and integrate a set of multi-mode data buffers around the array. The PE structure enables additional data bypassing and flexible data switching. Simultaneously, the multi-mode buffers facilitate fine-grained reallocation of on-chip memory resources, adapting to various dataflow requirements. ReDas can dynamically reconfigure to up to 129 different logical shapes and 3 dataflows for a 128 × 128 array. Finally, we propose an efficient mapper to generate appropriate configurations for each layer of DNN workloads. Compared to the conventional systolic array, ReDas can achieve about 4.6× speedup and 8.3× energy-delay product (EDP) reduction. Liang Wang 0020, Limin Xiao 0001, Tianhao Cai, Xiangrong Xu 0002 |
IEEE Trans. Computers | 3 |
| 2024 | ATA-Cache: Contention Mitigation for GPU Shared L1 Cache With Aggregated Tag ArrayabstractTo fully exploit the locality of GPU applications, the GPU shared L1 cache architecture, which shares L1 cache among multiple GPU cores, is a promising architecture while still suffering from high resource contentions. We present a GPU shared L1 cache architecture with an aggregated tag array that minimizes the L1 cache contentions and takes full advantage of inter-core locality. The key idea is to decouple and aggregate the tag arrays of multiple L1 caches so that the cache requests can be compared with all tag arrays in parallel to probe the replicated data in other caches. The GPU caches are only accessed by other GPU cores when replicated data exists, filtering out unnecessary cache accesses that cause high resource contentions. We also develop a two-level thread-block scheduling policy adapted for the shared L1 cache architecture to maximize the available locality. The experimental results show that GPU performance can be improved by 14.5% on average for applications with a high inter-core locality. Xiangrong Xu 0002, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Yuanqiu Lv, Xilong Xie, Hao Liu 0107 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Accelerating zk-SNARK with Group and Zone Optimization on GPUabstractZero-knowledge proof (ZKP) is a popular cryptographic strategy for building a trusted environment, which can be applied to blockchain, electronic voting, and other scenarios. However, ZKP involves a number of computationally intensive operations that limit its widespread adoption in time-sensitive practical applications. The multi-scalar multiplication (MSM) dominates the computations and takes over 70% of the total computation time. This paper proposes a GPU-based acceleration method for ZKP by designing several optimization techniques for MSM. First, this paper constructs a formal mathematical formula of the Pippenger algorithm, which provides a theoretical optimization framework for MSM. Second, by parallelizing the prefix sum, the time complexity of the bucket reduction part of MSM is reduced from $\mathcal{O}\left( {3 \times {2^C}} \right)$ to $\mathcal{O}\left( {2 \times {2^C}} \right)$. Finally, this paper also analyzes the influence of group size on the final calculation time under different data scales and gives a suitable range of group sizes. Compared to the state-of-the-art method, our method can achieve 1.01× to 1.12× for throughput. Runnan Shen, Liang Wang 0020, Haotian Luo, Jinqian Yang, Jinquan Wang, Qiancheng Sun, Limin Xiao 0001 |
ICPADS | 8 |
| 2023 | Towards Workload Trend Time Series Probabilistic Prediction via Probabilistic Deep LearningabstractThe workloads of autonomous driving traffic accident cloud data centers exhibit high variance and uncertainty. Accurate modeling and prediction of the variance and uncertainty of cloud workloads are crucial for the realization of reliable resource management in cloud data centers. Existing solutions are point prediction methods that can not capture the variance and uncertainty of the cloud workloads. In this paper, we propose a workload probabilistic prediction method with deep learning to model and predict the variance and uncertainty of cloud workload. Our method is a hybrid deep learning model which combines exponential smoothing, bidirectional long short-term memory (BLSTM) and quantile regression. First, a cloud workload pre-processing method based on exponential smoothing is proposed to smooth the high variance feature of cloud workloads. Then, a BLSTM based cloud workload algorithm is introduced. Finally, a differentiable quantile loss function is introduced into the prediction model to generate predictions of multiple quantiles. The experimental results on the Google cluster trace show that our method outperforms other four baseline models. Heng Guo 0007, Yunzhi Xue, Yuetiansi Ji, Limin Xiao 0001 |
SSTD | 6 |
| 2023 | CFIO: A conflict-free I/O mechanism to fully exploit internal parallelism for Open-Channel SSDs
Jinbin Zhu, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Guangjun Qin |
J. Syst. Archit. | 3 |
| 2023 | A high-bandwidth and low-cost data processing approach with heterogeneous storage architectures
Bing Wei 0002, Limin Xiao 0001, Wei Wei 0006, Baicheng Yan, Zhisheng Huo |
Pers. Ubiquitous Comput. | 2 |
| 2023 | QuickFPS: Architecture and Algorithm Co-Design for Farthest Point Sampling in Large-Scale Point CloudsabstractPoint clouds have been employed extensively in machine perception applications. Farthest point sampling (FPS) is a critical kernel for point cloud processing. With the rapid growth of point cloud scale, FPS introduces a large number of memory accesses, which become the bottleneck of the large-scale point cloud processing. In this article, we present QuickFPS, an architecture and algorithm co-design of FPS in large-scale point clouds. First, we systemically analyze the characteristics of FPS and put forward a bucket-based FPS algorithm. The algorithm introduces a two-level tree data structure to organize the large-scale point cloud into multiple buckets. By using two mechanisms named merged computation and implicit computation for the buckets, the external memory accesses and compute cost are significantly reduced. Then, we design an efficient domain-specific accelerator for FPS in large-scale point clouds. The accelerator takes advantage of different forms of parallelism and further improves the accelerator’s efficiency. Finally, we evaluate QuickFPS with several widely used point cloud datasets, which include small-scale and large-scale point clouds (up to 120 000 points). Overall, QuickFPS achieves performance speedups of$43.4\times$and$12.2\times$compared to GTX 1080Ti GPU and state-of-the-art point cloud accelerator PointAcc, respectively. Liang Wang 0020, Limin Xiao 0001, Hao Zhang 0205, Xiangrong Xu 0002, Jianfeng Zhu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | EBIO: An Efficient Block I/O Stack for NVMe SSDs With Mixed WorkloadsabstractWith the advent of high-performance nonvolatile memory express (NVMe) SSD, the overhead caused by the storage software stack becomes a significant bottleneck for exploiting the potential of NVMe SSD. Recent I/O isolation approaches eliminate CPU switching, I/O interference, and lock contention for I/O queues by pinning I/O threads in isolated and dedicated I/O paths. However, they degrade the overall performance of mixed workloads with heterogeneous I/O demands. The I/O-intensive workloads issue multiple I/O requests and quickly fill up their dedicated I/O queues, resulting in I/O wait. On the contrary, the non-I/O-intensive workloads cannot deliver enough I/O requests to saturate their associated I/O queues. Moreover, frequent allocations and deallocations of I/O request objects expose a significant overhead for I/O-intensive workloads. In this article, we propose EBIO, an efficient block I/O stack for NVMe SSD, to improve the overall performance of mixed workloads. EBIO contains an on-demand queue management strategy (ODQM) and a reusable I/O management strategy (RERM). Specifically, ODQM eliminates I/O wait by dynamically adding queues for I/O-intensive workloads and leverages I/O generation time to guarantee strong sequential consistency and fairness in scheduling I/O requests. RERM reduces the overhead caused by repeatedly allocating I/O request objects for I/O-intensive workloads by reusing the allocated objects. Experimental results show that, compared to the state-of-the-art approaches, EBIO improves input/output operations per second by up to 16.44% and reduces I/O latency by up to 32.76%. Jinbin Zhu, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Guangjun Qin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Cloud Workload Turning Points Prediction via Cloud Feature-Enhanced Deep LearningabstractCloud workload turning point is either a local peak point standing for workload pressure or a local valley point standing for resource waste. Predicting such critical points is important to give warnings to system managers to take precautionary measures aimed at achieving high resource utilization, quality of service (QoS), and profit of the investment. Existing researches mainly focus more on the workload's future point value prediction only, whereas trend-based turning point prediction is not considered. Moreover, one of the most critical challenges during the prediction is the fact that traditional trend prediction methods which succeed in financial and industrial areas, etc., have a weak ability to represent the cloud features, which means that they cannot describe the highly-variable cloud workloads time series. This article introduces a novel cloud workload turning point prediction approach based on cloud feature-enhanced deep learning. First, we establish a turning point prediction model of cloud server workload considering cloud workload features. Then, a cloud feature-enhanced deep learning model is designed for workload turning point prediction. Experiments on the most famous Google cluster demonstrate the effectiveness of our model compared with state-of-the-art models. To the best of our knowledge, this article is the first systematic research on turning point-based trend prediction of cloud workload time series by cloud feature-enhanced deep learning. Shaoning Li, Jiaxun Lv, Tianyuan Zhang 0004, Limin Xiao 0001, Haiguang Fang, Chunhao Wang, Yunzhi Xue |
IEEE Trans. Cloud Comput. | 6 |
| 2023 | Global Virtual Data Space for Unified Data Access Across Supercomputing CentersabstractIn the wide-area high-performance computing environment, heterogeneous storage resources are geographically distributed in different supercomputing centers, which leads to the barriers between applications and data. This paper proposes a global virtual data space, named GVDS, to meet the needs of unified data access across supercomputing centers. GVDS integrates the parallel/distributed file systems of supercomputing centers to present a virtual space with tremendous storage capability for users. GVDS organizes users into groups for easy management, which allows users to share, collaborate, and perform computations on the stored data. For failure tolerance, global metadata is replicated and distributed on multiple supercomputing centers, redundant I/O service components are deployed in each supercomputing center. GVDS uses adaptive prefetching, caching, and request merging to improve access performance. Experimental results running on real-world supercomputing centers show that, GVDS can deliver excellent I/O performance running micro-benchmark, real-world traces and applications. Bing Wei 0002, Limin Xiao 0001, Hanjie Zhou, Guangjun Qin |
IEEE Trans. Cloud Comput. | 2 |
| 2023 | Dynamic two-side matching of tasks and resources in wide-area distributed computing environments
Liang Wang 0020, Limin Xiao 0001, Runnan Shen, Jinquan Wang |
J. Supercomput. | 3 |
| 2022 | Smart scheduler: an adaptive NVM-aware thread scheduling approach on NUMA systems
Yuetao Chen, Keni Qiu, Haipeng Jia, Yunquan Zhang, Limin Xiao 0001, Lei Liu 0037 |
CCF Trans. High Perform. Comput. | 6 |
| 2022 | Publisher Correction: Smart scheduler: an adaptive NVM-aware thread scheduling approach on NUMA systems
Yuetao Chen, Keni Qiu, Haipeng Jia, Yunquan Zhang, Limin Xiao 0001, Lei Liu 0037 |
CCF Trans. High Perform. Comput. | 6 |
| 2022 | A self-tuning client-side metadata prefetching scheme for wide area network file systems
Bing Wei 0002, Limin Xiao 0001, Guangjun Qin, Jinbin Zhu, Baicheng Yan, Chaobo Wang, Zhisheng Huo |
Sci. China Inf. Sci. | 2 |
| 2022 | Evaluating performance variations cross cloud data centres using multiview comparative workload traces analysisabstractHow to evaluate the performance variations of large-scale cloud data centres is challenging due to diverse nature of cloud platforms. Classic methods such as profiling-based evaluating methods tend to only provide global statistics for a system compared with cloud tracing based approaches. However, existing tracing based research lacks a systematic comparative multiview analysis from architecure-view to job-view and task-view, etc.to evaluate cloud performance variations, together with a detailed case study. We introduce MuCoTrAna, a multiview comparative workload traces analysis approach to evaluate the performance variations of large-scale cloud data centres which assists the cloud platform performance managers and big trace analysts. The efficiency of the proposed approach is demonstrated via case studies in Alibaba 2018 trace and Google trace. The multifaceted analysis results of traces reveals the qualitative insights, performance bottlenecks, inferences and adequate suggestions from global view, machine view, job-task view, etc. Xiangrong Xu 0002, Limin Xiao 0001, Lei Ren 0001, Nasro Min-Allah, Yunzhi Xue |
Connect. Sci. | 3 |
| 2022 | Hypergraph-partitioning-based online joint scheduling of tasks and data
Liang Wang 0020, Limin Xiao 0001, Wei Wei 0006, Rafal Scherer, Guangjun Qin, Jinquan Wang |
J. Supercomput. | 3 |
| 2021 | UPM-DMA: An Efficient Userspace DMA-Pinned Memory Management Strategy for NVMe SSDs
Jinbin Zhu, Limin Xiao 0001, Liang Wang 0020, Guangjun Qin, Zhonglin Liu |
ICA3PP (1) | 2 |
| 2021 | Erasure-Coded Multi-Block Updates Based on Hybrid Writes and Common XORs FirstabstractErasure code is widely used in storage systems since it can offer higher reliability at lower redundancy than data replication. However, erasure coding based storage systems have to perform multi-block updates for partial writes of an erasure coding group, which leads to a large number of XOR operations. This paper presents an efficient approach, named ECMU, for erasure-coded multi-block update under a stringent latency by scheduling update sequences. ECMU takes a hybrid of reconstructed-write and read-modify-write for parity blocks of an erasure coding group, it dynamically selects the write scheme with the fewer XORs for each parity block to be updated, in order to reduce the number of XORs. ECMU iteratively retrieves the unmodified parity blocks to calculate the minimum XORs for each write scheme. For all parity blocks to be updated, after the write schemes are determined, ECMU performs the common XORs first, then it reuses the computational results to further reduce the number of XORs. ECMU caches a certain number of scheduling schemes to reduce the construction count of the scheduling schemes. Experimental results on real-world trace replaying show that the number of XORs and update time can be reduced significantly, compared with the state-of-the-art. Bing Wei 0002, Jigang Wu, Limin Xiao 0001 |
ICCD | 4 |
| 2021 | Data Delta Based Hybrid Writes for Erasure-Coded Storage Systems
Bing Wei 0002, Jigang Wu, Limin Xiao 0001 |
NPC | 5 |
| 2021 | Fine-grained management of I/O optimizations based on workload characteristics
Bing Wei 0002, Limin Xiao 0001, Bingyu Zhou, Guangjun Qin, Baicheng Yan, Zhisheng Huo |
Frontiers Comput. Sci. | 2 |
| 2020 | Incremental Throughput Allocation of Heterogeneous Storage With No Disruptions in Dynamic SettingabstractSolid-state drives (SSDs) have been added into storage systems for improving their performance, which will bring the heterogeneity into the storage medium. The throughput is one of the essential resources in heterogeneous storage systems, and how to allocate the throughput plays a crucial role in user performance. There are many types of research on the throughput allocation of heterogeneous storage systems. However, the throughput allocation of heterogeneous storage is facing new challenges in a dynamic setting, where users are not present in the system simultaneously, and enter the system dynamically. Drawing on economic gametheory, researchers have proposed many methods to tackle dynamic throughput allocation issues for heterogeneous storages, cross out enjoying Sharing Incentive (SI), Envy Freeness (EF), and Pareto Optimality (PO). However, they either relax constraints of fairness property to cause the allocation with weak fairness or interrupt some users present in the system to give up a piece of their allocations for new users entering the system, which will degrade these donors' performance. Moreover, all of existing methods will cause lower resource utilization due to constraints of users' dominant share equality. In this article, we propose a dynamic throughout allocation method based on gradual increase (DAGI), which can adapt to various workloads to make a fair allocation with a maximum resource utilization. Without relaxing constraints of fairness properties, when new users enter the system, DAGI can make a dynamic allocation with strong fairness by appropriately postponing the allocation of surplus throughputs, so this can provide an opportunity that DAGI can guarantee the final allocation with strong fairness when allocating remaining throughputs after all users are present in the system. Meanwhile, DAGI can gradually increase user allocation without reduction, which will not interrupt any users present in the system. Furthermore, DAGI can conduct a dynamic throughput allocation based on users' local bottleneck resources, which can adapt to various workloads of users to improve resource utilization. Extensive experiments are conducted to prove the effectiveness of DAGI. The experimental results show that DAGI can achieve higher resource utilization and performance than existing methods, and can satisfy desirable game-theoretic properties with guaranteeing the strong fairness. In addition, DAGI gradually increases the allocation of each user without interrupting any user to reduce its allocation to degrade its performance. Zhisheng Huo, Limin Xiao 0001, Minyi Guo, Xiaoling Rong |
IEEE Trans. Computers | 2 |
| 2019 | I/O Optimizations Based on Workload Characteristics for Parallel File Systems
Bing Wei 0002, Limin Xiao 0001, Bingyu Zhou, Guangjun Qin, Baicheng Yan, Zhisheng Huo |
NPC | 2 |