Jinquan Wang

dblp:40/1277 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 1 first-author · 9 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 RL-Paxos: Relieving the Leader's Burden with Efficient Task Offloading in Distributed Consensus
Jinquan Wang, Bing Wei 0002, Xiaojian Liao, Limin Xiao 0001
ICDE2
2026 A two-stage data placement strategy for cloud-edge-device collaborative environment
Runnan Shen, Jinquan Wang, Zhisheng Huo, Limin Xiao 0001, Shengyang Tan, Yuntong Li, Xiangrong Xu 0002, Liang Wang 0020
Comput. Commun.2
2026 MEIS: Optimizing deduplication system with efficient index structure
Runnan Shen, Jinquan Wang, Zhisheng Huo, Limin Xiao 0001, Jiantong Huo, Minyi Guo, Jing Shang 0001
J. Syst. Archit.2
2026 Accelerating LLM Inference via Low-Bit Fine-Grained Quantization Algorithm and Bit-Level Accelerator Co-Design
abstract
Large language models (LLMs) have emerged as one of the most impactful and transformative paradigms in natural language processing. Despite their remarkable success, the intensive computational demands and substantial memory footprint impose a significant barrier to efficient LLM inference.In this paper, we present a comprehensive solution to improve LLM inference performance under ultra-low weight precision, meticulously optimized through algorithm and architecture co-design. To achieve this, we first propose a fine-grained intra-cluster bit allocation method that partitions the weights into small clusters and explicitly considers the distribution of outliers and salient points within each cluster. Then, an intra-cluster protection mechanism is proposed to selectively preserve important weights during quantization, where an extended integer format and group-wise scale factor search are further introduced to mitigate accuracy degradation caused by aggressive bit-width reduction. Furthermore, we develop a memory-aligned encoding scheme to facilitate efficient memory access while enabling flexible identification of mixed-precision representations. Finally, we design a lightweight bit-level accelerator for low-bit LLM inference, offering simplified hardware design and enhanced adaptability through parallel bit-level computation. Compared to existing state-of-the-art quantization algorithms, our algorithm achieves higher model accuracy under ultra-low weight precision. Meanwhile, the proposed bit-level accelerator delivers speedups of 1.59×, 1.38×, and 1.61×, along with energy efficiency improvements of 1.52×, 1.42×, and 1.22× over ANT, OliVe, and FineQ, respectively.
Xilong Xie, Liang Wang 0020, Limin Xiao 0001, Tairan Zhang, Jinquan Wang, Yongyue Wang, Xiaojian Liao
IEEE Trans. Computers6
2025 CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
abstract
Large language models like GPT-4 are resource-intensive, but recent advancements suggest that smaller, specialized experts can outperform the monolithic models on specific tasks. The Collaboration-of-Experts (CoE) approach integrates multiple expert models, improving the accuracy of generated results and offering great potential for precision-critical applications, such as automatic circuit board quality inspection. However, deploying CoE serving systems presents challenges to memory capacity due to the large number of experts required, which can lead to significant performance overhead from frequent expert switching across different memory and storage tiers.
Jiashun Suo, Xiaojian Liao, Limin Xiao 0001, Jinquan Wang, Xiao Su 0002, Zhisheng Huo
ASPLOS (2)5
2025 Swift-Sim: A Modular and Hybrid GPU Architecture Simulation Framework
abstract
Simulation tools are critical for architects to quickly estimate the impact of aggressive new features of GPU architecture. Existing cycle-accurate GPU simulators are typically cumbersome and slow to run. We observe that it is time-consuming and unnecessary for cycle-accurate GPU simulators to perform detailed simulations for the entire GPU when exploring the design space of specific components. This paper proposes Swift-Sim, a modular and hybrid GPU simulation framework. With a highly modular design, our framework can choose appropriate modeling approaches for each component according to requirements. For components of interest to architects, we use cycle-accurate simulation to evaluate new GPU architectures. For other components, we use analytical modeling, which accelerates simulation speed with only minor and acceptable degradation in overall accuracy. Based on this simulation framework, we present two working examples of hybrid modeling that simulate the ALU pipeline and memory accesses using analytical models. We further implement two GPU performance simulators with different levels of simplification based on Swift-Sim and evaluate them using configurations from real GPUs. The results show that the two simulators achieve an 82.6x and 211.2x geometric mean speedup compared to Accel-Sim with insignificant accuracy degradation,
Xiangrong Xu 0002, Yuanqiu Lv, Liang Wang 0020, Limin Xiao 0001, Runnan Shen, Jinquan Wang
DATE7
2025 Amove: Accelerating LLMs through Mitigating Outliers and Salient Points via Fine-Grained Grouped Vectorized Data Type
abstract
The quantization of Large Language Models (LLMs) poses significant challenges due to the heterogeneous nature of feature point distributions in low-bit quantization scenarios, including salient points, normal outliers, and massive outliers.These challenges are particularly pronounced in supporting both weight-only and weight-activation quantization modes, as existing methods often focus on a single mode and fail to address the diverse feature characteristics holistically, resulting in suboptimal model accuracy and hardware efficiency trade-offs.To tackle these limitations, we introduce Amove, a novel codesign framework that synergistically integrates data type and hardware architecture design for efficient LLM quantization.Our approach is threefold: First, we conduct a comprehensive analysis of quantization granularity and propose a residual approximation mechanism that balances model accuracy and memory overhead under fine-grained quantization.Second, we design a flexible finegrained grouped vectorized data type, enabling seamless support for both weight-activation and low-bit weight-only quantization modes within a unified framework.Third, we implement the hardware architecture of Amove on both GPU tensor core and systolic arraybased architectures.The Amove-enhanced tensor core achieves an average speedup of 2.13× and a 1.70× reduction in energy consumption over the state-of-the-art OliVe design.Furthermore, an Amove-based accelerator achieves up to 2.67× speedup and 1.68× energy reduction over the state-of-the-art accelerator.
Xilong Xie, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Xiangrong Xu 0002, Jinquan Wang, Xiaojian Liao
MICRO7
2025 ICCG: low-cost and efficient consistency with adaptive synchronization for metadata replication
Liang Wang 0020, Jing Shang 0001, Zhiwen Xiao, Limin Xiao 0001, Bing Wei 0002, Runnan Shen, Jinquan Wang
Frontiers Comput. Sci.9
2025 MSF-GODE: Multi-scale frequency-domain learning in graph neural ODEs for accurate traffic flow forecasting
Jilong Tang, Xiaojiao Jiang, Jinquan Wang
Neurocomputing6
2025 Hierarchical Hashing: A Dynamic Hashing Method With Low Write Amplification and High Performance for Non-Volatile Memory
abstract
The hashing method is widely used as the index structure, which can be stored in NVM to improve the application performance. However, existing hashing methods may cause high extra write amplification to NVM and bring high additional storage overhead on NVM while providing low request performance. To solve these problems, we have proposed a dynamic hashing method calledHierarchical Hashing, whose basic idea is to leverage a novel hash collision resolution mechanism that can dynamically expand the size of the hash table.Hierarchical Hashingcan incur no extra write amplification to NVM when resolving hash collisions. Additionally, it can directly address all cells when resizing the hash table, thereby avoiding the additional storage overhead caused by non-addressable linked lists. Furthermore, the request performance can be improved as all cells of the hash table are addressable when resizing to resolve hash collisions. The experimental results demonstrate thatHierarchical Hashingbrings no extra write amplification to NVM and achieves nearly 90% space utilization and high request performance while providing 99% memory utilization, compared with existing representative hashing methods.
Jinquan Wang, Zhisheng Huo, Limin Xiao 0001, Jinqian Yang, Jiantong Huo, Minyi Guo
IEEE Trans. Computers1
2024 Minimizing the cost of periodically replicated systems via model and quantitative analysis
Liang Wang 0020, Limin Xiao 0001, Shixuan Jiang, Jinquan Wang, Bing Wei 0002, Guangjun Qin
Frontiers Comput. Sci.6
2023 Accelerating zk-SNARK with Group and Zone Optimization on GPU
abstract
Zero-knowledge proof (ZKP) is a popular cryptographic strategy for building a trusted environment, which can be applied to blockchain, electronic voting, and other scenarios. However, ZKP involves a number of computationally intensive operations that limit its widespread adoption in time-sensitive practical applications. The multi-scalar multiplication (MSM) dominates the computations and takes over 70% of the total computation time. This paper proposes a GPU-based acceleration method for ZKP by designing several optimization techniques for MSM. First, this paper constructs a formal mathematical formula of the Pippenger algorithm, which provides a theoretical optimization framework for MSM. Second, by parallelizing the prefix sum, the time complexity of the bucket reduction part of MSM is reduced from $\mathcal{O}\left( {3 \times {2^C}} \right)$ to $\mathcal{O}\left( {2 \times {2^C}} \right)$. Finally, this paper also analyzes the influence of group size on the final calculation time under different data scales and gives a suitable range of group sizes. Compared to the state-of-the-art method, our method can achieve 1.01× to 1.12× for throughput.
Runnan Shen, Liang Wang 0020, Haotian Luo, Jinqian Yang, Jinquan Wang, Qiancheng Sun, Limin Xiao 0001
ICPADS6
2023 Dynamic two-side matching of tasks and resources in wide-area distributed computing environments
Liang Wang 0020, Limin Xiao 0001, Runnan Shen, Jinquan Wang
J. Supercomput.5
2022 Hypergraph-partitioning-based online joint scheduling of tasks and data
Liang Wang 0020, Limin Xiao 0001, Wei Wei 0006, Rafal Scherer, Guangjun Qin, Jinquan Wang
J. Supercomput.7
2006 Partially Introducing Formal Methods into Object-Oriented Development: Case Studies Using a Metrics-Driven Approach
Yujun Zheng 0001, Jinquan Wang, Jinyun Xue
FM2