Junqing Lin

dblp:356/5104 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0007-1455-8725ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 83% Deep learning architectures and training · 17%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 84% GPUs and heterogeneous computing · 16%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.922026
CommitMoE: Efficient Fallback-Free MoE Inference with Offloading Under GPU Memory Constraints · AAAI 2026
Lua-LLM: Learning Unstructured-Sparsity Allocation for Large Language Models · NeurIPS 2025
Machine learning › Deep learning architectures and training › mixture of experts
mixture-of-experts inference
1.012026
CommitMoE: Efficient Fallback-Free MoE Inference with Offloading Under GPU Memory Constraints · AAAI 2026
Machine learning › Efficient and distributed learning
model offloading
1.012026
CommitMoE: Efficient Fallback-Free MoE Inference with Offloading Under GPU Memory Constraints · AAAI 2026
Machine learning › Efficient and distributed learning › model compression
large language model compression
0.912025
Lua-LLM: Learning Unstructured-Sparsity Allocation for Large Language Models · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression
pruning
0.912025
Lua-LLM: Learning Unstructured-Sparsity Allocation for Large Language Models · NeurIPS 2025
Compilers and program optimization
autotuning
0.812024
LO-SpMM: Low-cost Search for High-performance SpMM Kernels on GPUs · ACM Trans. Archit. Code Optim. 2024
Compilers and program optimization › domain-specific compilation
tensor algebra compilation
0.812024
LO-SpMM: Low-cost Search for High-performance SpMM Kernels on GPUs · ACM Trans. Archit. Code Optim. 2024
High-performance computing › sparse linear algebra
sparse matrix multiplication
0.812024
LO-SpMM: Low-cost Search for High-performance SpMM Kernels on GPUs · ACM Trans. Archit. Code Optim. 2024
High-performance computing › sparse linear algebra › sparse matrix multiplication
SpMM
0.812024
LO-SpMM: Low-cost Search for High-performance SpMM Kernels on GPUs · ACM Trans. Archit. Code Optim. 2024
GPUs and heterogeneous computing
GPU memory management
0.312026
CommitMoE: Efficient Fallback-Free MoE Inference with Offloading Under GPU Memory Constraints · AAAI 2026
Machine learning › Efficient and distributed learning › inference efficiency
sparse neural network inference
0.212024
LO-SpMM: Low-cost Search for High-performance SpMM Kernels on GPUs · ACM Trans. Archit. Code Optim. 2024

Methods — techniques the papers use, named apart from their topics

search space reduction · 2.3rank-based cost model · 2.3proxy estimation · 2.3expert prefetching · 2.0commit router · 2.0soft top-k operator · 0.9binary mask learning · 0.9
YearPublicationVenuePosition
2026 CommitMoE: Efficient Fallback-Free MoE Inference with Offloading Under GPU Memory Constraints
abstract
Mixture of Experts (MoE) models have emerged as a promising approach to scale language models efficiently by activating only a subset of parameters for each input. However, deploying these models under GPU memory constraints remains challenging, as existing offloading strategies incur significant overhead from CPU-GPU data transfers. While prior work has explored prefetching techniques to mitigate this bottleneck, these methods require costly fallback mechanisms when predictions fail. Since expert transfers cannot be canceled once initiated, the correct experts need to be loaded on demand sequentially, introducing additional latency. To address this, we present CommitMoE, a novel approach featuring a Commit Router that makes execution decisions based on expert predictions without fallback mechanisms. Our key insight reveals that router certainty strongly correlates with prediction accuracy, while in low-certainty scenarios, the model output demonstrates inherent robustness to expert selection. Leveraging this insight to design a systems-level solution, CommitMoE achieves 1.3× to 9.4× faster inference across different environments and datasets compared to state-of-the-art offloading frameworks while maintaining model quality.
Jingwei Sun 0001, Junqing Lin, Guangzhong Sun
AAAI3
2025 Lua-LLM: Learning Unstructured-Sparsity Allocation for Large Language Models
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their extensive parameter scales pose significant challenges for practical deployment. Unstructured pruning has emerged as an effective model compression strategy with minimal performance loss, which introduces fine-grained sparsity for weight parameters. While existing methods employ a layer-wise pruning strategy to avoid the complexity of global pruning for billion-scale LLMs, they require appropriate sparsity allocation for the layer-wise pruning objectives and often lead to suboptimal solutions for the overall model. In this paper, we propose Lua-LLM ($\textbf{L}$earning $\textbf{u}$nstructured-sparsity $\textbf{a}$llocation in LLMs), a learning-based global pruning framework that explores the optimal unstructured sparsity allocation. Unlike existing pruning methods, which primarily focus on allocating per-layer sparsity, Lua-LLM achieves flexible allocation for both layer-wise and intra-layer sparsity. Furthermore, Lua-LLM leverages a soft Top-K operator to approximate the importance-based mask selection mechanism, enabling efficient binary mask learning. Experimental results on LLaMA and OPT families demonstrate significant performance improvements over existing methods.
Mingge Lu, Jingwei Sun 0001, Junqing Lin, Zechun Zhou, Guangzhong Sun
NeurIPS3
2024 LO-SpMM: Low-cost Search for High-performance SpMM Kernels on GPUs
abstract
As deep neural networks (DNNs) become increasingly large and complicated, pruning techniques are proposed for lower memory footprint and more efficient inference. The most critical kernel to execute pruned sparse DNNs on GPUs is Sparse-dense Matrix Multiplication (SpMM). To maximize the performance of SpMM, despite the high-performance implementation generated from advanced tensor compilers, they often take a long time to iteratively search tuning configurations. Such a long time slows down the cycle of exploring better DNN architectures or pruning algorithms. In this article, we propose LO-SpMM to efficiently generate high-performance SpMM implementations for sparse DNN inference. Based on the analysis of nonzero elements’ layout, the characterization of the GPU architecture, and a rank-based cost model, LO-SpMM can effectively reduce the search space and eliminate possibly low-performance candidates. Besides, rather than generating complete SpMM implementations for evaluation, LO-SpMM constructs simplified proxies to quickly estimate performance, thereby substantially reducing compilation and execution costs. Experimental results show that LO-SpMM can reduce the search time by 281× at most, while the performance of generated SpMM implementations is comparable to or better than the state-of-the-art sparse tensor compiling solutions.
Junqing Lin, Jingwei Sun 0001, Honghe Zhang, Xianzhi Yu, Guangzhong Sun
ACM Trans. Archit. Code Optim.1
2023 EC-SpMM: Efficient Compilation of SpMM Kernel on GPUs
abstract
As deep neural networks (DNNs) become increasingly large and complicated, pruning techniques are proposed for lower memory footprint and more efficient inference. The most critical kernel to execute pruned sparse DNNs on GPUs is Sparse-dense Matrix Multiplication (SpMM). To maximize the performance of SpMM, despite the high-performance code generated from recent tensor compilers, they often take a long time for iteratively searching candidate configurations. Such a long time slows down the cycle of exploring better DNN architectures or pruning algorithms. In this paper, we propose EC-SpMM to efficiently generate high-performance SpMM kernels for sparse DNN inference. Based on the analysis of nonzero elements’ layout, the characterization of GPU architecture, and a rank-based cost model, EC-SpMM can effectively reduce the search space and eliminate possibly low-performance candidates. Experimental results show that EC-SpMM can reduce the compilation time by a factor of 35 ×, while the performance of generated SpMM kernels is comparable or even better, compared with the state-of-the-art sparse tensor compiling solution.
Junqing Lin, Honghe Zhang, Jingwei Sun 0001, Xianzhi Yu, Guangzhong Sun
ICPP1