VLDB 2026 Research / reviewers in the wild / expert
Qirui Zhou
dblp:259/2609
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0003-8646-1034ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 92% Program synthesis and code generation · 8% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 60% Hardware accelerators and domain-specific architectures · 30% GPUs and heterogeneous computing · 10% | |
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
code generation |
1.7 | 2 | 2025 | QiMeng-TensorOp: One-Line Prompt is Enough for High-Performance Tensor Operator Generation with Hardware Primitives · IJCAI 2025 QiMeng-GEMM: Automatically Generating High-Performance Matrix Multiplication Code by Exploiting Large Language Models · AAAI 2025 |
Compilers and program optimization › code generation › parallel code generation
GPU kernel generation |
1.0 | 1 | 2026 | QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation · AAAI 2026 |
Compilers and program optimization › code generation
high-performance code generation |
0.9 | 1 | 2025 | QiMeng-GEMM: Automatically Generating High-Performance Matrix Multiplication Code by Exploiting Large Language Models · AAAI 2025 |
Hardware accelerators and domain-specific architectures
accelerator programming models |
0.9 | 1 | 2025 | QiMeng-TensorOp: One-Line Prompt is Enough for High-Performance Tensor Operator Generation with Hardware Primitives · IJCAI 2025 |
High-performance computing › numerical linear algebra
matrix multiplication optimization |
0.9 | 1 | 2025 | QiMeng-GEMM: Automatically Generating High-Performance Matrix Multiplication Code by Exploiting Large Language Models · AAAI 2025 |
High-performance computing
performance optimization |
0.9 | 1 | 2025 | QiMeng-GEMM: Automatically Generating High-Performance Matrix Multiplication Code by Exploiting Large Language Models · AAAI 2025 |
Program synthesis and code generation
code generation with language models |
0.3 | 1 | 2026 | QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation · AAAI 2026 |
GPUs and heterogeneous computing
GPU kernel optimization |
0.3 | 1 | 2026 | QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation · AAAI 2026 |
Machine learning and data management
machine learning systems |
0.3 | 1 | 2025 | QiMeng-TensorOp: One-Line Prompt is Enough for High-Performance Tensor Operator Generation with Hardware Primitives · IJCAI 2025 |
Methods — techniques the papers use, named apart from their topics
large language model · 6.4prompt-based generation · 2.6parameter tuning · 2.6reinforcement learning · 2.0hierarchical strategy-implementation decoupling · 2.0prompt engineering · 1.7meta-prompt search · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel GenerationabstractDeveloping high-performance GPU kernels is critical for AI and scientific computing, but remains challenging due to its reliance on expert crafting and poor portability. While large language models (LLMs) offer promise for automation, both general-purpose and finetuned LLMs suffer from two fundamental and conflicting limitations: correctness and efficiency. The key reason is that existing LLM-based approaches directly generate the entire optimized low-level programs, requiring exploration of an extremely vast space encompassing both optimization policies and implementation codes. To address the challenge of exploring an intractable space, we propose Macro Thinking Micro Coding (MTMC), a hierarchical framework inspired by the staged optimization strategy of human experts. It decouples optimization strategy from implementation details, ensuring efficiency through high-level strategy and correctness through low-level implementation. Specifically, Macro Thinking employs reinforcement learning to guide lightweight LLMs in efficiently exploring and learning semantic optimization strategies that maximize hardware utilization. Micro Coding leverages general-purpose LLMs to incrementally implement the stepwise optimization proposals from Macro Thinking, avoiding full-kernel generation errors. Together, they effectively navigate the vast optimization space and intricate implementation details, enabling LLMs for high-performance GPU kernel generation. Comprehensive results on widely adopted benchmarks demonstrate the superior performance of MTMC on GPU kernel generation in both accuracy and running time. On KernelBench, MTMC achieves near 100% and 70% accuracy at Levels 1-2 and 3, over 50% than SOTA general-purpose and domain-finetuned LLMs, with up to 7.3× speedup over LLMs, and 2.2× over expert-optimized PyTorch Eager kernels. On the more challenging TritonBench, MTMC attains up to 59.64% accuracy and 34× speedup. All models and datasets will be made publicly available. Xinguo Zhu, Shaohui Peng, Jiaming Guo, Yunji Chen, Qi Guo 0001, Yuanbo Wen 0001, Hang Qin, Ruizhi Chen, Qirui Zhou, Ke Gao 0012, Ling Li 0001 |
AAAI | 9 |
| 2025 | QiMeng-GEMM: Automatically Generating High-Performance Matrix Multiplication Code by Exploiting Large Language ModelsabstractAs a crucial operator in numerous scientific and engineering computing applications, the automatic optimization of General Matrix Multiplication (GEMM) with full utilization of ever-evolving hardware architectures (e.g. GPUs and RISC-V) is of paramount importance. While Large Language Models (LLMs) can generate functionally correct code for simple tasks, they have yet to produce high-performance code. The key challenge resides in deeply understanding diverse hardware architectures and crafting prompts that effectively unleash the potential of LLMs to generate high-performance code. In this paper, we propose a novel prompt mechanism called QiMeng-GEMM which enables LLMs to comprehend the architectural characteristics of different hardware platforms and automatically search for the optimization combinations for GEMM. The key of QiMeng-GEMM is a set of informative, adaptive, and iterative meta-prompts. Based on this, a searching strategy for optimal combinations of meta-prompts is used to iteratively generate high-performance code. Extensive experiments conducted on 4 leading LLMs, various paradigmatic hardware platforms, and representative matrix dimensions unequivocally demonstrate QiMeng-GEMM’s superior performance in auto-generating optimized GEMM code. Compared to vanilla prompts, our method achieves a performance enhancement of up to 113×. Even when compared to human experts, our method can reach 115% of cuBLAS on NVIDIA GPUs and 211% of OpenBLAS on RISC-V CPUs. Notably, while human experts often take months to optimize GEMM, our approach reduces the development cost by over 240×. Qirui Zhou, Yuanbo Wen 0001, Ruizhi Chen, Ke Gao 0012, Weiqiang Xiong, Ling Li 0001, Qi Guo 0001, Yunji Chen |
AAAI | 1 |
| 2025 | α-SAV: Generalized Weighted Input Verification for Secure Aggregation in Federated LearningabstractFederated learning has found extensive application in the multimedia domain. However, due to its distributed nature, it is vulnerable to attacks such as Byzantine poisoning. To counteract malicious attacks, the secure aggregation process in federated learning requires input validation from participants. Existing input verification schemes, such as ACORN (USENIX Security 2023), ROFL (S&P 2023), et al., efficiently assess the validity of client inputs, but they fail to account for the impact of weights and do not support weighted secure aggregation. To address these issues, we propose α-SAV, an efficient weighted input verification scheme that utilizes Pedersen commitments to encrypt both privacy and weighted gradients. Our scheme incorporates a non-interactive zero-knowledge proof, the Sigma protocol, allowing clients to generate input proofs without interacting with the server. Verified inputs can then contribute to weighted aggregation. α-SAV is highly compatible, seamlessly integrating into existing federated learning frameworks with minimal additional cost. Experimental results demonstrate that the cost of α-SAV is linear. When trained on the MNIST dataset, the client computation time for α-SAV is 1.6 seconds, resulting in only 24% additional cost compared to ACORN and 3% compared to ROFL. Yuhao Long, Qirui Zhou, Mengyuan Zou, Songfeng Lu |
ICME | 3 |
| 2025 | QiMeng-TensorOp: One-Line Prompt is Enough for High-Performance Tensor Operator Generation with Hardware PrimitivesabstractComputation-intensive tensor operators constitute over 90% of the computations in Large Language Models (LLMs) and Deep Neural Networks. Automatically and efficiently generating high-performance tensor operators with hardware primitives is crucial for diverse and ever-evolving hardware architectures like RISC-V, ARM, and GPUs, as manually optimized implementation takes at least months and lacks portability. LLMs excel at generating high-level language codes, but they struggle to fully comprehend hardware characteristics and produce high-performance tensor operators. We introduce a tensor-operator auto-generation framework with a one-line user prompt (QiMeng-TensorOp), which enables LLMs to automatically exploit hardware characteristics to generate tensor operators with hardware primitives, and tune parameters for optimal performance across diverse hardware. Experimental results on various hardware platforms, SOTA LLMs, and typical tensor operators demonstrate that QiMeng-TensorOp effectively unleashes the computing capability of various hardware platforms, and automatically generates tensor operators of superior performance. Compared with vanilla LLMs, QiMeng-TensorOp achieves up to 1291× performance improvement. Even compared with human experts, QiMeng-TensorOp could reach 251% of OpenBLAS on RISC-V CPUs, and 124% of cuBLAS on NVIDIA GPUs. Additionally, QiMeng-TensorOp also significantly reduces development costs by 200× compared with human experts. Xuzhi Zhang, Shaohui Peng, Qirui Zhou, Yuanbo Wen 0001, Qi Guo 0001, Ruizhi Chen, Xinguo Zhu, Weiqiang Xiong, Haixin Chen, Congying Ma, Ke Gao 0012, Yunji Chen, Ling Li 0001 |
IJCAI | 3 |
| 2025 | Path-MGCN: a pathway activity-based multi-view graph convolutional network for determining spatial domainsabstractSpatial transcriptomics (ST) comprehensively measure the gene expression profiles while preserving the spatial information. Accumulated computational frameworks have been proposed to identify spatial domains, one of the fundamental tasks of ST data analysis, to understand the tissue architecture. However, current methods often overlook pathway-level functional context and struggle with data sparsity. Therefore, we develop Path-MGCN, a multi-view graph convolutional network (GCN) with attention mechanism, which integrates pathway information. We first calculate spot-level pathway activity scores via gene set variation analysis from gene expression and construct distinct adjacency graphs representing spatial and functional proximity. A multi-view GCN learns spatial, pathway, and shared embeddings adaptively fused by attention and followed by a Zero-inflated negative binomial decoder to retain the original transcriptome information. Comprehensive evaluations across diverse datasets (human dorsolateral prefrontal cortex, breast cancer and mouse brain) at various resolution demonstrate Path-MGCN's superior accuracy and robustness, significantly outperforming state-of-the-art methods and maintaining high performance across different pathway databases (Kyoto Encyclopedia of Genes and Genomes, Gene Ontology, Reactome). Crucially, Path-MGCN enhances biological interpretability, enabling the identification of Tertiary lymphoid structure-like regions and spatially resolved metabolic heterogeneity (hypoxia, glycolysis, AMP-activated protein kinase signaling) linked to tumor progression stages in human breast cancer. By effectively integrating functional context, Path-MGCN advances ST analysis, providing an accurate and interpretable framework to dissect tissue heterogeneity and enables detailed spatial mapping of molecular pathways that highlights potential targeted therapeutic strategies crucial for developing safe and effective synergistic anti-tumor therapies. Qirui Zhou, Chaowen Li, Songqing Gu, Weijun Sun, Zongmeng Zhang, Yishan Cai, Chao Yang 0005 |
Briefings Bioinform. | 1 |
| 2020 | Deep learning-based extraction of construction procedural constraints from construction regulations
Botao Zhong, Xuejiao Xing, Hanbin Luo, Qirui Zhou, Heng Li 0001, Timothy M. Rose, Weili Fang |
Adv. Eng. Informatics | 4 |