Mingjie Xing

dblp:118/8977 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 7 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BN-Guard: Batch-Normalization-Guided Activation Bounding for Reliable CNN Inference
Yuting Qian, Mingjie Xing
ICIC (9)2
2026 VecIntrinBench: Benchmarking Cross-Architecture Intrinsic Code Migration for RISC-V Vector
Liutong Han, Chu Kang, Mingjie Xing
ISCAS3
2025 HDCC: A Hierarchical Dataflow-Oriented CGRA Compiler for Complex Applications
abstract
CGRA(Coarse-Grained Reconfigurable Architecture) is characterized by high energy efficiency and reconfigurability, plays an important role in various complex applications. However, current CGRA compilers only handle simple inner loops and struggle with complex nested loops, making it hard to deploy large-scale complex applications on CGRA for acceleration. Therefore, we proposed a hierarchical dataflow compiler HDCC based on MLIR[1] to deal with complex loop structures in layers and developed a tool for generating dataflow graphs from MLIR intermediate code to capture the dataflow of complex loop structures. We also propose a method to deploy large-scale applications to CPU-CGRA with HDCC.The experimental results show that HDCC is capable of deploying large-scale complex applications such as neural network tasks and cryptographic tasks onto CGRA, achieving up to 11.4× performance improvement on these tasks.
Shangli Li, Mingjie Xing
ASP-DAC2
2025 Towards Efficient Compiler Auto-tuning: Leveraging Synergistic Search Spaces
abstract
Determining the optimal sequence of compiler optimization passes is challenging due to the extensive and intricate search space. Traditional auto-tuning techniques, such as iterative compilation and machine learning methods, are often limited by high computational costs and difficulties in generalizing to new programs. These approaches can be inefficient and may not fully address the varying optimization needs across different programs. This paper introduces a novel approach that leverages the synergistic relationships between optimization passes to effectively reduce the search space. By focusing on chained synergy pass pairs that jointly optimize a specific target, our method uses K-means clustering to capture common optimization patterns across programs and forms these pairs into coresets. Leveraging a supervised learning model trained on these coresets, we effectively predict the most beneficial coreset for new programs, streamlining the search for optimal sequences. By integrating various search strategies, our method quickly converges to near-optimal solutions. Our approach achieves state-of-the-art performance on ten benchmark datasets, including MiBench, CBench, NPB, and CHStone, demonstrating an average reduction of 7.5% in Intermediate Representation (IR) instruction count compared to Oz. Furthermore, this set of chained synergy pass pairs is also well-suited for iterative search studies by other researchers, as it enables achieving an average codesize reduction of 13.9% compared to Oz with a simple search strategy that takes only about 5 seconds, outperforming existing search-based techniques in the initial pass search space across five datasets.
Haolin Pan, Yuanyu Wei, Mingjie Xing, Chen Zhao 0024
CGO3
2025 A Method for Co-Design of Compiler and Large-Scale HPC System Architecture
abstract
The field of High-Performance Computing (HPC) is advancing towards large-scale systems, playing a crucial role in scientific research, engineering, and industrial applications by processing vast datasets and solving complex problems. It is essential to explore the design of large-scale HPC system architectures through simulation and to improve the performance of these designs using compilation techniques. To tackle the challenge of optimizing large-scale High-Performance Computing (HPC) system architectures and compiler designs, we propose a method that integrates Multi-Level Intermediate Representation (MLIR) with the Structural Simulation Toolkit (SST). This approach enables the concurrent design and simulation of large-scale HPC architectures and compilers, facilitating early-stage optimization. We extend an MLIR-based cryptographic algorithm compiler and validate the effectiveness of our method through its integration with SST. The example study demonstrates how this method can provide more comprehensive guidance in selecting the scale of high-performance architectures and task partitioning. This collaborative design process enables researchers and engineers to more efficiently develop HPC systems, advancing both compilation techniques and HPC system architectures.
Yuanyu Wei, Mingjie Xing
CSCWD2
2025 Breaking the Fusion Barrier: An Online Algorithm for Fused Normalization and Linear Layers
abstract
The performance of Large Language Model (LLM) inference is critically hindered by memory-bound operations, among which normalization layers are a primary bottleneck, especially during the latency-sensitive decoding phase. While deep learning compilers fuse normalization operations into a single kernel, a fundamental fusion barrier prevents them from merging the ubiquitous Normalization and subsequent Linear layer ($N \& L$) pattern. This barrier, rooted in a core data dependency, forces the execution of two separate kernels, incurring prohibitive kernel launch overhead and costly data round-trips to global memory. In this paper, we break this barrier by introducing FlashFusion, a novel online algorithm that reformulates the$N \& L$pattern to be mathematically equivalent to a single-pass computation. Our key insight is to decompose the computation into a set of independent parallel sums, allowing the normalization statistics to be calculated concurrently with the matrix multiplication, thus eliminating the core data dependency. We co-design a high-performance, hardware-aware GPU kernel that efficiently maps this algorithm to modern architectures, leveraging a tiling strategy to maximize the utilization of Tensor Cores and the memory hierarchy. FlashFusion significantly outperforms state-of-the-art compilers like PyTorch Inductor and TensorRT, achieving speedups of up to$3.03 \times$for the$N \& L$pattern in LLM decoding and effectively eliminating the normalization bottleneck.
Hanghang Cao, Shihao Gao, Quanyi Li, Mingjie Xing
ICPADS5
2025 Exploring the Feasibility of End-to-End Large Language Model as a Compiler
abstract
In recent years, end-to-end Large Language Model (LLM) technology has shown substantial advantages across various domains. As critical system software and infrastructure, compilers are responsible for transforming source code into target code. While LLMs have been leveraged to assist in compiler development and maintenance, their potential as an end-to-end compiler remains largely unexplored. This paper explores the feasibility of LLM as a Compiler (LaaC) and its future directions. We designed the CompilerEval†dataset and framework specifically to evaluate the capabilities of mainstream LLMs in source code comprehension and assembly code generation. In the evaluation, we analyzed various errors, explored multiple methods to improve LLM-generated code, and evaluated cross-platform compilation capabilities. Experimental results demonstrate that LLMs exhibit basic capabilities as compilers but currently achieve low compilation success rates. By optimizing prompts, scaling up the model, and incorporating reasoning methods, the quality of assembly code generated by LLMs can be significantly enhanced. Based on these findings, we maintain an optimistic outlook for LaaC and propose practical architectural designs and future research directions. We believe that with targeted training, knowledge-rich prompts, and specialized infrastructure, LaaC has the potential to generate high-quality assembly code and drive a paradigm shift in the field of compilation.
Shihao Gao, Mingjie Xing
IJCNN4
2025 HybridSIMD: A Super C++ SIMD Library with Integrated Auto-tuning Capabilities
abstract
Single Instruction, Multiple Data (SIMD) technology is crucial for enhancing computational efficiency in High-Performance Computing (HPC). While C++ SIMD libraries abstract away low-level complexities, their proliferation has led to a fragmented set of libraries, creating significant challenges in both performance and usability for developers. To overcome these library-level limitations, this paper introduces a new collaborative concept for SIMD library design. We present HybridSIMD, a C++ library to embody this principle, resolving fragmentation through a unified interface and an operator-level collaborative back-end that leverages the collective strengths of existing libraries. A built-in auto-tuning engine, featuring a hierarchical search strategy, automatically navigates the rich optimization space created by this collaborative approach to deliver maximum performance without manual intervention. Experimental results across six real-world HPC benchmarks on AVX2, AVX512, and NEON architectures demonstrate HybridSIMD’s superiority. Notably, the highest speedups achieved are 185.34× on AVX2, 97.80× on AVX512, and 71.32× on NEON, showcasing its effectiveness in resolving fragmentation while delivering state-of-the-art performance. Our artifact is available at https://github.com/Panhaolin2001/HybridSIMD.
Haolin Pan, Xulin Zhou, Mingjie Xing
ASE3
2025 Compiler-R1: Towards Agentic Compiler Auto-tuning with Reinforcement Learning
abstract
Compiler auto-tuning optimizes pass sequences to improve performance metrics such as Intermediate Representation (IR) instruction count. Although recent advances leveraging Large Language Models (LLMs) have shown promise in automating compiler tuning, two significant challenges still remain: the absence of high-quality reasoning datasets for agents training, and limited effective interactions with the compilation environment. In this work, we introduce Compiler-R1, the first reinforcement learning (RL)-driven framework specifically augmenting LLM capabilities for compiler auto-tuning. Compiler-R1 features a curated, high-quality reasoning dataset and a novel two-stage end-to-end RL training pipeline, enabling efficient environment exploration and learning through an outcome-based reward. Extensive experiments across seven datasets demonstrate Compiler-R1 achieving an average 8.46\% IR instruction count reduction compared to opt -Oz, showcasing the strong potential of RL-trained LLMs for compiler optimization. Our code and datasets are publicly available at https://github.com/Panhaolin2001/Compiler-R1.
Haolin Pan, Kaichun Yao, Libo Zhang 0001, Mingjie Xing
NeurIPS7
2025 MLProf: A Multi-Level Runtime Performance Profiling Framework for MLIR Operations
abstract
Optimizing performance within complex MLIR compilation pipelines necessitates a precise understanding of runtime behavior at the operation level.However, most existing performance analysis tools and methodologies are not specifically designed to accurately capture execution time across MLIR's multi-level hierarchical structure, and they typically lack automated mechanisms for identifying operation-level performance bottlenecks in a developer-friendly manner.Consequently, developers often need to perform manual instrumentation, which is a labor-intensive and error-prone process.In this paper, we present MLProf, a framework specifically designed to support developers in runtime performance analysis of MLIR operations.MLProf employs an automated instrumentation mechanism based on the MLIR infrastructure and an independent runtime data capture method to achieve accurate measurement and finegrained analysis of operation execution time.MLProf provides automated analysis of operation execution time and generates intuitive visualization reports, effectively highlighting runtime bottlenecks for developers.We evaluate MLProf on representative machine learning workloads, demonstrating its capability for multi-level operation analysis, and its efficiency in identifying performance bottlenecks.
Shihao Gao, Hanghang Cao, Mingjie Xing
SEKE5
2025 Navigating the SIMD Optimization Maze: A Reinforcement Learning Approach to Library and Compiler Co-Optimization
abstract
Single Instruction Multiple Data (SIMD) programs are crucial for performance, yet their optimization is complicated by hardware diversity and the varied behaviors of C++ SIMD libraries used to ensure portability.Standard compiler heuristics often struggle with the complex interactions between SIMD library implementations and optimization pass sequences, sometimes even leading to performance degradation compared to scalar code.The vast search space created by the need to cooptimize both SIMD library selection and compiler pass ordering makes manual tuning infeasible.To address this joint optimization challenge, we propose a Reinforcement Learning (RL) based auto-tuning framework specifically for SIMD programs.Our approach introduces three core contributions: (1) A unified C++ template interface seamlessly integrating seven distinct C++ SIMD libraries, facilitating switching between them.(2) An RL agent operating within a joint action space that simultaneously selects both the C++ SIMD libraries and LLVM optimization passes, enabling their co-optimization.(3) An assembly-level program representation designed to capture low-level SIMD characteristics, providing effective guidance for the RL agent.Using LLVM Intermediate Representation(IR) instruction count reduction as a stable proxy metric, experiments on nine SIMD benchmarks demonstrate that our method achieves an average 14.3% improvement over the LLVM opt -Oz baseline.This result highlights the efficacy of our approach in navigating the complex joint optimization space, outperforming traditional auto-tuning techniques.
Haolin Pan, Mingjie Xing
SEKE3
2025 COMPASS: An Agent for MLIR Compilation Pass Pipeline Generation
Shihao Gao, Mingjie Xing
TASE4
2024 Using Context and Hierarchy Features to Enhance Code Completion in Code Modification Scenarios
abstract
Automated software development has always been a research hotspot in the field of software engineering, and code completion can be regarded as a key technology.Currently, many studies on code completion regard the context of prediction as code that strictly appears before the cursor, and have not fully considered the use of code around the cursor for code modification scenarios.Meanwhile, the code serialization method has been developed from source code serialization to Abstract Syntax Tree (AST) serialization.However, AST serialization may lose original structure information.To solve these problems we propose contextual AST serialization (CAS) method, which can more accurately predict AST nodes by capturing the semantic association, context information and structure information between AST nodes.CAS not only converts AST nodes into sequences, but also connects AST sequences with the surrounding context to form richer representations.we conduct empirical studies using a standard Python dataset.The results demonstrate that CAS serialization method significantly improves the code prediction accuracy of the model, the MRR of various types of next token predictions is improved by 15.8% compared to TravTrans on average, and that of type node is improved by 19.4% on average.
Linhai Li, Mingjie Xing
SEKE2
2024 Time-Series based phase-ordering selection for code-size reduction
abstract
Recent advancements in networking, embedded systems, mobile computing, and artificial intelligence (AI) have brought renewed focus on code-size reduction.And within this domain, the sequence of optimization passes, known as phaseordering, considered as a critical determinant.While traditional methods often fall short in addressing complex optimization landscapes, reinforcement learning (RL) algorithms have evolved as a formidable tools to tackle these challenges.According to existing research, previously applied optimization passes can significantly influence subsequent decisions-highlighting a gap in traditional RL algorithms' ability to account for temporal dynamics inherent to the phase-ordering problem.In response to this, our paper proposes a strategy employing Long Short-Term Memory (LSTM) in conjunction with the Proximal Policy Optimization (PPO) RL algorithm, leveraging the temporal relationships between optimization passes to discern an enhanced phase-ordering scheme.The experimental results shows that our model reaches a 3000 iterations fewer in convergence and an additional 5.758% reduction in code size in entire benchmark than PPO alone.Furthermore, our approach also demonstrates robust generalization capabilities and superior performance across various benchmarks and RL algorithms.
Zhenbang Peng, Mingjie Xing
SEKE3
2012 A software memory partition approach for eliminating bank-level interference in multicore systems
abstract
Main memory system is a shared resource in modern multicore machines, resulting in serious interference, which causes performance degradation in terms of throughput slowdown and unfairness. Numerous new memory scheduling algorithms have been proposed to address the interference problem. However, these algorithms usually employ complex scheduling logic and need hardware modification to memory controllers, as a result, industrial venders seem to have some hesitation in adopting them.
Lei Liu 0030, Zehan Cui, Mingjie Xing, Yungang Bao, Mingyu Chen 0001, Chengyong Wu
PACT3