EDBT 2026 Demo / reviewers in the wild / expert
Huimin Cui
dblp:97/2112 · also Hui-Min Cui
· DBLP profile ↗
69ranked-venue papers
9as first author
40since 2021 · last 2026
0000-0002-2491-7679ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 6 first-author · 26 since 2021Software engineering, systems software and programming languages · 20 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Threads to Tiles: T2T, a Compiler for CUDA-to-NPU Translation via 2D VectorizationabstractCUDA’s programming model, exposing massive parallelism via fine-grained scalar threads, has become the de facto standard for GPU computing. Concurrently, NPUs are emerging as highly efficient accelerators, but their architecture is fundamentally different, relying on coarse-grained, explicit 2-D tile-based instructions. This creates a critical challenge: bridging the semantic gap "From Threads to Tiles". A direct translation is infeasible, as it requires lifting the implicit parallelism of CUDA’s scalar model into the explicit, multi-dimensional vector space of NPUs, a problem we formalize as a lifting challenge.This paper introduces T2T, a compiler framework that automates this "Threads to Tiles" translation via the 2-D Vectorization technique. T2T first transforms a CUDA kernel’s implicit SIMT parallelism into a structured, explicit loop nest via our Unified Parallelism Abstraction (UPA), making the parallelism analyzable. From this representation, T2T’s core vectorization engine systematically selects optimal pairs of loops and maps them onto the NPU’s 2-D tile instructions to maximize hardware utilization. To ensure correctness and handle performance-critical CUDA features, a final set of semantics-preserving optimizations is applied, including efficient control-flow management and vectorization of warp-level intrinsics.We implement T2T based on Polygeist and evaluate representative NPU architectures. On a diverse set of benchmarks, kernels translated by T2T achieve up to 73% of native CUDA performance on an A100 GPU and outperform baseline translation approaches by up to 6.9×. Our work demonstrates that a systematic, compiler-driven approach to 2-D vectorization is a principled and high-performance path for porting the rich CUDA ecosystem to the evolving landscape of NPU accelerators. Shuaijiang Li, Ying Liu 0055, Shuoming Zhang, Yijin Li, Yangyu Zhang, Runyu Zhou, Xiyu Shi, Chunwei Xia, Yuan Wen, Xiaobing Feng 0002, Huimin Cui |
CGO | 14 |
| 2026 | Progressive Low-Precision Approximation of Tensor Operators on GPUs: Enabling Greater Trade-Offs between Performance and AccuracyabstractRecent GPUs integrate specialized hardware for low-precision arithmetic (e.g., FP16, INT8), offering substantial speedups for tensor operations. However, existing methods typically rely on coarse, operator-level trial-and-error tuning, which restricts the performance–accuracy trade-off space and limits achievable gains.We present Platensor, a progressive low-precision approximation framework that expands this trade-off space through ne-grained, tile-level strategies. The key idea is to exploit the tiled computation patterns of GPUs to enable flexible precision control and richer optimization opportunities. Platensor performs a two-phase exploration: a fast rule-based pass that selects promising tile-level configurations, followed by an evolutionary search that refines them. It then automatically generates optimized kernels that combine tiles of different precisions.Experiments on GEMM operators and representative applications—including kNN, LLMs, and HPL-MxP—show that Platensor significantly broadens the attainable performance– accuracy trade-offs and more fully leverages low-precision arithmetic on modern GPUs compared to operator-level tuning. Fan Luo 0003, Guangli Li, Zhaoyang Hao, Xueying Wang 0003, Xiaobing Feng 0002, Huimin Cui, Jingling Xue |
CGO | 6 |
| 2026 | DACOS: Dependency-Aware Cross-Kernel Overlapping for Optimizing Short-Sequence Workloads in LLM Applications
Zhaoyang Hao, Guangli Li, Fan Luo 0003, Xueying Wang 0003, Huimin Cui, Jingling Xue |
Euro-Par (2) | 8 |
| 2026 | Symbiotic MLLM Serving: Dynamically Balancing Parallelism Across GPUs and Resources Within GPUs
Yangyu Zhang, Zhaolin Duan, Shuoming Zhang, Shuaijiang Li, Donglin Yu, Yuan Wen, Chunwei Xia, Xiyu Shi, Huimin Cui |
ISCA | 13 |
| 2026 | A Multi-Modal Retrieval-Augmented Framework for Compiler Backend Generation with LLMs
Ming Zhong 0016, Hongna Geng, Lulin Wang, Lei Qiu 0007, Huimin Cui, Xiaobing Feng 0002 |
SANER | 7 |
| 2026 | LEGO-compiler: enhancing neural compilation through translation composability
Shuoming Zhang, Qiuchu Yu, Chunwei Xia, Zheng Wang 0001, Yunji Chen, Xiaobing Feng 0002, Huimin Cui |
CCF Trans. High Perform. Comput. | 8 |
| 2026 | The new compiler stack: a survey on the synergy of LLMs and compilers
Shuoming Zhang, Qiuchu Yu, Chunwei Xia, Zheng Wang 0001, Xiaobing Feng 0002, Huimin Cui |
CCF Trans. High Perform. Comput. | 7 |
| 2026 | SYCL-MLU: unifying SIMT and SIMD in heterogeneous programming
Runyu Zhou, Yijin Li, En Shao, Ziyan Xie, Huimin Cui |
CCF Trans. High Perform. Comput. | 7 |
| 2026 | SparseZETA: Intelligent Auto-tuner for Designing High-Performance SpMV ProgramsabstractSparse matrix-vector multiplication (SpMV) is a crucial operation in scientific computing, graph analytics, and machine/deep learning. Its performance is highly sensitive to matrix sparsity patterns, necessitating tailored program designs. This paper introduces SparseZETA, an intelligent auto-tuner that generates high-performance, machine-designed SpMV programs by directly mimicking and composing human-expert actions. To efficiently navigate the vast design space, SparseZETA reformulates auto-tuning as a behavior-cloning problem: rather than costly exploration, it directly synthesizes programs by sequentially predicting actions in a one-pass decision-making process, guided by the real-time state of the evolving, partially constructed program designs. A novel self-training mechanism further accelerates the collection of training data for the prediction models. On NVIDIA A100 (and RTX 2080 Ti) GPUs, SparseZETA achieves average speedups of 1.27×–15.66× (1.44×–19.07×) over existing auto-tuners, human-designed programs, and a sparse compiler. SparseZETA substantially reduces the human effort required to design SpMV programs, including sparse format creation and kernel implementation, cutting the design time from days or even months to an average of 82.52ms per matrix via lightweight inference on only one CPU. Zhen Du, Ying Liu 0055, Xionghui Chen, Xiaobing Feng 0002, Huimin Cui, Jiajia Li 0001 |
Proc. ACM Program. Lang. | 6 |
| 2026 | MoonPoly: Bridging Code Generation and Adaptive Execution via Micro-Kernel Polymerization for Optimizing Dynamic-Shape Tensor OperatorsabstractThe prevalence of dynamic tensor shapes, driven by applications like language model serving with varying sequence lengths, is a defining characteristic of modern deep neural networks. This dynamism poses a fundamental challenge: reconciling the need for intensive, offline code generation to achieve peak performance with the demand for low-latency, adaptive execution to handle unpredictable runtime tensor shapes. Consequently, mainstream strategies are ineffective. Vendor-provided libraries, while highly optimized for a subset of common shapes, suffer performance degradation on unconventional ones. Static tensor compilers are hamstrung by prohibitive just-in-time compilation overheads for each new shape. While recent dynamic-shape compilers offer an alternative, they rely on predefined shape ranges, making them brittle when inputs fall outside these bounds. To resolve this tension, we present MoonPoly , a dynamic-shape tensor compiler that introduces micro-kernel polymerization . Our approach decouples these conflicting requirements through a two-stage process. In the offline stage, it performs intensive auto-tuning to generate a set of micro-kernels and corresponding performance models. The online stage then performs adaptive execution, rapidly assembling a near-optimal tensor operator on-the-fly, guided by a lightweight cost model. Evaluated on an NVIDIA A100 GPU, MoonPoly achieves an average operator-level speedup of 1.27× over the cuBLAS library across a diverse set of operators and data types, which in turn yields end-to-end inference acceleration for a variety of models, including BERT, the Vision Transformer, and large language models. Yangyu Zhang, Guangli Li, Feng Yu 0019, Fan Luo 0003, Qianqi Sun, Xueying Wang 0003, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
ACM Trans. Archit. Code Optim. | 8 |
| 2026 | BePilot: An AI Programming Assistant for Compiler Backend DevelopmentabstractCompiler backends are tasked with generating executable machine code for various processors. As the diversity of processors continues to grow, it is imperative for programmers to tailor specific compiler backends to accommodate each one. However, compiler backend development remains a labor-intensive and time-consuming process, with limited automation tools available. Although large language models (LLMs) have demonstrated strong abilities in code completion and code generation tasks, the lack of appropriate datasets for compiler backend development limits the application of LLMs in this field. In this article, we introduce ComBack++, a multilingual dataset covering C/C++, machine description, and TableGen, with 184 backends from GCC and LLVM, four backend-specific tasks. Based on ComBack++, we present BePilot, a compiler backend-specific LLM available in two sizes: BePilot-1.5B and BePilot-7B. We also introduce CB-Retriever , a retriever that constructs few-shot prompts via in-context learning to improve vanilla LLM performance in resource-constrained settings. Experimental results show that BePilot-1.5B and BePilot-7B achieve significantly higher accuracy across four tasks in ComBack++ compared to 12 baseline LLMs (125M–34B parameters). In addition, CB-Retriever consistently boosts the accuracy of six mainstream LLMs. Both BePilot-1.5B and BePilot-7B, as well as vanilla LLMs augmented with CB-Retriever , outperform the traditional manual compiler backend development approach (Fork-Flow) in efficiency across all four tasks in ComBack++. Furthermore, human evaluation by four experienced compiler backend developers confirms that BePilot not only improves development efficiency over Fork-Flow but also surpasses commercial AI programming assistants such as GPT-4o-mini and Gemini2-Flash in terms of code quality. These findings confirm that BePilot and CB-Retriever can substantially enhance compiler backend development efficiency. Ming Zhong 0016, Lulin Wang, Hongna Geng, Lei Qiu 0007, Huimin Cui, Xiaobing Feng 0002 |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2025 | Qiwu: Exploiting Ciphertext-Level SIMD Parallelism in Homomorphic Encryption ProgramsabstractFully Homomorphic Encryption (FHE), particularly the CKKS scheme, enables computation on encrypted data, facilitating secure task offloading to untrusted servers. CKKS allows packing multiple complex values into a single ciphertext, crucial for fixed-point arithmetic in machine learning, while leveraging SIMD parallelism at the plaintext level. However, operations such as reductions can degrade performance by creating a large number of bubbles (or gaps) in intermediate ciphertexts, leading to wasted computational resources. We introduce Qiwu, a ciphertext-level vectorization approach that enhances performance by fusing multiple ciphertexts containing bubbles. Qiwu uses a DSL to specify zero bubbles in input ciphertexts and nonzero bubbles in output ciphertexts, employs data-flow analysis to track them, and formulates a fusion plan guided by a cost-benefit assessment. Implemented in an existing FHE compiler, Qiwu was evaluated on four applications (including three machine learning tasks) and three kernels. It achieves speedups of up to 18.0× on CPUs, averaging 3.4× (geometric mean), compared to the state-of-the-art compiler that exploit only plaintext-level parallelism. Zhongcheng Zhang, Ying Liu 0055, Zhenchuan Chen, Xiaobing Feng 0002, Huimin Cui, Jingling Xue |
CGO | 7 |
| 2025 | VEGA: Automatically Generating Compiler Backends using a Pre-trained Transformer ModelabstractWe introduce VEGA, an AI-driven system aimed at easing the development of compiler backends for new targets. Our approach involves categorizing functions from existing backends into function groups, each comprising various target-specific implementations of a standard compiler interface function, abstracted as a single function template. Therefore, generating a new backend involves customizing these function templates to specific target requirements. To capitalize on AI's capabilities in code generation, VEGA maps statements in a target-specific version of a function template into feature vectors, distinguishing between target-independent and target-specific properties. Leveraging a pre-trained model, VEGA can efficiently auto-generate a version of each function template tailored to a specific target, thereby enabling the construction of a complete compiler backend for a new target based solely on its target description files. We evaluated VEGA on three distinct targets: a CPU processor (RISC-V), a customized processor with instruction extensions (RI5CY), and an IoT processor (xCORE). VEGA demonstrated high efficiency, generating compiler backends under an hour, which can substantially enhance developer productivity. Across the three targets, VEGA achieved accuracy rates of 71.5%, 73.2%, and 62.2% for all generated functions, significantly outperforming the traditional fork-flow method, which yielded less than 8% accuracy. Moreover, VEGA provides explicit confidence scores for generated functions and statements, allowing developers to easily identify areas requiring minimal manual intervention. This research has the potential to improve the effectiveness of traditional compiler backend development. Ming Zhong 0016, Lulin Wang, Lei Qiu 0007, Ying Liu 0055, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
CGO | 7 |
| 2025 | TopServe: Task-Operator Co-scheduling for Efficient Multi-DNN Inference Serving on GPUs
Guangli Li, Feng Yu 0019, Xueying Wang 0003, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
Euro-Par (2) | 6 |
| 2025 | Light-FP: Analyze Floating-Point Error in a Highly Condensed ApproachabstractApproximate computing is emerging as a promising paradigm of High-Performance Computing (HPC) to increase application performance, with mixed-precision computing Jiazhi Mi, Ruixiang Gao, Ronghong Shen, You Fu, Huimin Cui |
ICS | 9 |
| 2025 | SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMsabstractRecent multimodal large language models (MLLMs) marry modality-specific
vision or audio encoders with a shared text decoder. While the encoder is compute-
intensive but memory-light, the decoder is the opposite, yet state-of-the-art serving
stacks still time-multiplex these complementary kernels, idling SMs or HBM in
turn. We introduce SpaceServe, a serving system that space-multiplexes MLLMs:
it decouples all modality encoders from the decoder, and co-locates them on the
same GPU using fine-grained SM partitioning available in modern runtimes. A
cost-model-guided Space-Inference Scheduler (SIS) dynamically assigns SM slices,
while a Time-Windowed Shortest-Remaining-First (TWSRFT) policy batches en-
coder requests to minimise completion latency and smooth decoder arrivals.
Evaluation shows that SpaceServe reduces time-per-output-token by 4.81×
on average and up to 28.9× on Nvidia A100 GPUs. SpaceServe is available at
https://github.com/gofreelee/SpaceServe Shuoming Zhang, Xiyu Shi, Yangyu Zhang, Shuaijiang Li, Donglin Yu, Zheming Yang, Yuan Wen, Huimin Cui |
NeurIPS | 11 |
| 2025 | IR-OptSet: An Optimization-Sensitive Dataset for Advancing LLM-Based IR OptimizerabstractCompiler optimization is essential for improving program performance, yet modern compilers still depend on manually crafted transformation rules over intermediate representations (IRs). As compilers grow in complexity, maintaining these rule-based optimizations becomes increasingly labor-intensive and difficult to scale. Recent advances in large language models (LLMs) offer a promising alternative, but their effectiveness in compiler optimization remains limited—primarily due to the lack of IR-oriented datasets that expose models to diverse transformation samples in real-world scenarios (optimization-sensitive samples), hindering LLMs from learning rich and generalizable optimization strategies.In this paper, we introduce IR-OptSet, the first public optimization-sensitive dataset for advancing LLM-based IR optimizers. It comprises 170K LLVM IR samples from open-source repositories across 8 representative optimization domains. IR-OptSet defines two core tasks: Code Analysis and Optimized Code Generation, and provides tools for correctness verification, performance evaluation, and dataset expansion. In our experiments, fine-tuning three representative LLMs on IR-OptSet leads to significant accuracy improvements across both tasks. Moreover, the LLM fine-tuned with IR-OptSet outperforms traditional compiler with the -O3 option in 64 test cases in terms of performance. Further analysis reveals that IR-OptSet provides greater transformation diversity and representativeness than three widely used IR-oriented datasets, highlighting its potential to drive model-based IR optimization. IR-OptSet is publicly available at https://huggingface.co/datasets/YangziResearch/IR-OptSet. Lei Qiu 0007, Fang Lyu, Ming Zhong 0016, ZhiLei Chai, Haojie Zhou, Huimin Cui, Xiaobing Feng 0002 |
NeurIPS | 7 |
| 2025 | Beehive: A Scalable Disaggregated Memory Runtime Exploiting Asynchrony of Multithreaded Programs
Quanxi Li, Ying Liu 0055, Yanwen Xia, Jie Zhang 0048, Mosong Zhou, Xiaobing Feng 0002, Huimin Cui, Yizhou Shan, Chenxi Wang 0005 |
NSDI | 8 |
| 2025 | TensorMD: Molecular Dynamics Simulation with Ab Initio Accuracy of 50 Billion AtomsabstractMolecular dynamics simulation emerges as an important area that HPC+AI helps to investigate the physical properties, with machine-learning interatomic potentials (MLIPs) being used. General-purpose machine-learning (ML) tools have been leveraged in MLIPs, but they are not perfectly matched with each other, since many optimization opportunities in MLIPs have been missed by ML tools. This inefficiency arises from the fact that HPC+AI applications work with far more computational complexity compared with pure AI scenarios. This paper has developed an MLIP, named TensorMD, independently from any ML tool. TensorMD has been evaluated on two supercomputers and scaled to 51.8 billion atoms, i.e., ~ 3× compared with state-of-the-art. Yucheng Ouyang, Ying Liu 0055, Honghui Shang, Zhenchuan Chen, Jiahao Shan, Huimin Cui, Xiaobing Feng 0002, Xingyu Gao 0003, Haifeng Song 0003, Xin Chen 0023, Rongfen Lin |
PPoPP | 6 |
| 2025 | Boosting Large Language Models for System Software Retargeting: A Preliminary StudyabstractSystem software bridges hardware platforms and high-level applications. As new hardware platforms emerge, developers must customize code to support various system software, a process known as “retargeting”. This process is time-consuming and poorly automated. While large language models (LLMs) are proficient in general code generation tasks, their effectiveness in retargeting is limited by code complexity and abstract function descriptions. This paper presents TeSyn, a novel framework to enhance the code generation capabilities for system software retargeting. TeSyn comprises three steps: target-specific value extraction, common code clustering, and template synthesis. To evaluate TeSyn's effectiveness, we intro-duce SysRetar, the first dataset for system software retargeting, covering four types of system software and 195 hardware platforms. In our experiments, we select five LLMs and fine-tune CodeLLaMA-7B-Instruct on SysRetar to create SysRetar-LLM. Results show that TeSyn significantly enhances retargeting performance across five LLMs. Furthermore, code generated by SysRetar- LLM requires substantially less modification than the manual retargeting approach (Fork-Flow), suggesting potential improvements in efficiency. Given these promising results, we outline future research directions for advancing retargeting through LLMs. The dataset and code are publicly available at https://huggingface.co/doczll05/SysRetar-LLM. Ming Zhong 0016, Lulin Wang, Lei Qiu 0007, Hongna Geng, Huimin Cui, Xiaobing Feng 0002 |
SANER | 6 |
| 2025 | TENSORMD: Accelerating Molecular Dynamics with a High-Performance Machine Learning Interatomic PotentialabstractAI has been integrated into HPC across various scientific fields, significantly enhancing performance. In molecular dynamics simulations, HPC+AI facilitates the investigation of atomic-scale physical properties using machine-learning interatomic potentials (MLIPs). However, general-purpose ML tools (e.g., TensorFlow) used in MLIPs are not optimally matched, leading to missed optimization opportunities due to the higher computational complexity and greater diversity of HPC+AI applications compared to pure AI scenarios. To address this, we introduce TensorMD, an MLIP independent of existing ML tools, enabling flexible optimizations that standard ML frameworks cannot support. TensorMD outperforms a state-of-the-art MLIP—winner of the 2020 Gordon Bell Prize and built on an ML tool—by 1.88 × on NVIDIA A100 GPU. Additionally, TensorMD was evaluated on two supercomputers with different architectures, achieving significantly reduced time-to-solution and supporting molecular dynamics simulations at scales beyond 50 billion atoms. Yucheng Ouyang, Ying Liu 0055, Xin Chen 0023, Honghui Shang, Zhenchuan Chen, Rongfen Lin, Xingyu Gao 0003, Jiahao Shan, Haifeng Song 0003, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
SC | 13 |
| 2025 | Orthrus: Efficient and Timely Detection of Silent User Data Corruption in the Cloud with Resource-Adaptive Computation ValidationabstractEven with substantial endeavors to test and validate processors, computational errors may still arise post-installation. One particular category of CPU errors transpires discreetly, without crashing applications or triggering hardware warnings. These elusive errors pose a significant threat by undermining user data, and their detection is challenging. This paper introduces Orthrus, a solution for the timely detection of silent user data corruption caused by post-installation CPU errors. Orthrus safeguards user data in cloud applications by providing simple annotations and compiler support for users to identify data operators and validating these operators asynchronously across cores while maintaining a low overhead (2%–6%), making it practical for production deployment. Our evaluation, using carefully injected errors, demonstrates that Orthrus can detect 87% of data corruptions with just a single core dedicated to validation, increasing to 91% and 96% when two and four cores are used, respectively. Chenxiao Liu, Zhenting Zhu, Quanxi Li, Yanwen Xia, Yifan Qiao 0002, Xiangyun Deng, Youyou Lu, Tao Xie 0001, Huimin Cui, Zidong Du, Guoqing Harry Xu, Chenxi Wang 0005 |
SOSP | 9 |
| 2025 | Scalable tasking runtime with parallelized builders for explicit message passing architectures
Xiran Gao, Huimin Cui, Xiaobing Feng 0002 |
Parallel Comput. | 4 |
| 2025 | SRSparse: Generating Codes for High-Performance Sparse Matrix-Vector Semiring ComputationsabstractSparse matrix-vector semiring computation is a key operation in sparse matrix computations, with performance strongly dependent on both program design and the features of the sparse matrices. Given the diversity of sparse matrices, designing a tailored program for each matrix is challenging. To address this, we propose SRSparse, 1 a program generator that creates tailored programs by automatically combining program designing methods to fit specific input matrices. It provides two components: the problem definition configuration , which declares the computation, and the scheduling language , which can be leveraged by an auto-tuner to specify the program designs. The two are lowered to the intermediate representations of SRSparse, the Format IR and Kernel IR , which respectively generate format conversion routine and kernel code. We evaluate SRSparse on four representative sparse kernels and three format conversion routines. For sparse kernels, SRSparse achieves median speedups over handwritten programs: COO (3.50×), CSR-Adaptive (5.36×), CSR5 (2.06×), ELL (1.63×), Gunrock (1.57×), and GraphBLAST (1.96×); over an auto-tuner: AlphaSparse (1.16×); and over a compiler: TACO (1.71×). For format conversion routines, SRSparse achieves median speedups over handwritten implementations: Intel MKL (7.60×), SPARSKIT (2.61×), CUSP (2.77×), and Ginkgo (1.74×); and over a compiler: TACO (4.04×). Zhen Du, Ying Liu 0055, Ninghui Sun, Huimin Cui, Xiaobing Feng 0002, Jiajia Li 0001 |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | OptiFX: Automatic Optimization for Convolutional Neural Networks with Aggressive Operator Fusion on GPUsabstractConvolutional Neural Networks (CNNs) are fundamental to advancing computer vision technologies. As CNNs become more complex and larger, optimizing model inference remains a critical challenge in both industry and academia. On modern GPU platforms, CNN operators are typically memory-bound, leading to significant performance degradation due to memory wall effects. While recent advancements have utilized operator fusion–merging multiple operators into one–to enhance inference performance, the fusion of multiple region-based operators like convolution is seldom addressed. This article introduces AFusion , a novel operator fusion technique aimed at improving inference performance, and OptiFX, an automatic optimization framework based on this approach. OptiFX employs a cost-based backtracking search to identify optimal sub-graphs for fusion and utilizes template-based code generation to create efficient kernels for these fused sub-graphs. We evaluate OptiFX across seven prominent CNN architectures–GoogLeNet, ResNet, DenseNet, MobileNet, SqueezeNet, NasNet, and UNet–on Nvidia A6000 Ada, RTX 4090, and Jetson AGX Orin platforms. Our results demonstrate that OptiFX significantly outperforms existing methods, achieving average speedups of \(2.91\times\) , \(3.30\times\) , and \(2.09\times\) in accelerating inference performance on these platforms, respectively. Xueying Wang 0003, Shigang Li 0002, Fan Luo 0003, Zhaoyang Hao, Tong Wu 0024, Ruiyuan Xu, Huimin Cui, Xiaobing Feng 0002, Guangli Li, Jingling Xue |
ACM Trans. Archit. Code Optim. | 8 |
| 2024 | Optimizing Deep Learning Inference via Global Analysis and Tensor ExpressionsabstractOptimizing deep neural network (DNN) execution is important but becomes increasingly difficult as DNN complexity grows. Existing DNN compilers cannot effectively exploit optimization opportunities across operator boundaries, leaving room for improvement. To address this challenge, we present Souffle, an open-source compiler that optimizes DNN inference across operator boundaries. Souffle creates a global tensor dependency graph using tensor expressions, traces data flow and tensor information, and partitions the computation graph into subprograms based on dataflow analysis and resource constraints. Within a subprogram, Souffle performs local optimization via semantic-preserving transformations, finds an optimized program schedule, and improves instruction-level parallelism and data reuse. We evaluated Souffle using six representative DNN models on an NVIDIA A100 GPU. Experimental results show that Souffle consistently outperforms six state-of-the-art DNN optimizers by delivering a geometric mean speedup of up to 3.7× over TensorRT and 7.8× over Tensorflow XLA. Chunwei Xia, Qianqi Sun, Zheng Wang 0001, Yuan Wen, Xiaobing Feng 0002, Huimin Cui |
ASPLOS (1) | 8 |
| 2024 | Optimizing Dynamic-Shape Neural Networks on Accelerators via On-the-Fly Micro-Kernel PolymerizationabstractIn recent times, dynamic-shape neural networks have gained widespread usage in intelligent applications to address complex tasks, introducing challenges in optimizing tensor programs due to their dynamic nature. As the operators' shapes are determined at runtime in dynamic scenarios, the compilation process becomes expensive, limiting the practicality of existing static-shape tensor compilers. To address the need for effective and efficient optimization of dynamic-shape neural networks, this paper introduces MikPoly, a novel dynamic-shape tensor compiler based on micro-kernel polymerization. MikPoly employs a two-stage optimization approach, dynamically combining multiple statically generated micro-kernels using a lightweight cost model based on the shape of a tensor operator known at runtime. We evaluate the effectiveness of MikPoly by employing popular dynamic-shape operators and neural networks on two representative accelerators, namely GPU Tensor Cores and Ascend NPUs. Our experimental results demonstrate that MikPoly effectively optimizes dynamic-shape workloads, yielding an average performance improvement of 1.49× over state-of-the-art vendor libraries. Feng Yu 0019, Guangli Li, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
ASPLOS (2) | 4 |
| 2024 | ComBack: A Versatile Dataset for Enhancing Compiler Backend Development EfficiencyabstractCompiler backends are tasked with generating executable machine code for processors. With the proliferation of diverse processors, it is imperative for programmers to tailor specific compiler backends to accommodate each one. Meanwhile, compiler backend development is a laborious and time-consuming task, lacking effective automation methods. Although language models have demonstrated strong abilities in code related tasks, the lack of appropriate datasets for compiler backend development limits the application of language models in this field.In this paper, we introduce ComBack, the first public dataset designed for improving compiler backend development capabilities of language models. ComBack includes 178 backends for mainstream compilers and three tasks including statement-level completion, next-statement suggestion and code generation, representing common development scenarios. We conducted experiments by fine-tuning six pre-trained language models with ComBack, demonstrating its effectiveness in enhancing model accuracy across the three tasks. We further evaluated the top-performing model(CodeT5+) across the three tasks for new targets, comparing its accuracy with conventional methods (Fork-Flow), ChatGPT-3.5-Turbo, and Code-LLaMA-34B-Instruct. Remarkably, fine-tuned CodeT5+ with only 220M parameters on ComBack outperformed Fork-Flow methods significantly and surpassed ChatGPT and Code-LLaMA. This suggests potential efficiency improvements in compiler development. ComBack is avaliable at https://huggingface.co/datasets/docz1105/ComBack. Ming Zhong 0016, Fang Lyu, Lulin Wang, Hongna Geng, Lei Qiu 0007, Huimin Cui, Xiaobing Feng 0002 |
NeurIPS | 6 |
| 2024 | A Tale of Two Paths: Toward a Hybrid Data Plane for Efficient Far-Memory Applications
Chenxi Wang 0005, Yifan Qiao 0002, Zhe Wang 0017, Chenggang Wu 0002, Youyou Lu, Xiaobing Feng 0002, Huimin Cui, Shan Lu 0001, Guoqing Harry Xu |
OSDI | 10 |
| 2024 | Pushing the Limit of Quantum Mechanical Simulation to the Raman Spectra of a Biological System with 100 Million AtomsabstractRaman spectroscopy offers invaluable insights into the chemical composition and structural characteristics of various materials, making it a powerful tool for structural analysis. However, accurate quantum mechanical simulations of Raman spectra for large systems, such as biological materials, have been limited due to immense computational costs and technical challenges. In this study, we developed efficient algorithms and optimized implementations on heterogeneous computing architectures to enable fast and highly scalable ab initio simulations of Raman spectra for large-scale biological systems with up to 100 million atoms. Our simulations have achieved nearly linear strong and weak scaling on two cutting-edge high-performance computing systems, with peak FP64 performances reaching 400 PFLOPS on 96,000 nodes of new Sunway supercomputer and 85 PFLOPS on 6,000 node of ORISE supercomputer. These advances provide promising prospects for extending quantum mechanical simulations to biological systems. Honghui Shang, Ying Liu 0055, Zhikun Wu, Zhenchuan Chen, Jinfeng Liu 0004, Meiyue Shao, Yingzhou Li, Bowen Kan, Huimin Cui, Xiaobing Feng 0002, Yunquan Zhang, Donald G. Truhlar, Hong An, Xiao He 0004, Jinlong Yang 0003 |
SC | 9 |
| 2023 | Occamy: Elastically Sharing a SIMD Co-processor across Multiple CPU CoresabstractSIMD extensions are widely adopted in multi-core processors to exploit data-level parallelism. However, when co-running workloads on different cores, compute-intensive workloads cannot take advantage of the underutilized SIMD lanes allocated to memoryintensive workloads, reducing the overall performance. This paper proposes Occamy, a SIMD co-processor that can be shared by multiple CPU cores, so that their co-running workloads can spatially share its SIMD lanes. The key idea is to enable elastic spatial sharing by dynamically partitioning all the SIMD lanes across different workloads based on their phase behaviors, so that each workload may execute in variable-length SIMD mode. We also introduce an Occamy compiler to support such variable-length vectorization by analyzing such phase behaviors and generating the vectorized code that works with varying vector lengths. We demonstrate that Occamy can improve SIMD utilization, and consequently, performance over three representative SIMD architectures, with negligible chip area cost. Zhongcheng Zhang, Yan Ou, Ying Liu 0055, Chenxi Wang 0005, Yongbin Zhou, Yucheng Ouyang, Jiahao Shan, Ying Wang 0001, Jingling Xue, Huimin Cui, Xiaobing Feng 0002 |
ASPLOS (3) | 12 |
| 2023 | Honeycomb: Secure and Efficient GPU Executions via Static Validation
Haohui Mai, Hongren Zheng, Zibin Liu, Mingyu Gao 0001, Huimin Cui, Xiaobing Feng 0002, Christoforos E. Kozyrakis |
OSDI | 8 |
| 2023 | Portable and Scalable All-Electron Quantum Perturbation Simulations on Exascale SupercomputersabstractQuantum perturbation theory is pivotal in determining the critical physical properties of materials. The first-principles computations of these properties have yielded profound and quantitative insights in diverse domains of chemistry and physics. In this work, we propose a portable and scalable OpenCL implementation for quantum perturbation theory, which can be generalized across various high-performance computing (HPC) systems. Optimal portability is realized through the utilization of a cross-platform unified interface and a collection of performance-portable heterogeneous optimizations. Exceptional scalability is attained by addressing major constraints on memory and communication, employing a locality-enhancing task mapping strategy and a packed hierarchical collective communication scheme. Experiments on two advanced supercomputers demonstrate that our implementation exhibits remarkably performance on various material systems, scaling the system to 200,000 atoms with all-electron precision. This research enables all-electron quantum perturbation simulations on substantially larger molecular scales, with a potentially significant impact on progress in material sciences. Zhikun Wu, Yangjun Wu, Ying Liu 0055, Honghui Shang, Yingxiang Gao, Zhongcheng Zhang, Yingchi Long, Xiaobing Feng 0002, Huimin Cui |
SC | 10 |
| 2023 | Automatic Target Description File Generation
Hongna Geng, Fang Lyu, Ming Zhong 0016, Huimin Cui, Jingling Xue, Xiaobing Feng 0002 |
J. Comput. Sci. Technol. | 4 |
| 2023 | Reinvent Cloud Software Stacks for Resource Disaggregation
Chenxi Wang 0005, Yi-Zhou Shan, Pengfei Zuo, Huimin Cui |
J. Comput. Sci. Technol. | 4 |
| 2023 | VTensor: Using Virtual Tensors to Build a Layout-Oblivious AI Programming Framework
Feng Yu 0019, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
J. Comput. Sci. Technol. | 3 |
| 2022 | Scaling Poisson Solvers on Many Cores via MMEwaldabstractThe Poisson solver for the calculation of the electrostatic potential is an essential primitive in quantum mechanics calculations. In this article, we adopt the Ewald method and propose a highly-optimized and scalable framework for Poisson solver, MMEwald, on the new generation Sunway supercomputer, capable of utilizing the collection of 390-core accelerators it uses. The MMEwald is based on a grid adapted cut-plane approach to partition the points into batches and distribute the batch to the processors. Furthermore, we propose a set of architecture-specific optimizations to efficiently utilize the memory bandwidth and computation capacity of the supercomputer. Experimental results demonstrate the efficiency of the MMEwald in providing strong and weak scaling performance. Mingchuan Wu, Yangjun Wu, Honghui Shang, Ying Liu 0055, Huimin Cui, Xiaohui Duan, Yunquan Zhang, Xiaobing Feng 0002 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | NRHI: A Concurrent Non-Rehashing Hash Index for Persistent MemoryabstractPersistent memory (PM) featured with data persistence, byte-addressability, and DRAM-like performance has been commercially available with the advent of Intel®Optane™ DC persistent memory. The DRAM-like performance and disk-like persistence invite shifting hashing-based index schemes, which are important building blocks of today’s internet service infrastructures to provide fast queries, from DRAM onto persistent memory. Numerous hash indexes for persistent memory have been proposed to optimize writes and crash consistency, but with poor scalability under resizing. Generally, resizing consists of allocating a new hash table and rehashing items from the old table into the new one. We argue that resizing with rehashing performed in either blocking or non-blocking way can degrade the overall performance and limit the scalability.In order to mitigate the limitation of resizing, this paper proposes a Non-Rehashing Hash Index (NRHI) scheme to perform resizing with no necessity of rehashing items. NRHI leverages a layered structure to link hash tables without moving key-value pairs across layers, thus reducing the time spent on rehashing in blocking way and alleviating slots contention occurred in non-blocking way. Furthermore, the compare-and-swap primitive is utilized to support concurrent lock-free hashing operations. Experimental results on real PM hardware show that NRHI outperforms the state-of-the-art PM hash indexes by 1.7× to 3.59×, and scales linearly with the number of threads. Huimin Cui, Lei Liu 0030 |
ICCD | 2 |
| 2021 | Accelerating all-electron ab initio simulation of raman spectra for biological systemsabstractRaman spectroscopy provides chemical and compositional information that can serve as a structural fingerprint for various materials. Therefore, simulations of Raman spectra, including both quantum perturbation analyses and ground-state calculations are of significant interest. However, highly accurate full quantum mechanical (QM) simulations of Raman spectra have previously been confined to small systems. For large systems such as biological materials, the computational cost of full QM simulations is extremely high, and their extension to such systems remains challenging. In the work described here, by employing robust new algorithms and advances in implementation for the many-core architectures, we are able to perform fast, accurate, and massively parallel full ab initio simulations of the Raman spectra of biological systems with excellent strong and weak scaling, thereby providing a starting point for applying QM approaches to structural studies of such systems. Honghui Shang, Yunquan Zhang, Ying Liu 0055, Mingchuan Wu, Yangjun Wu, Di Wei, Huimin Cui, Xin Liu 0081, Fei Wang 0096, Yuxi Ye, Yingxiang Gao, Shuang Ni, Xin Chen 0023, Dexun Chen |
SC | 9 |
| 2021 | Unified Holistic Memory Management Supporting Multiple Big Data Processing Frameworks over Hybrid MemoriesabstractTo process real-world datasets, modern data-parallel systems often require extremely large amounts of memory, which are both costly and energy inefficient. Emerging non-volatile memory (NVM) technologies offer high capacity compared to DRAM and low energy compared to SSDs. Hence, NVMs have the potential to fundamentally change the dichotomy between DRAM and durable storage in Big Data processing. However, most Big Data applications are written in managed languages and executed on top of a managed runtime that already performs various dimensions of memory management. Supporting hybrid physical memories adds a new dimension, creating unique challenges in data replacement. This article proposes Panthera, a semantics-aware, fully automated memory management technique for Big Data processing over hybrid memories. Panthera analyzes user programs on a Big Data system to infer their coarse-grained access patterns, which are then passed to the Panthera runtime for efficient data placement and migration. For Big Data applications, the coarse-grained data division information is accurate enough to guide the GC for data layout, which hardly incurs overhead in data monitoring and moving. We implemented Panthera in OpenJDK and Apache Spark. Based on Big Data applications’ memory access pattern, we also implemented a new profiling-guided optimization strategy, which is transparent to applications. With this optimization, our extensive evaluation demonstrates that Panthera reduces energy by 32–53% at less than 1% time overhead on average. To show Panthera’s applicability, we extend it to QuickCached, a pure Java implementation of Memcached. Our evaluation results show that Panthera reduces energy by 28.7% at 5.2% time overhead on average. Chenxi Wang 0005, John N. Zigman, Haris Volos 0001, Onur Mutlu, Xiaobing Feng 0002, Guoqing Harry Xu, Huimin Cui |
ACM Trans. Comput. Syst. | 11 |
| 2020 | Bandwidth-Aware Loop Tiling for DMA-Supported Scratchpad MemoryabstractScratchpad Memory (SPM) is widely used in emerging domain-specific architectures and accelerators for improving energy efficiency and time predictability. Typically, SPM-based architectures use DMA for fetching data from off-chip memory and global load instructions for loading fine-grained data directly into registers. For such architectures, neither capacity-only nor bandwidth-only loop tiling can efficiently use the bandwidth and SPM. This paper introduces a bandwidth-aware loop tiling approach that enables a tradeoff between SPM space utilization and bandwidth utilization to be made, by leveraging a runtime tiling framework and a cross-host-kernel IPA. Experimental results demonstrate that our approach can achieve the performance improvement of up to 4x, with a geometric average of 26%. Mingchuan Wu, Ying Liu 0055, Huimin Cui, Qingfu Wei, Quanfeng Li, Jingling Xue, Xiaobing Feng 0002 |
PACT | 3 |
| 2020 | VTensor: Using Virtual Tensors to Build a Layout-oblivious AI Programming FrameworkabstractTensors are a popular programming interface for developing AI algorithms. Representative AI programming frameworks require developers to be always aware of tensor layouts, thereby reducing their productivity in integrating an existing operation with a new library and/or writing a new operation. We propose VTensor, a layout-oblivious virtual tensor programming interface, together with a global layout inference mechanism to resolve the layout required by virtual tensors. Furthermore, VTensor leverages a layout-oriented optimization to globally minimize the number of layout conversion operations, together with a straggler-ware scheduling algorithm and a pool-based memory allocation scheme to globally allocate resources. VTensor yields significant speedup and LOC (Lines of Codes) reduction compared to TensorFlow. Feng Yu 0019, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
PACT | 3 |
| 2020 | Referee: A Pattern-Guided Approach for Auto Design in Compiler-Based Analyzers
Lei Wang 0004, Ying Liu 0055, Huimin Cui, Jingling Xue, Xiaobing Feng 0002 |
SANER | 5 |
| 2020 | DNNTune: Automatic Benchmarking DNN Models for Mobile-cloud ComputingabstractDeep Neural Networks (DNNs) are now increasingly adopted in a variety of Artificial Intelligence (AI) applications. Meantime, more and more DNNs are moving from cloud to the mobile devices, as emerging AI chips are integrated into mobiles. Therefore, the DNN models can be deployed in the cloud, on the mobile devices, or even mobile-cloud coordinate processing, making it a big challenge to select an optimal deployment strategy under specific objectives. This article proposes a DNN tuning framework, i.e., DNNTune, that can provide layer-wise behavior analysis across a number of platforms. Using DNNTune, this article further selects 13 representative DNN models, including CNN, LSTM, and MLP, and three mobile devices ranging from low-end to high-end, and two AI accelerator chips to characterize the DNN models on these devices to further assist users finding opportunities for mobile-cloud coordinate computing. Our experimental results demonstrate that DNNTune can find a coordinated deployment achieving up to 1.66× speedup and 15× energy saving comparing with mobile-only and cloud-only deployment. Chunwei Xia, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
ACM Trans. Archit. Code Optim. | 3 |
| 2019 | PPOpenCL: a performance-portable OpenCL compiler with host and kernel thread code fusionabstractOpenCL offers code portability but no performance portability. Given an OpenCL program X specifically written for one platform P, existing OpenCL compilers, which usually optimize its host and kernel codes individually, often yield poor performance for another platform Q. Instead of obtaining a performance-improved version of X for Q via manual tuning, we aim to achieve this automatically by a source-to-source OpenCL compiler framework, PPOpenCL. By fusing X's host and kernel thread codes (with the operations in different work-items in the same work-group represented explicitly), we are able to apply data flow analyses, and subsequently, performance-enhancing optimizations on a fused control flow graph specifically for platform Q. Validation against OpenCL benchmarks shows that PPOpenCL (implemented in Clang 3.9.1) can achieve significantly improved portable performance on seven platforms considered. Ying Liu 0055, Mingchuan Wu, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
CC | 4 |
| 2019 | Panthera: holistic memory management for big data processing over hybrid memoriesabstractModern data-parallel systems such as Spark rely increasingly on in-memory computing that can significantly improve the efficiency of iterative algorithms. To process real-world datasets, modern data-parallel systems often require extremely large amounts of memory, which are both costly and energy-inefficient. Emerging non-volatile memory (NVM) technologies offers high capacity compared to DRAM and low energy compared to SSDs. Hence, NVMs have the potential to fundamentally change the dichotomy between DRAM and durable storage in Big Data processing. However, most Big Data applications are written in managed languages (e.g., Scala and Java) and executed on top of a managed runtime (e.g., the Java Virtual Machine) that already performs various dimensions of memory management. Supporting hybrid physical memories adds in a new dimension, creating unique challenges in data replacement and migration. Chenxi Wang 0005, Huimin Cui, John N. Zigman, Haris Volos 0001, Onur Mutlu, Xiaobing Feng 0002, Guoqing Harry Xu |
PLDI | 2 |
| 2018 | Revisiting Loop Tiling for Datacenters: Live and Let LiveabstractAs DNNs gain popularity in modern datacenters, it becomes imperative to revisit compiler optimizations for DNNs in a colocation scenario. Loop tiling turns out to be the most significant compiler optimization, since DNNs typically apply a series of matrix computations iteratively to a massive amount of data. Huimin Cui, Jingling Xue, Xiaobing Feng 0002 |
ICS | 2 |
| 2018 | On Retargeting the AI Programming Framework to New Hardwares
Yisong Chang, Denghui Li, Chunwei Xia, Huimin Cui, Ke Zhang 0017, Xiaobing Feng 0002 |
NPC | 5 |
| 2018 | Lazygraph: lazy data coherency for replicas in distributed graph-parallel computationabstractReplicas 1 of a vertex play an important role in existing distributed graph processing systems which make a single vertex to be parallel processed by multiple machines and access remote neighbors locally without any remote access. However, replicas of vertices introduce data coherency problem. Existing distributed graph systems treat replicas of a vertex v as an atomic and indivisible vertex, and use an eager data coherency approach to guarantee replicas atomicity. In eager data coherency approach, any changes to vertex data must be immediately communicated to all replicas of v, thus leading to frequent global synchronizations and communications. Lei Wang 0004, Liangji Zhuang, Junhang Chen, Huimin Cui, Ying Liu 0055, Xiaobing Feng 0002 |
PPoPP | 4 |
| 2018 | NVM Streaker: a fast and reconfigurable performance simulator for non-volatile memory-based memory architecture
Danqi Hu, Chenxi Wang 0005, Huimin Cui, Lei Wang 0004, Ying Liu 0055, Xiaobing Feng 0002 |
J. Supercomput. | 4 |
| 2017 | Parallel Incremental Frequent Itemset Mining for Large Data
Yu-Geng Song, Huimin Cui, Xiaobing Feng 0002 |
J. Comput. Sci. Technol. | 2 |
| 2016 | Articulation points guided redundancy elimination for betweenness centralityabstractBetweenness centrality (BC) is an important metrics in graph analysis which indicates critical vertices in large-scale networks based on shortest path enumeration. Typically, a BC algorithm constructs a shortest-path DAG for each vertex to calculate its BC score. However, for emerging real-world graphs, even the state-of-the-art BC algorithm will introduce a number of redundancies, as suggested by the existence of articulation points. Articulation points imply some common sub-DAGs in the DAGs for different vertices, but existing algorithms do not leverage such information and miss the optimization opportunity. Lei Wang 0004, Fan Yang 0049, Liangji Zhuang, Huimin Cui, Xiaobing Feng 0002 |
PPoPP | 4 |
| 2016 | Predicting Cross-Core Performance Interference on Multicore Processors with Regression AnalysisabstractDespite their widespread adoption in cloud computing, multicore processors are heavily under-utilized in terms of computing resources. To avoid the potential for negative and unpredictable interference, co-location of a latency-sensitive application with others on the same multicore processor is disallowed, leaving many cores idle and causing low machine utilization. To enable co-location while providing QoS guarantees, it is challenging but important to predict performance interference between co-located applications. We observed that the performance degradation of an application can be represented as a piecewise predictor function of the aggregate pressures on shared resources from all cores. Based on this observation, we propose to adopt regression analysis to build a predictor function for an application. Furthermore, the prediction model thus obtained for an application is able to characterize its contentiousness and sensitivity. Validation using a large number of single-threaded and multi-threaded benchmarks and nine real-world datacenter applications on two different platforms shows that our approach is also precise, with an average error not exceeding 0.4 percent. Huimin Cui, Jingling Xue, Xiaobing Feng 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Hadoop+: Modeling and Evaluating the Heterogeneity for MapReduce Applications in Heterogeneous ClustersabstractDespite the widespread adoption of heterogeneous clusters in modern data centers, modeling heterogeneity is still a big challenge, especially for large-scale MapReduce applications. In a CPU/GPU hybrid heterogeneous cluster, allocating more computing resources to a MapReduce application does not always mean better performance, since simultaneously running CPU and GPU tasks will contend for shared resources. Wenting He, Huimin Cui, Binbin Lu, Shengmei Li, Gong Ruan, Jingling Xue, Xiaobing Feng 0002, Wensen Yang, Youliang Yan |
ICS | 2 |
| 2015 | Global μ-stability of impulsive reaction-diffusion neural networks with unbounded time-varying delays and bounded continuously distributed delays
Huimin Cui, Jianxin Feng, Tingfeng Wang |
Neurocomputing | 1 |
| 2015 | A novel single multiplicative neuron model trained by an improved glowworm swarm optimization algorithm for time series prediction
Huimin Cui, Jianxin Feng, Tingfeng Wang |
Knowl. Based Syst. | 1 |
| 2015 | WiseThrottling: a new asynchronous task scheduler for mitigating I/O bottleneck in large-scale datacenter servers
Lei Liu 0030, Huimin Cui, Lei Wang 0004, Ying Liu 0055, Xiaobing Feng 0002, Pen-Chung Yew |
J. Supercomput. | 3 |
| 2014 | Specializing Compiler Optimizations through Programmable Composition for Dense Matrix ComputationsabstractGeneral purpose compilers aim to extract the best average performance for all possible user applications. Due to the lack of specializations for different types of computations, compiler attained performance often lags behind those of the manually optimized libraries. In this paper, we demonstrate a new approach, programmable composition, to enable the specialization of compiler optimizations without compromising their generality. Our approach uses a single pass of source-level analysis to recognize a common pattern among dense matrix computations. It then tags the recognized patterns to trigger a sequence of general-purpose compiler optimizations specially composed for them. We show that by allowing different optimizations to adequately communicate with each other through a set of coordination handles and dynamic tags inserted inside the optimized code, we can specialize the composition of general-purpose compiler optimizations to attain a level of performance comparable to those of manually written assembly code by experts, thereby allowing selected computations in applications to benefit from similar levels of optimizations as those manually applied by experts. Qing Yi, Huimin Cui |
MICRO | 3 |
| 2014 | Dynamic I/O-Aware Scheduling for Batch-Mode Applications on Chip Multiprocessor Systems of Cluster Platforms
Huimin Cui, Lei Wang 0004, Lei Liu 0030, Chenggang Wu 0002, Xiaobing Feng 0002, Pen-Chung Yew |
J. Comput. Sci. Technol. | 2 |
| 2013 | An empirical model for predicting cross-core performance interference on multicore processorsabstractDespite their widespread adoption in cloud computing, multicore processors are heavily under-utilized in terms of computing resources. To avoid the potential for negative and unpredictable interference, co-location of a latency-sensitive application with others on the same multicore processor is disallowed, leaving many cores idle and causing low machine utilization. To enable co-location while providing QoS guarantees, it is challenging but important to predict performance interference between co-located applications. This research is driven by two key insights. First, the performance degradation of an application can be represented as a predictor function of the aggregate pressures on shared resources from all cores, regardless of which applications are co-running and what their individual pressures are. Second, a predictor function is piecewise rather than non-piecewise as in prior work, thereby enabling different types of dominant contention factors to be more accurately captured by different subfunctions in its different subdomains. Based on these insights, we propose to adopt a two-phase regression approach to efficiently building a predictor function. Validation using a large number of benchmarks and nine real-world datacenter applications on three different platforms shows that our approach is also precise, with an average error not exceeding 0.4%. When applied to the nine datacenter applications, our approach improves overall resource utilization from 50% to 88% at the cost of 10% QoS degradation. Xiaobing Feng 0002, Huimin Cui, Youliang Yan, Jingling Xue, Wensen Yang |
PACT | 3 |
| 2013 | Layout-oblivious compiler optimization for matrix computationsabstractMost scientific computations serve to apply mathematical operations to a set of preconceived data structures, e.g., matrices, vectors, and grids. In this article, we use a number of widely used matrix computations from the LINPACK library to demonstrate that complex internal organizations of data structures can severely degrade the effectiveness of compiler optimizations. We then present a data-layout-oblivious optimization methodology, where by isolating an abstract representation of the computations from complex implementation details of their data, we enable these computations to be much more accurately analyzed and optimized through varying state-of-the-art compiler technologies. We evaluated our approach on an Intel 8-core platform using two source-to-source compiler infrastructures, Pluto and EPOD. Our results show that while the efficiency of a computational kernel differs when using different data layouts, the alternative implementations typically benefit from a common set of optimizations on the operations. Therefore separately optimizing the operations and the data layout of a computation could dramatically enhance the effectiveness of compiler optimizations compared with the conventional approaches of using a unified representation. Huimin Cui, Qing Yi, Jingling Xue, Xiaobing Feng 0002 |
ACM Trans. Archit. Code Optim. | 1 |
| 2012 | Layout-oblivious optimization for matrix computationsabstractMost scientific computations serve to apply mathematical operations to a set of preconceived data structures, e.g., matrices, vectors, and grids. In this paper, we use a number of widely used matrix computations from the LINPACK library to demonstrate that complex internal organizations of data structures can severely degrade the effectiveness of compilers optimizations. We then present a data layout oblivious optimization methodology, where by isolating an abstract representation of computations from complex implementation details of their data, we enable these computations to be much more accurately analyzed and optimized through varying state-of-the-art compiler technologies. Huimin Cui, Qing Yi, Jingling Xue, Xiaobing Feng 0002 |
PACT | 1 |
| 2012 | A Highly Parallel Reuse Distance Analysis Algorithm on GPUsabstractReuse distance analysis is a runtime approach that has been widely used to accurately model the memory system behavior of applications. However, traditional reuse distance analysis algorithms use tree-based data structures and are hard to parallelize, missing the tremendous computing power of modern architectures such as the emerging GPUs. This paper presents a highly-parallel reuse distance analysis algorithm (HP-RDA) to speedup the process using the SPMD execution model of GPUs. In particular, we propose a hybrid data structure of hash table and local arrays to flatten the traditional tree representation of memory access traces. Further, we use a probabilistic model to correct any loss of precision from a straightforward parallelization of the original sequential algorithm. Our experimental results show that using an NVIDIA GPU, our algorithm achieves a factor of 20 speedup over the traditional sequential algorithm with less than 1% loss in precision. Huimin Cui, Qing Yi, Jingling Xue, Lei Wang 0004, Xiaobing Feng 0002 |
IPDPS | 1 |
| 2012 | A Hybrid Circular Queue Method for Iterative Stencil Computations on GPUs
Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
J. Comput. Sci. Technol. | 2 |
| 2012 | Extendable pattern-oriented optimization directivesabstractAlgorithm-specific, that is, semantic-specific optimizations have been observed to bring significant performance gains, especially for a diverse set of multi/many-core architectures. However, current programming models and compiler technologies for the state-of-the-art architectures do not exploit well these performance opportunities. In this article, we propose a pattern-making methodology that enables algorithm-specific optimizations to be encapsulated into “optimization patterns”. Such optimization patterns are expressed in terms of preprocessor directives so that simple annotations can result in significant performance improvements. To validate this new methodology, a framework, named EPOD, is developed to map these directives into the underlying optimization schemes for a particular architecture. It is difficult to create an exact performance model to determine an optimal or near-optimal optimization scheme (including which optimizations to apply and in which order) for a specific application, due to the complexity of applications and architectures. However, it is trackable to build individual optimization components and let compiler developers synthesize an optimization scheme from these components. Therefore, our EPOD framework provides an Optimization Programming Interface (OPI) for compiler developers to define new optimization schemes. Thus, new patterns can be integrated into EPOD in a flexible manner. We have identified and implemented a number of optimization patterns for three representative computer platforms. Our experimental results show that a pattern-guided compiler can outperform the state-of-the-art compilers and even achieve performance as competitive as hand-tuned code. Therefore, such a pattern-making methodology represents an encouraging direction for domain experts' experience and knowledge to be integrated into general-purpose compilers. Huimin Cui, Jingling Xue, Lei Wang 0004, Xiaobing Feng 0002, Dongrui Fan |
ACM Trans. Archit. Code Optim. | 1 |
| 2011 | Extendable pattern-oriented optimization directivesabstractCurrent programming models and compiler technologies for multi-core processors do not exploit well the performance benefits obtainable by applying algorithm-specific, i.e., semantic-specific optimizations to a particular application. In this work, we propose a pattern-making methodology that allows algorithm-specific optimizations to be encapsulated into “optimization patterns” that are expressed in terms of pre-processor directives so that simple annotations can result in significant performance improvements. To validate this new methodology, a framework, named EPOD, is developed to map such directives to the underlying optimization schemes. We have identified and implemented a number of optimization patterns for three representative computer platforms. Our experimental results show that a pattern-guided compiler can outperform the state-of-the-art compilers and even achieve performance as competitive as hand-tuned code. Thus, such a pattern-making methodology represents an encouraging direction for domain experts' experience and knowledge to be integrated into general-purpose compilers. Huimin Cui, Jingling Xue, Lei Wang 0004, Xiaobing Feng 0002, Dongrui Fan |
CGO | 1 |
| 2011 | Automatic Library Generation for BLAS3 on GPUsabstractHigh-performance libraries, the performance-critical building blocks for high-level applications, will assume greater importance on modern processors as they become more complex and diverse. However, automatic library generators are still immature, forcing library developers to manually tune library to meet their performance objectives. We are developing a new script-controlled compilation framework to help domain experts reduce much of the tedious and error-prone nature of manual tuning, by enabling them to leverage their expertise and reuse past optimization experiences. We focus on demonstrating improved performance and productivity obtained through using our framework to tune BLAS3 routines on three GPU platforms: up to 5.4x speedups over the CUBLAS achieved on NVIDIA GeForce 9800, 2.8x on GTX285, and 3.4x on Fermi Tesla C2050. Our results highlight the potential benefits of exploiting domain expertise and the relations between different routines (in terms of their algorithms and data structures). Huimin Cui, Lei Wang 0004, Jingling Xue, Xiaobing Feng 0002 |
IPDPS | 1 |
| 2010 | An adaptive task creation strategy for work-stealing schedulingabstractWork-stealing is a key technique in many multi-threading programming languages to get good load balancing. The current work-stealing techniques have a high implementation overhead in some applications and require a large amount of memory space for data copying to assure correctness. They also cannot handle many application programs that have an unbalanced call tree or have no definitive working sets. Lei Wang 0004, Huimin Cui, Yuelu Duan, Xiaobing Feng 0002, Pen-Chung Yew |
CGO | 2 |
| 2010 | Landing Stencil Code on Godson-T
Huimin Cui, Lei Wang 0004, Dongrui Fan, Xiaobing Feng 0002 |
J. Comput. Sci. Technol. | 1 |