Jianmin Pang

dblp:45/5874 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Security and privacy · 2Software engineering, systems software and programming languages · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Optimization of General Matrix Multiplication for RISC-V Processors
abstract
General Matrix Multiplication (GEMM) is a fundamental operation in high-performance computing and artificial intelligence, with its performance directly impacting the overall efficiency of higher-level applications. This paper optimizes GEMM by leveraging the characteristics of the RISC-V architecture, achieving significant performance improvements on the XuanTie C908 and C910 processors. It is also observed that traditional optimization method become a bottleneck for small-scale GEMM performance on RISC-V. To address this, we propose an analytical model-based approach that effectively enhances small-scale GEMM performance. In performance evaluation, we tested square matrices and irregular matrices (derived from convolutional neural networks) to validate the optimized GEMM across various scenarios and assess its performance in realworld applications. Compared to the GEMM routines in popular open-source BLAS (Basic Linear Algebra Subprograms) libraries (OpenBLAS and BLIS), the optimized GEMM achieves average speedups of$1.29 \times$and$2.20 \times$, with maximum speedups of$6.51 \times$and$27.85 \times$on the C908 platform. On the C910 platform, the average speedups are$2.38 \times$and$5.15 \times$, with maximum speedups of$17.15 \times$and$84.16 \times$.
Bei Zhou 0004, Jianmin Pang, Fei Li 0045, Hongru Yang, Mengyao Duan, Jinchen Xu
HPCC3
2023 AMPT: Automatic Mixed-Precision Tuning Based on Accuracy Gain
abstract
Mixed-precision tuning is a valuable approach for striking a balance between the accuracy and performance of floating-point computations. However, selecting suitable configurations from numerous mixed-precision configurations can be challenging. The discrete and finite nature of floating-point numbers introduces unavoidable rounding errors that are difficult to predict. To address these challenges, we propose an accuracy gain evaluation function for expressions. This evaluation function quantifies the extent of accuracy improvement achieved by adjusting the precision of the operations within the expression. This is achieved through a thorough examination of the impact of each operation's condition number and implementation error on the final error. Based on this evaluation function, we present AMPT, a mixed-precision tuning method specifically for expressions. AMPT not only identifies the optimal mixed-precision configuration but also automatically generates the corresponding code. Experimental results demonstrate that AMPT accurately assesses the accuracy gains of different mixed-precision configurations, and the mixed-precision code generated by AMPT can achieve low-error computation of expressions with low performance overhead.
Jiangwei Hao, Yuanyuan Xia, Fei Li 0045, Hongru Yang, Zongjiang Yi, Bei Zhou 0004, Jianmin Pang
APSEC7
2023 Few-shot short-text classification with language representations and centroid similarity
Wenfu Liu, Jianmin Pang
Appl. Intell.2
2023 Binary Vulnerability Similarity Detection Based on Function Parameter Dependency
abstract
Many existing works compute the binary vulnerability similarity based on binary procedure, which has coarse detection granularity and cannot locate the vulnerability trigger position accurately, and have a higher false positive rate, so a new binary vulnerability similarity detection method based on function parameter dependency in hazard API is proposed. First, convert the instructions of different architectures into an intermediate language, and use the compiler with a back-end optimizer to optimize and normalize the binary procedure. Then, locate the hazard API that appears in the binary procedure, and perform the function parameters dependency analysis to generate a set of parameter slices on the hazard API. Experiments show that the method has a higher recall rate (up to 14.3% better than the baseline model) in real-world scenarios, and not only locates the triggering position of the vulnerability but also identifies the fixed vulnerability.
Qudong He, Fudong Liu, Jianmin Pang, Ruinan Yang, Jiabin Yin, Yunxiang Ge
Int. J. Semantic Web Inf. Syst.5
2022 Automatically Generating High-performance Matrix Multiplication Kernels on the Latest Sunway Processor
abstract
We present an approach to the automatic generation of efficient matrix multiplication code on the latest Sunway processor, which will be employed by the next-generation machine of Sunway TaihuLight, one of the fastest supercomputers on earth. The method allows users to write simple C code and automatically generates high-performance matrix multiplication kernels. It uses polyhedral transformations to implement rapid compute decomposition, data exchanges across memory hierarchy and memory latency hiding. An assembly routine is finally integrated into the generated kernels. While achieving up to 90.14% of the theoretical peak performance, our method surpasses a highly tuned library by 9.44%. Compared with existing techniques, our approach reduces the software development life cycle to generate efficient matrix code from months to seconds. We also take into account batched matrix multiplication and some fusion patterns for deep learning (DL), outperforming the library-based implementations by 1.30 × and 1.67 ×.
Xiaohan Tao, Jinlong Xu, Jianmin Pang, Jie Zhao 0002
ICPP5
2022 Revisiting split tiling for stencil computations in polyhedral compilation
Jianmin Pang
J. Supercomput.3
2021 Dual-axial self-attention network for text classification
Xiaochuan Zhang, Xipeng Qiu, Jianmin Pang, Fudong Liu, Xingwei Li
Sci. China Inf. Sci.3
2021 Compiler-directed scratchpad memory data transfer optimization for multithreaded applications on a heterogeneous many-core architecture
abstract
Abstract The heterogeneous many-core architecture plays an important role in the fields of high-performance computing and scientific computing. It uses accelerator cores with on-chip memories to improve performance and reduce energy consumption. Scratchpad memory (SPM) is a kind of fast on-chip memory with lower energy consumption compared with a hardware cache. However, data transfer between SPM and off-chip memory can be managed only by a programmer or compiler. In this paper, we propose a compiler-directed multithreaded SPM data transfer model (MSDTM) to optimize the process of data transfer in a heterogeneous many-core architecture. We use compile-time analysis to classify data accesses, check dependences and determine the allocation of data transfer operations. We further present the data transfer performance model to derive the optimal granularity of data transfer and select the most profitable data transfer strategy. We implement the proposed MSDTM on the GCC complier and evaluate it on Sunway TaihuLight with selected test cases from benchmarks and scientific computing applications. The experimental result shows that the proposed MSDTM improves the application execution time by 5.49 $$\times$$ × and achieves an energy saving of 5.16 $$\times$$ × on average.
Xiaohan Tao, Jianmin Pang, Jinlong Xu
J. Supercomput.2
2021 Enhancing Dynamic Binary Translation in Mobile Computing by Leveraging Polyhedral Optimization
abstract
Dynamic binary translation (DBT) is gaining importance in mobile computing. Mobile Edge Computing (MEC) augments mobile devices with powerful servers, whereas edge servers and smartphones are usually based on heterogeneous architecture. To leverage high‐performance resources on servers, code offloading is an ideal approach that relies on DBT. In addition, mobile devices equipped with multicore processors and GPU are becoming ubiquitous. Migrating x86_64 application binaries to mobile devices by using DBT can also make a contribution to providing various mobile applications, e.g., multimedia applications. However, the translation efficiency and overall performance of DBT for application migration are not satisfactory, because of runtime overhead and low quality of the translated code. Meanwhile, traditional DBT systems do not fully exploit the computational resources provided by multicore processors, especially when translating sequential guest applications. In this work, we focus on leveraging ubiquitous multicore processors to improve DBT performance by parallelizing sequential applications during translation. For that, we propose LLPEMU, a DBT framework that combines binary translation with polyhedral optimization. We investigate the obstacles of adapting existing polyhedral optimization in compilers to DBT and present a feasible method to overcome these issues. In addition, LLPEMU adopts static‐dynamic combination to ensure that sequential binaries are parallelized while incurring low runtime overhead. Our evaluation results show that LLPEMU outperforms QEMU significantly on the PolyBench benchmark.
Jianmin Pang, Fudong Liu, Jun Wang 0070, Jie Tan 0002
Wirel. Commun. Mob. Comput.2
2018 Automatic Benchmark Generation Framework for Malware Detection
abstract
To address emerging security threats, various malware detection methods have been proposed every year. Therefore, a small but representative set of malware samples are usually needed for detection model, especially for machine-learning-based malware detection models. However, current manual selection of representative samples from large unknown file collection is labor intensive and not scalable. In this paper, we firstly propose a framework that can automatically generate a small data set for malware detection. With this framework, we extract behavior features from a large initial data set and then use a hierarchical clustering technique to identify different types of malware. An improved genetic algorithm based on roulette wheel sampling is implemented to generate final test data set. The final data set is only one-eighteenth the volume of the initial data set, and evaluations show that the data set selected by the proposed framework is much smaller than the original one but does not lose nearly any semantics.
Guanghui Liang, Jianmin Pang, Zheng Shan, Runqing Yang
Secur. Commun. Networks2
2017 Planning virtual infrastructures for time critical applications with multiple deadline constraints
Arie Taal, Paul Martin 0002, Yang Hu 0013, Huan Zhou 0006, Jianmin Pang, Cees T. A. M. de Laat, Zhiming Zhao
Future Gener. Comput. Syst.6
2015 On-Demand Self-Adaptivity of Service Availability for Cloud Multi-tier Applications
abstract
Cloud data centers are increasingly becoming the first choice for many Internet enterprises (especially medium-sized and small ones) to deploy their online application. However, scaling service availability autonomously is a critical issue for Internet applications. Surplus service supply may take a lot of unnecessary money and insufficient resources reserve will result in denial of service when meeting sudden massive requests. A general idea to address the issue is increasing the resource utilization through workload balancing and dynamic resources management. In this paper, we propose a self-adaptive approach that is suitable for the multi-tier cloud applications. The approach tries to scale the applications' service availability on demand and reduce infrastructure costs by improving utilization of the resources that are already billed. Moreover, in order to cope with the unexpected requests, an evaluation method is adopted to estimate the trend of requests' development and then decide to add or remove working servers. Finally, we compare the proposed algorithms with the workload consolidation method by a quantitative analysis, showing the superiorities of costs and performance in some situations.
Jianmin Pang, Ning Qi
CCGRID2
2013 Convergence and Scalarization in Whole Function Vectorization
abstract
When implementing SPMD programs on multi core platforms, whole function vectorization is an important optimization method. SPMD program has drawback that lots of instructions across multi threads are redundant which is sustained in vectorization. This paper proposes to alleviate this overhead by detecting scalar operations and extract them out in vectorization instructions. An algorithm is designed to deal with control flow and data flow synchronously in which convergent and invariance analysis is employed to statically identify convergent execution and invariant values or instructions. Our algorithm is effectively on implementing SPMD programs on multi core platforms. The experiments show our method could improve the execution efficiency by 13.3%.
Jianmin Pang, Jiuzhen Jin, Chao Dai
DASC2
2008 A Consistency Combination Algorithm for Global Dynamic Computation and Data Decomposition
abstract
The speed of processor accessing local memory is much faster than that of accessing remote memory by communication on distributed memory machines. To reduce the cost brought by the communication the parallel recognition compiler must give efficient computation partition and data distribution, and guarantee that the data needed to visit during computation is kept in the local memories. While in the process of parallel recognition, we observed that in many cases there is no global consistency decomposition; and using multiple fashions of data distribution may improve the performance of parallelism. A consistency combination algorithm for dynamic decomposition that allows data reorganization was presented in this paper. The algorithm starts from the solving of decomposition in single loop nest, and then it fuses different decomposition fashions from the sets of decomposition results and using linear transformation for global data distribution consistence. Our algorithm also takes the structure and the control flow of programs into account to direct the priority order in the process of linear transformation. Effectiveness of our algorithm was shown by verification in the end of this paper.
Rongcai Zhao, Jianmin Pang
CISIS3
2007 Parameter and Return-value Analysis of Binary Executables
abstract
The recovery of parameter and return-value plays an important role in decompilation, reverse engineering, binary translation and software maintenance etc. Furthermore, related approaches are very useful to inter- procedural analyzing and slicing of binary executable. However, the operations on parameters and return- values always appear obscure after the optimizing phases of a compiler, which will make the recovery hard to realize. In this paper, we present a flow-insensitive but context-sensitive algorithm based on data dependence analysis to get back parameters and return- values. In addition, we discuss our experimental results obtained by applying our techniques to a static binary translation framework. Evidence shows that our method performs well in analyzing the parameters and return-values of executables. We use an IA-64 executable for demonstration, but our techniques are not limited to any particular architecture.
Rongcai Zhao, Jianmin Pang
COMPSAC (1)3
2005 LFTOP: An LF-Based Approach to Domain-Specific Reasoning
Jianmin Pang, Paul Callaghan, Zhaohui Luo
J. Comput. Sci. Technol.1