EDBT 2026 Demo / reviewers in the wild / expert
Weixing Ji
dblp:45/2857
· DBLP profile ↗
58ranked-venue papers
5as first author
24since 2021 · last 2026
0000-0002-3250-0435ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 39 · 4 first-author · 15 since 2021Software engineering, systems software and programming languages · 12 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ARROW: Adaptive Row Reorganization for Warp-Balanced SpMV
Jianhua Gao 0001, Weixing Ji |
APPT | 3 |
| 2026 | ApproxRAG: A Systematic Framework for Mitigating I/O Overheads in Retrieval-Augmented Generation
Danying Ge, Qizhi Jiang, Jianhua Gao 0001, Weixing Ji |
Euro-Par (2) | 5 |
| 2026 | JDExtractor: an automated approach for efficient extraction of defect-related methods in Java projects
Jiawei Ye, Weixing Ji |
Autom. Softw. Eng. | 3 |
| 2026 | HRPF: A parallel programming framework for recursive algorithms on heterogeneous CPU-GPU systems
Yizhuo Wang 0001, Senhao Shao, Jianhua Gao 0001, Weixing Ji, Hongbo Xing |
Parallel Comput. | 5 |
| 2025 | Adaptive point cloud compression based on precision-aware floating-point encoding
Yanpeng Han, Yizhuo Wang 0001, Fawang Liu, Jianhua Gao 0001, Weixing Ji |
CCF Trans. High Perform. Comput. | 5 |
| 2025 | Deep learning-based software engineering: progress, challenges, and opportunitiesabstractAbstract Researchers have recently achieved significant advances in deep learning techniques, which in turn has substantially advanced other research disciplines, such as natural language processing, image processing, speech recognition, and software engineering. Various deep learning techniques have been successfully employed to facilitate software engineering tasks, including code generation, software refactoring, and fault localization. Many studies have also been presented in top conferences and journals, demonstrating the applications of deep learning techniques in resolving various software engineering tasks. However, although several surveys have provided overall pictures of the application of deep learning techniques in software engineering, they focus more on learning techniques, that is, what kind of deep learning techniques are employed and how deep models are trained or fine-tuned for software engineering tasks. We still lack surveys explaining the advances of subareas in software engineering driven by deep learning techniques, as well as challenges and opportunities in each subarea. To this end, in this study, we present the first task-oriented survey on deep learning-based software engineering. It covers twelve major software engineering subareas significantly impacted by deep learning techniques. Such subareas spread out through the whole lifecycle of software development and maintenance, including requirements engineering, software development, testing, maintenance, and developer collaboration. As we believe that deep learning may provide an opportunity to revolutionize the whole discipline of software engineering, providing one survey covering as many subareas as possible in software engineering can help future research push forward the frontier of deep learning-based software engineering more systematically. For each of the selected subareas, we highlight the major advances achieved by applying deep learning techniques with pointers to the available datasets in such a subarea. We also discuss the challenges and opportunities concerning each of the surveyed software engineering subareas. Xiangping Chen, Xing Hu 0008, Yuan Huang 0002, He Jiang 0001, Weixing Ji, Yanjie Jiang, Yanyan Jiang 0001, Bo Liu 0094, Hui Liu 0003, Xiaoli Lian, Guozhu Meng, Xin Peng 0001, Hailong Sun 0001, Lin Shi 0006, Bo Wang 0050, Chong Wang 0013, Jifeng Xuan, Xin Xia 0001, Yibiao Yang, Yixin Yang 0006, Li Zhang 0029, Yuming Zhou, Lu Zhang 0023 |
Sci. China Inf. Sci. | 5 |
| 2025 | RaNAS: Resource-Aware Neural Architecture Search for Edge ComputingabstractNeural architecture search (NAS) for edge devices is often time-consuming because of long-latency deploying and testing on edge devices. The ability to accurately predict the computation cost and memory requirement for convolutional neural networks (CNNs) in advance holds substantial value. Existing work primarily relies on analytical models, which can result in high prediction errors. This article proposes a resource-aware NAS (RaNAS) model based on various features. Additionally, a new graph neural network is introduced to predict inference latency and maximum memory requirements for CNNs on edge devices. Experimental results show that, within the error bound of ±1%, RaNAS achieves an accuracy improvement of approximately 8% for inference latency prediction and about 25% for maximum memory occupancy prediction over the state-of-the-art approaches. Jianhua Gao 0001, Zeming Liu, Yizhuo Wang 0001, Weixing Ji |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | PTPS: Precision-Aware Task Partitioning and Scheduling for SpMV on CPU-FPGA Heterogeneous PlatformsabstractThe CPU-FPGA heterogeneous computing architecture is extensively employed in the embedded domain due to its low cost and power efficiency, with numerous sparse matrix-vector multiplication (SpMV) acceleration efforts already targeting this architecture. However, existing work rarely includes collaborative SpMV computations between CPU and FPGA, which limits the exploration of hybrid architectures that could potentially offer enhanced performance and flexibility. This article introduces an FPGA architecture design that supports multiprecision SpMV computations, including FP16, FP32, and FP64. Building on this, PTPS, a precision-aware SpMV task partitioning and dynamic scheduling algorithm tailored for the CPU-FPGA heterogeneous architecture, is proposed. The core idea of PTPS is lossless partitioning of sparse matrices across multiple precisions, prioritizing low-precision SpMV computations on the FPGA and high-precision computations on the CPU. PTPS not only leverages the strengths of CPU and FPGA for collaborative SpMV computations but also reduces data transmission overhead between them, thereby improving the overall computational efficiency. Experimental evaluation demonstrates that the proposed approach offers an average speedup of$1.57\times $over the CPU-only approach and$2.58\times $over the FPGA-only approach. Jianhua Gao 0001, Xingze Huang, Yizhuo Wang 0001, Weixing Ji |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | Automated Recommendation of Extracting Local Variable RefactoringsabstractExtracting local variable refactoring is frequently employed to replace one or more occurrences of a complex expression with simple accesses to a newly introduced variable. To facilitate refactoring, most IDEs can automate the extract local variable refactorings when the to-be-extracted expressions are selected by developers. However, refactoring tools usually replace all expressions that are lexically identical to the selected one without a comprehensive analysis of the safety of the refactoring. The automatically conducted refactorings may lead to serious software defects. Besides that, existing refactoring tools rely heavily on software developers to spot to-be-extracted expressions although it is often challenging for inexperienced developers and maintainers to make the selection. To this end, in this article, we propose an automated approach, called ValExtractor+ , to recommending extract local variable refactoring opportunities and to automatically and safely conduct the refactorings. ValExtractor+ is composed of two parts, i.e., solutionAdvisor and opportunityAdvisor . Given a to-be-extracted expression, solutionAdvisor leverages lightweight static source code analysis to validate potential side effects of the expression, and to identify expressions that could be extracted together with the selected expression as a single variable without changing the semantics of the program or introducing any new exceptions. The static code analysis significantly improves the safety of automated extraction of local variables. To free programmers from manually selecting to-be-extracted expressions, opportunityAdvisor leverages solutionAdvisor to automatically retrieve all expressions that could be extracted safely as well as their refactoring solutions. It then leverages a learning-based classifier to predict which of the retrieved expressions should be extracted. Evaluations on open-source applications suggest that solutionAdvisor successfully avoided all defects (more than two hundred) caused by extracting local variable refactorings conducted by Eclipse (243 defects) or IntelliJ IDEA (263 defects). Additionally, opportunityAdvisor was able to effectively recommend expressions for extraction, achieving 307 true positives (TP) and 21,121 true negatives (TN). Four pull requests from our work (PR IDs: 66, 333, 439, and 360) were successfully merged into the Eclipse community repository, showcasing the practical impact and robustness of our approach as recognized by the wider developer community. Yanjie Jiang, Xiaye Chi, Yuxia Zhang, Weixing Ji, Guangjie Li, Weixiao Wang, Yunni Xia, Lu Zhang 0023, Hui Liu 0003 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2025 | An Empirical Study on the Relationship between Defects and Source Code's UnnaturalnessabstractNatural languages are “natural” in that texts in natural languages are repetitive and predictable. Recent research indicates that programming languages share similar characteristics (naturalness), with source code displaying patterns of repetition and predictability. Notably, studies have shown that buggy code deviates from these natural patterns in that buggy code is significantly less natural than bug-free one. In this article, we conduct a large-scale and extensive empirical study to investigate whether code defects lead to unnaturalness of source code. Different from existing studies, we leverage multiple large-scale and high-quality bug repositories where bug-irrelevant changes in bug-fixing commits have been explicitly excluded. The leveraged software applications cover different programming languages, and the empirical study involves real-world software defects as well as defects injected automatically with well-known mutation operators. On the one side, our evaluation results confirm existing studies in that buggy source code lines are often less natural than bug-free ones. On the other side, our evaluation reveals some interesting new findings. First, fixing bugs does not significantly improve the naturalness of code lines and the fixed lines on average are as unnatural as buggy ones. This finding may suggest that software defects are not the root causes of source code’s unnaturalness although there does existing statistically significant correlation between software defects and source code’s naturalness. Second, defects in different programming languages have similar effect on source code’s naturalness. The conclusions (i.e., buggy code is less natural but fixing the bugs cannot improve source code’s naturalness) hold regardless of the programming languages. Third, injecting defects automatically by well-known mutation operators does not significantly reduce the naturalness of involved source code lines. This suggests that automatically injected defects may have a similar impact on the naturalness of source code as real-world defects inadvertently introduced by developers. Fourth, the detects’ impact on source code’s naturalness varies slightly among different categories of software defects. Although fixing bugs on average does not significantly improve the naturalness of involved source code, fixing “checking” related bugs does significantly improve the naturalness of source code. Finally, locating buggy code lines according to naturalness alone is inaccurate, resulting in extremely low precision (less than one percent). Yanjie Jiang, Hui Liu 0003, Yuxia Zhang, Weixing Ji, Hao Zhong 0001, Lu Zhang 0023 |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2024 | A Multi-View Deep Learning Method for Predicting Blood-Brain Barrier Permeability of PeptidesabstractThe blood-brain barrier (BBB) plays a crucial role in protecting brain health by acting as a barrier between the brain and blood vessels. This barrier also presents challenges for delivering peptide drugs to brain targets. There is a pressing need for computational methods to accurately predict the permeability of peptides across the BBB. However, existing approaches face challenges due to limited real experimentally data and incomplete molecular information within peptide sequences. In this paper, we introduce MultiB3Pred, a multi-view deep learning method designed to address these challenges. Our method makes three key contributions. Firstly, we employ a effective amino acid replacement strategy for data augmentation. Secondly, We utilize sequence embeddings from a biologically pretrained model ProtT5 [1], further refined by a Transformer to capture dependencies on our specific dataset, leading to better sequence representations for the sequence predictor. Lastly, we derive SMILES from the sequences and train a novel SMILES learner. Precisely, the physicochemical properties of the molecules with the graph representation captured by the graph neural network from the molecular graphs are integrated through multilayer perceptron. The predicted probabilities from two sub-predictors are averaged to obtain the final result. Experiments demonstrate that MultiB3Pred achieves state-of-the-art accuracy and Matthews correlation coefficient of 94.4% and 89.9% respectively, showcasing its excellent performance in predicting blood-brain barrier penetration. At the same time, the stability of the model is confirmed by the good results of the 5-fold crossover experiment. Yizhuo Wang 0001, Chunfeng Li, Weixing Ji |
BIBM | 4 |
| 2024 | Load Balancing Optimizations for Distributed GMRES Algorithm
Shuaizhe Guo, Jianhua Gao 0001, Weixing Ji, Yizhuo Wang 0001 |
ICA3PP (6) | 4 |
| 2024 | JLeaks: A Featured Resource Leak Repository Collected From Hundreds of Open-Source Java ProjectsabstractHigh-quality defect repositories are vital in defect detection, localization, and repair. However, existing repositories collected from open-source projects are either small-scale or inadequately labeled and packed. This paper systematically summarizes the programming APIs of system resources (i.e., file, socket, and thread) in Java. Additionally, this paper demonstrates the exceptions that may cause resource leaks in the chained and nested streaming operations. A semi-automatic toolchain is built to improve the efficiency of defect extraction, including automatic building for large legacy Java projects. Accordingly, 1,094 resource leaks were collected from 321 open-source projects on GitHub. This repository, named JLeaks, was built by round-by-round filtering and cross-validation, involving the review of approximately 3,185 commits from hundreds of projects. JLeaks is currently the largest resource leak repository, and each defect in JLeaks is well-labeled and packed, including causes, locations, patches, source files, and compiled bytecode files for 254 defects. We have conducted a detailed analysis of JLeaks for defect distribution, root causes, and fix approaches. We compare JLeaks with two well-known resource leak repositories, and the results show that JLeaks is more informative and complete, with high availability, uniqueness, and consistency. Additionally, we show the usability of JLeaks in two application scenarios. Future studies can leverage our repository to encourage better design and implementation of defect-related algorithms and tools. Weixing Ji, Wuhuang Yao, Yizhuo Wang 0001, Hui Liu 0003, Haiyang Peng |
ICSE | 2 |
| 2024 | ConvDarts: a fast and exact convolutional algorithm selector for deep learning frameworks
Weixing Ji, Qinyuan Li, Xilai Yao, Wanyi Zhu |
CCF Trans. High Perform. Comput. | 2 |
| 2024 | pSpMv: precision-based sparse matrix partition and SpMV optimizationabstractAbstract The new generation of computing devices tends to support multiple floating-point formats and different computing precision. Besides single and double precision, half precision is embraced and widely supported by new computing devices. Low-precision representations have compact memory size and lightweight computing strength, and they also bring opportunities to the optimization of BLAS routines. This paper proposes a new sparse matrix partition approach based on IEEE 754 standard floating-point format. An input sparse matrix in double precision is partitioned and transformed into several sub-matrices in different precision without loss of accuracy. Most non-zero elements can be stored in half or single precision, if the most significant bits of exponent and the least significant bits of mantissa are zeros in double-precision representation. Based on this mixed-precision representation of sparse matrix, we also present a new SpMV algorithm pSpMV for GPU devices. pSpMV not only reduces the memory access overhead, but also reduces the computing strength of floating-point numbers. Experimental results on two GPU devices show that pSpMV achieves a geometric mean speedup of 1.39x on Tesla V100 and 1.45x on Tesla P100 over double-precision SpMV for 2,554 sparse matrices. Yizhuo Wang 0001, Jianhua Gao 0001, Weixing Ji |
CCF Trans. High Perform. Comput. | 4 |
| 2024 | Revisiting thread configuration of SpMV kernels on GPU: A machine learning based approach
Jianhua Gao 0001, Weixing Ji, Yizhuo Wang 0001, Feng Shi 0009 |
J. Parallel Distributed Comput. | 2 |
| 2024 | Optimization of Large-Scale Sparse Matrix-Vector Multiplication on Multi-GPU SystemsabstractSparse matrix-vector multiplication (SpMV) is one of the important kernels of many iterative algorithms for solving sparse linear systems. The limited storage and computational resources of individual GPUs restrict both the scale and speed of SpMV computing in problem-solving. As real-world engineering problems continue to increase in complexity, the imperative for collaborative execution of iterative solving algorithms across multiple GPUs is increasingly apparent. Although the multi-GPU-based SpMV takes less kernel execution time, it also introduces additional data transmission overhead, which diminishes the performance gains derived from parallelization across multi-GPUs. Based on the non-zero elements distribution characteristics of sparse matrices and the tradeoff between redundant computations and data transfer overhead, this article introduces a series of SpMV optimization techniques tailored for multi-GPU environments and effectively enhances the execution efficiency of iterative algorithms on multiple GPUs. First, we propose a two-level non-zero elements-based matrix partitioning method to increase the overlap of kernel execution and data transmission. Then, considering the irregular non-zero elements distribution in sparse matrices, a long-row-aware matrix partitioning method is proposed to hide more data transmissions. Finally, an optimization using redundant and inexpensive short-row execution to exchange costly data transmission is proposed. Our experimental evaluation demonstrates that, compared with the SpMV on a single GPU, the proposed method achieves an average speedup of 2.00× and 1.85× on platforms equipped with two RTX 3090 and two Tesla V100-SXM2, respectively. The average speedup of 2.65× is achieved on a platform equipped with four Tesla V100-SXM2. Jianhua Gao 0001, Weixing Ji, Yizhuo Wang 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | Optimization of Sparse Matrix Computation for Algebraic Multigrid on GPUsabstractAMG is one of the most efficient and widely used methods for solving sparse linear systems. The computational process of AMG mainly consists of a series of iterative calculations of generalized sparse matrix-matrix multiplication (SpGEMM) and sparse matrix-vector multiplication (SpMV). Optimizing these sparse matrix calculations is crucial for accelerating solving linear systems. In this paper, we first focus on optimizing the SpGEMM algorithm in AmgX, a popular AMG library for GPUs. We propose a new algorithm called SpGEMM-upper, which achieves an average speedup of 2.02× on Tesla V100 and 1.96× on RTX 3090 against the original algorithm. Next, through experimental investigation, we conclude that no single SpGEMM library or algorithm performs optimally for most sparse matrices, and the same holds true for SpMV. Therefore, we build machine learning-based models to predict the optimal SpGEMM and SpMV used in the AMG calculation process. Finally, we integrate the prediction models, SpGEMM-upper, and other selected algorithms into a framework for adaptive sparse matrix computation in AMG. Our experimental results prove that the framework achieves promising performance improvements on the test set. Yizhuo Wang 0001, Fangli Chang, Bingxin Wei, Jianhua Gao 0001, Weixing Ji |
ACM Trans. Archit. Code Optim. | 5 |
| 2023 | An Automatic Deployment Method for Hybrid Cloud Simulation Platform
Xilai Yao, Yizhuo Wang 0001, Weixing Ji, Qiurui Chen |
APPT | 3 |
| 2023 | An Automated Approach to Extracting Local VariablesabstractExtract local variable is a well-known and widely used refactoring. It is frequently employed to replace one or more occurrences of a complex expression with simple accesses to a newly added variable. Although most IDEs provide tool support for extract local variables, such tools without deep analysis of the refactorings may result in semantic errors. To this end, in this paper, we propose a novel and more reliable approach, called ValExtractor, to conduct extract variable refactorings automatically. The major challenge of automated extract local variable refactorings is how to efficiently and accurately identify the side effect of the extracted expressions and the potential interaction between the extracted expressions and their contexts without time-consuming dynamic execution of the involved programs. To resolve this challenge, ValExtractor leverages a lightweight static source code analysis to validate the side effect of the selected expression, and to identify which occurrences of the selected expression could be extracted together without changing the semantics of the program or introducing potential new exceptions. Our evaluation results on open-source Java applications suggest that Eclipse and IntelliJ IDEA, the state-of-the-practice refactoring engines, resulted in a large number of faulty extract variable refactorings whereas ValExtractor successfully avoided all such errors. The proposed approach has been merged into (and distributed with) Eclipse to improve the safety of extract local variable refactoring. Xiaye Chi, Hui Liu 0003, Guangjie Li, Weixiao Wang, Yunni Xia, Yanjie Jiang, Yuxia Zhang, Weixing Ji |
ESEC/SIGSOFT FSE | 8 |
| 2022 | Do bugs lead to unnaturalness of source code?abstractTexts in natural languages are highly repetitive and predictable because of the naturalness of natural languages. Recent research validated that source code in programming languages is also repetitive and predictable, and naturalness is an inherent property of source code. It was also reported that buggy code is significantly less natural than bug-free one, and bug fixing substantially improves the naturalness of the involved source code. In this paper, we revisit the naturalness of buggy code and investigate the effect of bug-fixing on the naturalness of source code. Different from the existing investigation, we leverage two large-scale and high-quality bug repositories where bug-irrelevant changes in bug-fixing commits have been explicitly excluded. Our evaluation results confirm that buggy lines are often less natural than bug-free ones. However, fixing bugs could not significantly improve the naturalness of involved code lines. Fixed lines on average are as unnatural as buggy ones. Consequently, bugs are not the root cause of the unnaturalness of source code, and it could be inaccurate to identify buggy code lines solely by the naturalness of source code. Our evaluation results suggest that the naturalness-based buggy line detection results in extremely low precision (less than one percentage). Yanjie Jiang, Hui Liu 0003, Yuxia Zhang, Weixing Ji, Hao Zhong 0001, Lu Zhang 0023 |
ESEC/SIGSOFT FSE | 4 |
| 2022 | TaiChi: A Hybrid Compression Format for Binary Sparse Matrix-Vector Multiplication on GPUabstractBinary Sparse Matrix-Vector Multiplication (SpMV) is a heavy computational kernel in weblink analysis, integer factorization, compressed sensing, spectral graph theory, and other domains. Testing several popular GPU-based SpMV implementations on 400 sparse matrices, we observed that data transfer to GPU memory accounts for a large part of the total computation time. The transfer of constant value “1”s can be easily eliminated for binary sparse matrices. However, compressing index arrays has always been a great challenge. This article proposes a new compression format TaiChi to further reduce index data copies and improve the performance of SpMV, especially for diagonally dominant binary sparse matrices. Input matrices are first partitioned into relatively dense and ultra-sparse areas. Then the dense areas are encoded inversely by marking “0”s, while the ultra-sparse area is encoded by marking “1”s. We also designed a new SpMV algorithm only using addition and subtraction for binary matrices based on our partition and encoding format. Evaluation results on real-world binary sparse matrices show that our hybrid encoding for binary matrix significantly reduces the data transfer and speeds up the kernel execution. It achieves the highest transfer and kernel execution speedups of 5.63x and 3.84x on GTX 1080 Ti, 3.39x and 3.91x on Tesla V100. Jianhua Gao 0001, Weixing Ji, Zhaonian Tan, Yizhuo Wang 0001, Feng Shi 0009 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | AMF-CSR: Adaptive Multi-Row Folding of CSR for SpMV on GPUabstractSpMV is a cost-dominant operation used in many iterative methods for solving large-scale sparse linear systems. However, irregular memory access of SpMV to the multiplied vector leads to low data locality and then harms the performance. This paper presents an adaptive multi-row folding of CSR (AMF-CSR) format for SpMV calculation on GPU. This new storage format supports the folding of the variable number of rows in order to achieve better load balancing in computation. AMF-CSR not only increases the density of non-zero elements in a folded row, thereby improving the access locality of the multiplied vector, but also merges an approximately equal number of nonzero elements in a folded row, hence achieving load balancing. The performance evaluation using 28 sparse matrices shows that the proposed SpMV algorithm based on AMF-CSR achieves the highest speedup of 4.11x and 3.62x on GTX 1080 Ti and Tesla V100 respectively against a fixed multi-row folding-based SpMV algorithm. Evaluation results using 450 regular sparse matrices and 450 irregular sparse matrices also show that AMF-CSR is superior to other SpMV implementations. Jianhua Gao 0001, Weixing Ji, Senhao Shao, Yizhuo Wang 0001, Feng Shi 0009 |
ICPADS | 2 |
| 2021 | Towards Optimal Fast Matrix Multiplication on CPU-GPU Platforms
Senhao Shao, Yizhuo Wang 0001, Weixing Ji, Jianhua Gao 0001 |
PDCAT | 3 |
| 2020 | Identification of Misleading Location Information in Compiler DiagnosesabstractThe location of compilation errors are usually reported by compilers to facilitate quick fixing of such compiler errors. However, sometimes such location information could be incorrect or misleading, which significantly reduces the chance of quick fixing. To this end, in this paper, we propose an automated approach, called iMiLi, to identify misleading location information in compiler diagnoses. We generate potentially illegal programs (called mutants) by mutating legal programs, and compile such mutants. If the compiler generates error diagnoses on a mutant, we extract the location information from the resulting diagnoses. The location information is suspicious if it does not point to the source code where the associated mutation is conducted. Then we propose heuristics for each kind of mutation operators to exclude such suspicious but correct location information. We evaluate the proposed approach on a state-of-the-practice compiler (i.e., Eclipse Compiler for Java, known as ECJ). iMiLi successfully identifies seven categories of incorrect/misleading location information in diagnoses of ECJ. Miaoying Wang, Weixing Ji, Dejiang Jing, Hui Liu 0003 |
APSEC | 2 |
| 2020 | MMSparse: 2D partitioning of sparse matrix based on mathematical morphology
Zhaonian Tan, Weixing Ji, Jianhua Gao 0001, Yueyan Zhao, Akrem Benatia, Yizhuo Wang 0001, Feng Shi 0009 |
Future Gener. Comput. Syst. | 2 |
| 2020 | Cube-based incremental outlier detection for streaming computing
Jianhua Gao 0001, Weixing Ji, Anmin Li, Yizhuo Wang 0001, Zongyu Zhang |
Inf. Sci. | 2 |
| 2020 | Attentive boundary aware network for multi-scale skin lesion segmentation with adversarial training
Zenghui Wei, Feng Shi 0009, Weixing Ji, Guanghui Han |
Multim. Tools Appl. | 4 |
| 2018 | Exploiting Task-Based Parallelism for Parallel Discrete Event SimulationabstractToday large-scale simulation applications are becoming common in research and industry. A significant fraction of them run on multi-core clusters. Current parallel simulation kernels use multi-process and multi-thread to exploit inter-node parallelism and intra-node parallelism on multi-core clusters. We exploit task-base parallelism in parallel discrete event simulation (PDES) kernels, which is more fine-grained than thread-level and process-level parallelism. In our system, every simulation event is wrapped to a task. Work-stealing task scheduling scheme is applied to achieve dynamic load balancing among the multi-cores, and a graph partitioning approach is applied in partitioning simulation entities among the cluster nodes. Experimental results show that our PDES kernel outperforms existing PDES kernels by fully exploiting task parallelism. Yizhuo Wang 0001, Zhiwei Gao 0001, Weixing Ji, Duzheng Qing |
PDP | 3 |
| 2018 | BestSF: A Sparse Meta-Format for Optimizing SpMV on GPUabstractThe Sparse Matrix-Vector Multiplication (SpMV) kernel dominates the computing cost in numerous scientific applications. Many implementations based on different sparse formats were proposed to improve this kernel on the recent GPU architectures. However, it has been widely observed that there is no “best-for-all” sparse format for the SpMV kernel on GPU. Indeed, serious performance degradation of an order of magnitude can be observed without a careful selection of the sparse format to use. To address this problem, we propose in this article BestSF (Best Sparse Format), a new learning-based sparse meta-format that automatically selects the most appropriate sparse format for a given input matrix. To do so, BestSF relies on a cost-sensitive classification system trained using Weighted Support Vector Machines (WSVMs) to predict the best sparse format for each input sparse matrix. Our experimental results on two different NVIDIA GPU architectures using a large number of real-world sparse matrices show that BestSF achieved a noticeable overall performance improvement over using a single sparse format. While BestSF is trained to select the best sparse format in terms of performance (GFLOPS), our further experimental investigations revealed that using BestSF also led, in most of the test cases, to the best energy efficiency (MFLOPS/W). To prove its practical effectiveness, we also evaluate the performance and energy efficiency improvement achieved when using BestSF as a building block in a GPU-based Preconditioned Conjugate Gradient (PCG) iterative solver. Akrem Benatia, Weixing Ji, Yizhuo Wang 0001, Feng Shi 0009 |
ACM Trans. Archit. Code Optim. | 2 |
| 2017 | Exploring grouped coherence for clustered hierarchical cache
Sensen Hu, Feng Shi 0009, Weixing Ji, Xu Chen 0016, Shahnawaz Talpur |
J. Supercomput. | 3 |
| 2016 | Machine Learning Approach for the Predicting Performance of SpMV on GPUabstractSparse Matrix-Vector Multiplication (SpMV) kernel dominates the computing cost in numerous scientific applications. Many implementations based on different sparse formats were proposed recently for optimizing this kernel on the GPU side. Since the performance of the SpMV varies significantly according to the sparsity characteristics of the input matrix and the hardware features, developing an accurate performance model for this kernel is a challenging task. The traditional approach of building such models by analytical modeling is difficult in practice and requires a thorough understanding of the interaction between the GPU hardware and the sparse code. In this paper, we propose to use a machine learning approach to predict the performance of the SpMV kernel using several sparse formats (COO, CSR, ELL, and HYB) on GPU. We used two popular machine learning algorithms, Support Vector Regression (SVR) and Multilayer Perceptron neural network (MLP). Our experimental results on two different GPUs (Fermi GTX 512 and Maxwell GTX 980 Ti) show that the SVR models deliver the best accuracy with average prediction error ranging between 7% and 14%. Akrem Benatia, Weixing Ji, Yizhuo Wang 0001, Feng Shi 0009 |
ICPADS | 2 |
| 2016 | Sparse Matrix Format Selection with Multiclass SVM for SpMV on GPUabstractSparse Matrix-Vector Multiplication (SpMV) kernel dominates the computing cost in numerous scientific applications. Many implementations based on different sparse formats were proposed recently for this kernel on the GPU side. Since the performance of these sparse formats varies significantly according to the sparsity characteristics of the input matrix and the hardware specifications, no one of them can be considered as the best one to use for every sparse matrix. In this paper, we address the problem of selecting the best representation for a given sparse matrix on GPU by using a machine learning approach. First, we present some interesting and easy to compute features for characterizing the sparse matrices on GPU. Second, we use a multiclass Support Vector Machine (SVM) classifier to select the best format for each input matrix. We consider in this paper four popular formats (COO, CSR, ELL, and HYB), but our work can be extended to support more sparse representations. Experimental results on two different GPUs (Fermi GTX 580 and Maxwell GTX 980 Ti) show that we achieved more than 98% of the performance possible with a perfect selection. Akrem Benatia, Weixing Ji, Yizhuo Wang 0001, Feng Shi 0009 |
ICPP | 2 |
| 2016 | Profiling and analysis of object lazy allocation in Java programsabstractLazy allocation strategy allows the memory management system to defer the space allocation action of objects until they are being accessed. This paper investigates the potential benefits of a lazy allocator for Java applications. A heap tracing tool is implemented by instrumenting an existing Java virtual machine HotSpot, which records useful object manipulating events at runtime. By profiling and analyzing a large number of benchmarks, we show the potential dynamic memory management optimization opportunity in Java programs. We also designed a simulation system to demonstrate the actual effects of a lazy allocator. Weixing Ji, Yujin Gao, Duzheng Qing |
SNPD | 2 |
| 2015 | Task Parallel Implementation of Matrix Multiplication on Multi-socket Multi-core Architectures
Yizhuo Wang 0001, Weixing Ji, Xu Chen 0016, Sensen Hu |
ICA3PP (3) | 2 |
| 2015 | Memory-Aware NoC Application Mapping Based on Adaptive Genetic Algorithm
Yizhuo Wang 0001, Zhibiao Zhang, Lifu Huang, Weixing Ji |
ICA3PP (1) | 4 |
| 2015 | Refactoring for Separation of Concurrent Concerns
Yang Zhang 0037, Dongwen Zhang, Weixing Ji, Yizhuo Wang 0001 |
ICA3PP (3) | 3 |
| 2014 | Dynamic enforcement of determinism in a parallel scripting languageabstractDeterminism is an appealing property for parallel programs, as it simplifies understanding, reasoning and debugging. It is particularly appealing in dynamic (scripting) languages, where ease of programming is a dominant design goal. Some existing parallel languages use the type system to enforce determinism statically, but this is not generally practical for dynamic languages. In this paper, we describe how determinism can be obtained---and dynamically enforced/verified---for appropriate extensions to a parallel scripting language. Specifically, we introduce the constructs of Deterministic Parallel Ruby (DPR), together with a run-time system (Tardis) that verifies properties required for determinism, including correct usage of reductions and commutative operators, and the mutual independence (data-race freedom) of concurrent tasks. Experimental results confirm that DPR can provide scalable performance on multicore machines and that the overhead of Tardis is low enough for practical testing. In particular, Tardis significantly outperforms alternative data-race detectors with comparable functionality. We conclude with a discussion of future directions in the dynamic enforcement of determinism. Weixing Ji, Michael L. Scott |
PLDI | 2 |
| 2014 | An adaptive and hierarchical task scheduling scheme for multi-core clusters
Yizhuo Wang 0001, Yang Zhang 0037, Xiaojun Wang 0005, Xu Chen 0016, Weixing Ji, Feng Shi 0009 |
Parallel Comput. | 6 |
| 2013 | A work-stealing scheduling framework supporting fault toleranceabstractFault tolerance and load balancing are critical points for executing long-running parallel applications on multicore clusters. This paper addresses both fault tolerance and load balancing on multicore clusters by presenting a novel work-stealing task scheduling framework which supports hardware fault tolerance. In this framework, both transient and permanent faults are detected and recovered at task granularity. We incorporate task-based fault detection and recovery mechanisms into a hierarchical work-stealing scheme to establish the framework. This framework provides low-overhead fault-tolerance and optimal load balancing by fully exploiting task parallelism. Yizhuo Wang 0001, Weixing Ji, Feng Shi 0009, Qi Zuo |
DATE | 2 |
| 2012 | Exploring Object-Level Parallelism on Chip Multi-processors
Weixing Ji, Yizhuo Wang 0001, Junqing Zhao |
ICA3PP (2) | 1 |
| 2012 | Knowledge-Based Adaptive Self-Scheduling
Yizhuo Wang 0001, Weixing Ji, Feng Shi 0009, Qi Zuo, Ning Deng 0002 |
NPC | 2 |
| 2012 | A Hierarchical Work-Stealing Framework for Multi-core ClustersabstractWork-stealing has been widely used in task-based parallel programming for dynamic load balancing. The overhead of work-stealing on distributed memory systems is much higher than that on shared memory systems. To minimize the overhead of work-stealing on a multi-core cluster, we propose a hierarchical work-stealing framework, in which work-stealing is performed inside a node before across the node boundary. Two key techniques used in our framework to reduce the inter-node steals are: a) adaptive initial partitioning for different task parallel patterns; b) centralized control for inter-node work-stealing, which improves the efficiency of victim selection and termination detection. We compare our technique to the classical work-stealing scheme and a state-of-the-art work-stealing scheme [1] for multi-core clusters. Our technique outperforms them by 19% and 8% respectively. Yizhuo Wang 0001, Weixing Ji, Qi Zuo, Feng Shi 0009 |
PDCAT | 2 |
| 2012 | A scalable method-level parallel library and its improvement
Yang Zhang 0037, Weixing Ji |
J. Supercomput. | 2 |
| 2011 | A Semi-automatic Scratchpad Memory Management Framework for CMP
Ning Deng 0002, Weixing Ji, Qi Zuo |
APPT | 2 |
| 2011 | Floorplanning exploration and performance evaluation of a new Network-on-ChipabstractThe Network-on-Chip (NoC) paradigm has emerged as a revolutionary methodology in current System-on-Chips (SoCs) for integrating a large number of processing elements in a single die. It has the advantage of enhanced performance, scalability and modularity, compared with previous bus-based communication architectures. Recently, A new Triplet-based Hierarchical Interconnection Network (THIN) has been proposed. In this paper, we explore the three-dimensional (3D) floor-planning of THIN and present two different floorplanning and routing methods using both the Manhattan routing and the Y-architecture routing architectures. A cycle-accurate simulator is developed based on Noxim NoC simulator and ORION 2.0 energy model. The latency, power consumption and area requirement of both THIN and Mesh are evaluated. The experimental results indicate that the proposed design provides 24.95% reduction in average power consumption and 16.84% improvement in area requirement. Licheng Xue, Weixing Ji, Qi Zuo |
DATE | 2 |
| 2011 | Dynamic and adaptive SPM management for a multi-task environment
Weixing Ji, Ning Deng 0002, Feng Shi 0009, Qi Zuo |
J. Syst. Archit. | 1 |
| 2009 | Group-caching for NoC based multicore cache coherent systemsabstractMost CMPs use on-chip networks to connect cores and tend to integrate more simple cores on a single die. Low-radix networks, such as 2D-MESH, are widely used in tiled CMPs since they can be mapped to on-chip networks efficiently. However, low-radix networks introduce high network latency caused by long diameter. In this paper, we propose the use of group-caching design in NoC based multicore cache coherent systems. In our design, on-chip L2 banks are organized to form multiple groups. Each cache group behaves like a shared L2 cache for the cores inside cache group while the cache coherence between cache groups is maintained by coherence messages. Besides, group-caching also adopts the new cache replacement policy to improve the inefficient use of the aggregate L2 cache capacity. Compared to banked and shared L2 design, as most L2 accesses are served by local cache group, the hop count is significantly reduced. Experiment results based on full-system simulation show that for 2D-MESH, group-caching can increase the performance by 2%∼8% compared to banked and shared L2 design, with network energy consumption reduced by 11%∼13%. Experiment results also show that the communication overhead inside cache group plays an important role in the performance of groupcaching. Feng Shi 0009, Qi Zuo, Weixing Ji, Ning Deng 0002, Licheng Xue, Yu-an Tan 0001 |
DATE | 4 |
| 2009 | Storage Architecture for an On-chip Multi-core ProcessorabstractModern multi-core processor architectures strive for the highest possible performance of various applications. This paper discusses a triple-based multi-core architecture which supports object-oriented methodology and applications in hardware level. However, the Memory Wall is still the bottleneck which should be resolved to decrease the disparity between how fast a CPU can operate on data and how fast it can get data. We present hierarchical shared memory system architecture (HSM) which is hierarchically constructed memory shared by multi-cores. Moreover, we propose a new approach mapping data among different levels of cache and memory, which is called partially-inclusive cache mapping policy that facilitates the coherence of shared memory. This paper focus on object-oriented systems combined with the HSM and partially-inclusive policy and presents a new objects management model. The analysis based on comparisons between our objects management and link structured object organization methods shows that our method is predominant in spatial and temporal aspects on memory parallel access efficiency and costs less storage space to organize objects. Mengxiao Liu, Weixing Ji, Xing Pu |
DSD | 2 |
| 2009 | N-port memory mapping for LUT-based FPGAsabstractAs current FPGAs grow in logic capacity, they are widely used to implement entire systems. In some specific applications, such as our embedded multi-core processor TriBA[1],user memory models are not limited to single-port or dual-port. Thus, we need a cost-effective way to realize N-port memory on FPGA since most commercial products do not provide N-port physical arrays. In this paper, we propose a hierarchical N-port memory architecture for LUT-based FPGAs. The principle of this architecture is to create a two-level memory hierarchy formed by different resources. We map the memory resources inside LUTs as 1-port memory banks, and interleave these banks to create N-port L1 memory. We also interleave physical dual-port arrays to build N-port L2 memory. We also provide the data transfer between L1 and L2 memories and assume that such data transfer is managed by software control just like the strategy used by SPM. Compared to L1 memory, L2 memory has the advantage in cost and also has several disadvantages, such as longer access time and higher conflict probability. If most accesses are served by its L1 memory portion, hierarchical memory architecture will achieve both goals in cost and access time. We implement this architecture on Xilinx Virtex-II chips to measure its cost and also use the memory trace collected from multi-core simulator to measure its average access time. The product of cost and average access time shows that, hierarchical memory architecture is a cost-effective way to realize N-port memory on FPGA. Feng Shi 0009, Qi Zuo, Weixing Ji, Mengxiao Liu |
FPGA | 4 |
| 2009 | A Parallel Memory System Model for Multi-core ProcessorabstractModern multi-core processors are predominant in improving performance of parallel applications. This paper discusses a triple-based multi-core architecture which provides native support for object-oriented methodology and applications in hardware level. The model explicitly represents objects and supports messaging-based communication, which maps well to the standard style of interaction in object oriented languages. However, the Memory Wall is still the bottleneck which should be resolved to decrease the disparity between how fast a CPU can operate on data and how fast it can get data. A hierarchy shared memory system (HSM) working with the partially-inclusive cache mapping policy is proposed. And a new object management model is presented, which use object table and recycle stack scheme to supports explicit dynamic object management. Our cache design presents an innovative solution to handling the costs of cache coherence by allowing applications to control the amount of sharing between cores. Experimental analysis based on comparisons between our objects management and other common link structured object organization methods shows that our method is predominant in spatial and temporal aspects on memory parallel access efficiency and costs less storage space to organize objects. Mengxiao Liu, Weixing Ji, Xing Pu |
NAS | 2 |
| 2009 | A Novel Adaptive Scratchpad Memory Management StrategyabstractScratchpad Memory (SPM) is a fast and small software-managed SRAM. Its current extensive uses in embedded processors are motivated by the advantages of power saving, small area and low access time compared with cache. However, existing SPM management methods depend heavily on profiling and compilers. The dependence on compiler also makes embedded applications hard to transplant. This paper presents a novel strategy to manage the scratchpad memory without compiler support. Based on the memory reference locality theory, a hardware random sampling module is adopted to dynamically identify the frequently accessed addresses at runtime. The consequential data movement and address redirection are handled by software operation with the assistance of memory management unit (MMU). We evaluate our method on 10 typical embedded applications and compare the results to a cache reference system. Experimental results show that, on average, our scheme can achieve 33:5% reduction in energy consumption with only slight (<1%) decrease in throughput versus the reference system. Ning Deng 0002, Weixing Ji, Feng Shi 0009, Yizhuo Wang 0001 |
RTCSA | 2 |
| 2008 | A state machine approach for problem detection in large-scale distributed systemabstractEfficient problem detection methods play an important role in system management. In this paper, a formal method is described for problem detection in large scale and distributed enterprise IT environment. Events from distributed system components are collected, filtered and correlated. Leveraging these correlated events, the behavior of a distributed system is presented as a problem detection state machine (PDSM). PDSM is built up automatically from system logs without any specification of the target system. This approach combines logs from multi-sources and does not require any human involved or experimental instructions. It is generally applicable to a large class of distributed systems. Experimental results show that the implementation of PDSM performs problem detection efficiently in typical distributed enterprise systems. Kewei Sun, Jie Qiu 0001, Ying Li 0012, Ying Chen 0004, Weixing Ji |
NOMS | 5 |
| 2007 | The Design of a Novel Object-oriented Processor : OOMIPSabstractA novel object-oriented processor is proposed in this paper, which provides support for object addressing, message passing and dynamic memory management. Object running on this processor has its own control thread and communicates with others via messages. A virtual addressed object cache that reduces the indirection overhead while maintaining the efficiency of object relocation is presented. Object table that maintains the handles is used to obtain the actual object location on an object cache miss. Hardware support for explicit dynamic memory management is provided. Object allocation and deletion is strictly bounded in time. Moreover, a new concurrently dynamic memory management algorithm is proposed, which enables the processor to freely access heap during memory compaction and the applications will not be suspended for the completion of memory compaction. Weixing Ji, Feng Shi 0009 |
ASAP | 1 |
| 2007 | A Triplet-based Computer Architecture Supporting Parallel Object ComputingabstractA real scalable triplet-based computer architecture TriBA is proposed in this paper. TriBA is an object-oriented chip multi-processor that supports truly parallel execution of objects from hardware. Cores on the same chip are connected via triplet-based hierarchical interconnection network (THIN), which has simple topology and computing locality characteristic. A distributed deterministic routing algorithm (DDRA) is elaborated, already proposed for THIN. Runtime objects are mapped to processor cores according to their coupling degree which can also transfer to idle cores if needed. TriBA achieves the unification of software architecture and computer, and also relieves the burden of parallel programming. Feng Shi 0009, Weixing Ji, Haroon-ul-Rashid |
ASAP | 2 |
| 2007 | A self-maintained memory module supporting DMMabstractThe memory intensive nature of object-oriented languages such as C++ and Java has created the need of a high-performance dynamic memory management (DMM); however, it is a challenging task to provide efficient reliable system without violating real time performance constraints. Hardware approach emerges as one of the candidate in improving the performance of DMM. This paper presents an efficient design for explicit dynamic memory management which exploits the high speed of a pure hardware implementation. Object allocation and deletion are strictly bounded in time. The whole heap space is divided into two semi-spaces, and a concurrent bidirectional memory compaction algorithm is proposed. So that memory compaction can be done while mutator process is running on the processor. A small built in object-based cache memory is available to avoid indirect object addressing inefficiencies. Experiments show that this hardware scheme can greatly improve the speed and predictability of DMM. Weixing Ji, Feng Shi 0009 |
CASES | 1 |
| 2007 | THIN: A New Hierarchical Interconnection Network-on-Chip for SOC
Feng Shi 0009, Weixing Ji |
ICA3PP | 3 |
| 2007 | Performance Evaluation of a Self-Maintained Memory ModuleabstractHardware approach emerges as one of the candidate in improving the performance of dynamic memory management. This paper presents measurements of a self-maintained memory module subjected to several different workloads. This memory module supporting explicit dynamic memory management takes advantage of the high speed of a pure hardware implementation. Object allocation and deletion are strictly bounded in time. The whole heap space is divided into two semi-spaces, and a concurrent bidirectional memory compaction algorithm is exploited, so that memory compaction can be done while mutator process is running on the processor concurrently. Reported measurements demonstrate that hardware-assisted memory management is a viable alternative to traditional explicit memory management techniques. Experimental results show that more than 60% of memory traffic is saved by the proposed memory compaction scheme compared to software-only approach. Both processor delay and program execution time are greatly reduced. Weixing Ji, Feng Shi 0009, Qi Zuo |
RTSS | 1 |