VLDB 2026 Research / reviewers in the wild / expert
ZhiLei Chai
dblp:43/2840 · also Zhilei Chai
· DBLP profile ↗
36ranked-venue papers
2as first author
22since 2021 · last 2026
0000-0003-3822-1653ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 10 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 2Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Configurable Streaming Accelerator for LUT-Based Super-Resolution on FPGA
Xuzhuo Hu, Jianyang Ding, Huachen Zhang, Tianshuo Lu, ZhiLei Chai |
APPT | 6 |
| 2026 | L-PCN: A Point Cloud Accelerator Exploiting Spatial Locality through Octree-Based Islandization
Jieming Yin, ZhiLei Chai, Jiliang Zhang 0011, Herman Lam |
ISCA | 5 |
| 2026 | RACP: An Efficient RISC-V Domain-Specific Processor for Arbitrary-Size Kernel CNNs
Jianyang Ding, Tianshuo Lu, Huachen Zhang, Xuzhuo Hu, ZhiLei Chai |
ISCAS | 6 |
| 2026 | An efficient RISC-V processor with customized instruction set for sparse DNN acceleration on embedded system
Jianyang Ding, Huachen Zhang, Tianshuo Lu, ZhiLei Chai |
J. Syst. Archit. | 5 |
| 2025 | Hybrid-SANet: Hybrid Self-attention Transformer for Efficient Image Super-Resolution
Jianyang Ding, Huachen Zhang, Nachuan Zhang, Tianshuo Lu, ZhiLei Chai |
CGI (3) | 6 |
| 2025 | EVO-QNN: Efficient Mixed-Precision Quantization Inference on RISC-V-Based Edge DeviceabstractMixed-Precision Quantized Neural Network (MPQNN) helps balance inference precision and efficiency under resource constraints, while most of them lack high-energy-efficiency hardware acceleration solutions. To address these challenges, we propose a SW/HW co-design framework termed EVO-QNN for low-energy and low-latency inference. Specifically, EVO-QNN manages bit-widths of operators and extends SIMD instructions based on customized RISC-V core. Experimental results demonstrate that our framework can achieve performance improvement ranging from 1.23x to 1.58x, with only a 1.02% and 6.61% increasing in area and power consumption for 2–8 bit convolution operators. Tianshuo Lu, Jianyang Ding, Huachen Zhang, ZhiLei Chai |
FCCM | 6 |
| 2025 | RV-ESMC: Efficient Sparse Matrix Convolution Processor based on RISC-V Custom instructions for Edge PlatformsabstractAs the demand for deep neural network (DNN) inference on edge platforms grows, deploying compute-intensive DNNs on resource-constrained devices remains challenging. This paper proposes a novel sparse convolution acceleration processor, RV-ESMC, based on RISC-V architecture, with custom instructions to enable efficient edge DNN inference. RV-ESMC provides flexibility by supporting inline assembly calls in C programming. Experimental results indicate that RV-ESMC can reduce execution time by over 70% in DNNs with convolution operations compared to conventional instruction sets. The functionality of RV-ESMC is validated on an FPGA platform and its performance is comprehensively evaluated based on a 55nm CMOS process. The results show that RV-ESMC can achieve a peak energy efficiency of 675 GOPS/W. Huachen Zhang, Jianyang Ding, Tianshuo Lu, ZhiLei Chai |
FCCM | 6 |
| 2025 | Multi-modal Information Enhancement for Long-Tailed Recognition
Shengnan Fan, ZhiLei Chai, Xiangyu Cheng, Yuying Pan |
ICIC (11) | 2 |
| 2025 | Lightweight Transformer with Enhanced Inverted Residual Blocks for Bird Sound RecognitionabstractBird sound recognition relies on capturing unique characteristics of avian vocalizations to achieve accurate species identification. Recent advances in deep learning have significantly improved classification accuracy and Transformer-based models stand out due to their superior ability to model long-range dependencies. However, most of them suffer from high computational complexity, posing challenges for deployment on edge devices in the wild. To address this issue, this paper proposes a lightweight Transformer-based model incorporating enhanced Inverted Residual Blocks (IRB). To this end, we first replace original Multilayer Perceptron (MLP) modules with IRB. Additionally, we propose a stage-wise incremental strategy for setting expansion factors to reduce redundancy. At the same time, we incorporate residual connections before and after depthwise convolutions to maintain model performance. Furthermore, we refine attention configurations throughout various phases of the network. We strategically minimize redundant attention in the initial stages, while intensifying its application in the crucial stages. This approach achieves an optimal balance between computation and accuracy. Experimental results on the Bird-CLEF2023, DCASE2020, and Birdsdata demonstrate that our proposed model can achieve approximately a 5-fold reduction in parameters, an 85% decrease in computational load, and a 2.7-fold increase in inference speed on Jetson AGX Xavier, while preserving high accuracy. Xiangyu Cheng, Shengnan Fan, Jianyang Ding, ZhiLei Chai |
IJCNN | 6 |
| 2025 | IR-OptSet: An Optimization-Sensitive Dataset for Advancing LLM-Based IR OptimizerabstractCompiler optimization is essential for improving program performance, yet modern compilers still depend on manually crafted transformation rules over intermediate representations (IRs). As compilers grow in complexity, maintaining these rule-based optimizations becomes increasingly labor-intensive and difficult to scale. Recent advances in large language models (LLMs) offer a promising alternative, but their effectiveness in compiler optimization remains limited—primarily due to the lack of IR-oriented datasets that expose models to diverse transformation samples in real-world scenarios (optimization-sensitive samples), hindering LLMs from learning rich and generalizable optimization strategies.In this paper, we introduce IR-OptSet, the first public optimization-sensitive dataset for advancing LLM-based IR optimizers. It comprises 170K LLVM IR samples from open-source repositories across 8 representative optimization domains. IR-OptSet defines two core tasks: Code Analysis and Optimized Code Generation, and provides tools for correctness verification, performance evaluation, and dataset expansion. In our experiments, fine-tuning three representative LLMs on IR-OptSet leads to significant accuracy improvements across both tasks. Moreover, the LLM fine-tuned with IR-OptSet outperforms traditional compiler with the -O3 option in 64 test cases in terms of performance. Further analysis reveals that IR-OptSet provides greater transformation diversity and representativeness than three widely used IR-oriented datasets, highlighting its potential to drive model-based IR optimization. IR-OptSet is publicly available at https://huggingface.co/datasets/YangziResearch/IR-OptSet. Lei Qiu 0007, Fang Lyu, Ming Zhong 0016, ZhiLei Chai, Haojie Zhou, Huimin Cui, Xiaobing Feng 0002 |
NeurIPS | 5 |
| 2025 | Accelerating large-scale multi-scalar multiplication in Zk-SNARK through exploiting its multilevel parallelism
Ning Wang 0039, Pengcheng Hua, ZhiLei Chai |
Integr. | 5 |
| 2025 | MaxSwap-Enhanced Knowledge Consistency Learning for long-tailed recognition
Shengnan Fan, ZhiLei Chai, Yuying Pan, Xiangyu Cheng |
Image Vis. Comput. | 2 |
| 2025 | Hybrid Self-Aligned Fusion With Dual-Weight Attention Network for Alzheimer's DetectionabstractDementia, particularly Alzheimer's disease (AD), affects millions of elderly individuals worldwide. Traditionally, interview data, including audio recordings and transcripts, is used to train Artificial Intelligence models for the automatic detection of AD patterns. In this work, we introduce a novel attention-weighted image set, where each image integrates text-image relevance with focused areas from the Cookie Theft picture, derived from the corresponding description. Furthermore, we propose a novel multimodal architecture, Hybrid Self-Aligned Fusion with Dual-Weight Attention Network (HSAF-DWAN), to predict AD, using audio recordings, transcripts, and corresponding attention-weighted images. This architecture consists of two key modules: an Intra-Modality Self-Alignment (IMSA) module, which captures relationships within a single modality, and a Dual-Weight Cross-Modality Attention (DW-CMA) module, which effectively fuses cross-modality data through a dual-weight mechanism, incorporating an optimized cross-attention and secondary weighting. Extensive experiments conducted on the Cookie Theft corpus from DementiaBank demonstrate that our method outperforms state-of-the-art models, achieving an accuracy of 86.71% and an F1 score of 88.15%. Ning Wang 0039, ZhiLei Chai |
IEEE Signal Process. Lett. | 4 |
| 2025 | Optimizing Sparse Matrix Convolution on RISC-V Core: Custom Instructions for Embedded SystemabstractWith the increasing demand for deep neural network (DNN) inference tasks on embedded platforms, deploying compute-intensive DNNs on resource-constrained embedded platforms faces challenges. While sparsification technology offers a potential solution, its implementation on edge platforms still faces difficulties. In this article, we propose a novel sparse convolution acceleration processor based on RISC-V architecture, and design specialized custom instructions to enable efficient edge DNN inference. To this end, we mainly address three technical issues. In response to numerical characteristics of sparse convolution, the designed processor can implement a hardware-friendly architecture that transforms convolutions into sparse matrix multiplication. Additionally, it employs a column-major and element-level parallel strategy to optimize load imbalance issues present in the Gustavson algorithm, thereby enhancing sparse matrix computations. To further improve computational efficiency, our work is designed by incorporating efficient execution units that reduce instruction execution overheads while minimizing memory access frequency. Compared to traditional accelerators, our work supports custom instruction formats in the C programming language, offering superior flexibility. Extensive experimental results indicate that our work can reduce execution time over 70% when running most DNNs with convolution operations compared to conventional instruction sets. Moreover, the functionality of our work is validated on an FPGA platform, and its performance is comprehensively evaluated based on a 55 nm CMOS process. The results show that our work can achieve a peak energy efficiency of 675 GOPS/W in most network inference tasks, demonstrating exceptional computational performance and energy efficiency. Huachen Zhang, Jianyang Ding, Tianshuo Lu, ZhiLei Chai |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2025 | ISRLUT: Integer-Only FHD Image Super-Resolution Based on Neural Lookup Table and Near-Memory ComputingabstractWhile Deep Neural Networks (DNNs) have achieved remarkable progress in Image Super-Resolution (SR) task, they face significant challenges for edge processing FHD images. Complex DNN operators lead to high hardware resource consumption and latency. Computational inefficiency of FPU increases energy consumption, while DDR access overhead and on-chip memory overflow further constrain real-time capabilities. To address this, we propose ISRLUT, a novel accelerator architecture focused on integer-only inference and near-memory computing. Its core contributions include: (1) Fusion of Neural LUT arithmetic with reconfigurable compute units, transforming unified LUT operators from DNN operators and enhancing hardware utilization; (2) An integer-only inference and parallel architecture, eliminating floating-point dependencies and significantly reducing energy consumption; (3) An innovative internal operator memory management scheme coupled with Tile-based Buffer Overlap and Private Cache Mechanism. We deploy ISRLUT on FPGA and ASIC platforms. Experiments demonstrate that ISRLUT achieves efficient performance: For 4 \(\times\) upscaling, it requires only 36.9 KB of storage and achieves a PSNR of 30.21 dB on Set5. Hardware implementation using a 55 nm ASIC consumes merely 0.0337 W power, delivers an energy efficiency of 7278.6 Mpixels/s/W, and achieves a real-time frame rate of 118 FPS for 4 \(\times\) FHD processing, validating its superiority in energy efficiency and hardware utilization. Tianshuo Lu, Jianyang Ding, Huachen Zhang, ZhiLei Chai |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2024 | GCMLP: A Lightweight Network for Gamut Compression
Xiaokai Du, ZhiLei Chai |
PRCV (8) | 5 |
| 2024 | Joint learning of foreground, background and edge for salient object detection
ZhiLei Chai, Guodong Guo |
Comput. Vis. Image Underst. | 3 |
| 2023 | GAMA: Geometric analysis based motion-aware architecture for moving object segmentation
Bangwu Xu, ZhiLei Chai, Xueliang Guo, Jianbo Shi |
Comput. Vis. Image Underst. | 3 |
| 2022 | Scene text detection by adaptive feature selection with text scale-aware loss
Wenli Luo, ZhiLei Chai, Guodong Guo |
Appl. Intell. | 3 |
| 2022 | Multi-scale feature aggregation and boundary awareness network for salient object detection
Jianzhe Wang, ZhiLei Chai, Guodong Guo |
Image Vis. Comput. | 3 |
| 2022 | Fine-Grained Image Classification With Global Information and Adaptive Compensation LossabstractFine-grained image classification differs from traditional image classification in that the former needs to divide subclasses under a basic level of categories. Previous works always focus on how to locate discriminative parts of objects, but we find that the global and background information of objects neglected by them is also valuable in some situations. This letter proposes a method to combine the global information and discriminative parts information of objects to do classification, which includes three modules: (1) Activation map based crop-erase module localizes objects while avoiding localization bias due to excessive bias of the network to learn one discriminative part. (2) Part attention module helps learning discriminative part features of objects. (3) Two-level fusion module gives consideration to the global and local information of objects and some potentially effective background information. Meanwhile, we propose an adaptive compensation loss to distinguish easily confused categories. Experiments show that our method achieves state-of-the-art performance on three open benchmarks. Shuting Miao, ZhiLei Chai, Guodong Guo |
IEEE Signal Process. Lett. | 3 |
| 2021 | A novel framework for UAV returning based on FPGA
Qunfang He, Danping Zou, ZhiLei Chai |
J. Supercomput. | 4 |
| 2020 | Crowd counting by the dual-branch scale-aware network with ranking loss constraintsabstractImage crowd counting is a challenging problem. This study proposes a new deep learning method that estimates crowd counting for the congested scene. The proposed network is composed of two major components: the first ten layers of VGG16 are used as the backbone network, and a dual‐branch (named as Branch_S and Branch_D) network is proposed to be the second part of the network. Branch_S extracts low‐level information (head blob) through a shallow fully convolutional network and Branch_D uses a deep fully convolutional network to extract high‐level context features (faces and body). Features learnt from the two different branches can handle the problem of scale variation due to perspective effects and image size differences. Features of different scales extracted from the two branches are fused to generate predicted density map. On the basis of the fact that an original graph must contain more or equal number of persons than any of its sub‐images, a ranking loss function utilising the constraint relationship inside an image is proposed. Moreover, the ranking loss is combined with Euclidean loss as the final loss function. Our approach is evaluated on three benchmark datasets, and better results are achieved compared with the state‐of‐the‐art works. Fangfang Yan, ZhiLei Chai, Guodong Guo |
IET Comput. Vis. | 3 |
| 2019 | A PYNQ-compliant Online Platform for Zynq-based DNN DevelopersabstractThe Zynq heterogeneous SoC from Xilinx is able to supporting software/hardware co-designing in one single chip, making it possible to take advantage of software flexibility and hardware acceleration at the same time. PYNQ project from Xilinx is trying to take advantage of high performance and low power consumption of Zynq while improve its programmability. In order to improve the ecosystem of PYNQ and help more embedded AI applications use the Zynq-based high-efficiency computational engine, this paper proposes a PYNQ-compliant online platform (OpenHEC-PYNQ) that integrates all necessary factors for the Zynq-based DNN developer. This platform makes HDL/HLS designers able to access all resources they needed via the Internet and finish all jobs one-stop. To show effectiveness of this platform, a YOLOv2 FPGA acceleration library is implemented based on OpenHEC-PYNQ. Jun Xia 0003, Wenmin Yang, ZhiLei Chai |
FPGA | 5 |
| 2019 | Severe Convective Weather Classification in Remote Sensing Images by Semantic Segmentation
ZhiLei Chai, Wenlai Zhao |
ICANN (3) | 2 |
| 2018 | Taking advantage of multi-regions-based diagonal texture structure descriptor for image retrieval
Wei Song 0008, Yubing Zhang, Fei Liu 0001, ZhiLei Chai, Feng Ding 0001, Xuezhong Qian, Soon Cheol Park |
Expert Syst. Appl. | 4 |
| 2017 | FingerVoice: A Syllable Based Input System Via Fingers TouchingabstractThis paper proposes a novel input system via fingers touching, called FingerVoice. Instead of traditional letter based input system, it is based on syllables. Fingers on the left hand stand for consonants while fingers on the right hand for vowels. Fingers from both hands touch each other to form a syllable. We use a pair of gloves with conductive finger caps to detect the touch. As an example, the implementation of Chinese is presented in details. By connecting it to a smart phone via Bluetooth, a utility speaking system is implemented to help the speech-impaired to "speak" via smart phone. Our experimental result shows that this syllable based input device is faster than letter based input devices. Yangyang Ma, ZhiLei Chai, Mingsong Chen 0001 |
ASSETS | 3 |
| 2017 | An FPGA-Based Real-Time Moving Object Tracking Approach
Yangyang Ma, ZhiLei Chai, Mingsong Chen 0001, Daojing He |
ICA3PP | 3 |
| 2016 | FPGA-Based Parallel Implementation of SURF AlgorithmabstractSURF (Speeded up robust features) detection is used extensively in object detection, tracking and matching. However, due to its high complexity, it is usually a challenge to perform such detection in real time on a general-purpose processor. This paper proposes a parallel computing algorithm for the fast computation of SURF, which is specially designed for FPGAs. By efficiently exploiting the advantages of the architecture of an FPGA, and by appropriately handling the inherent parallelism of the SURF computation, the proposed algorithm is able to significantly reduce the computation time. Our experimental results show that, for an image with a resolution of 640x480, the processing time for computing using SURF is only 0.047 seconds on an FPGA (XC6SLX150T, 66.7 MHz), which is 13 times faster than when performed on a typical i3-3240 CPU (with a 3.4 GHz main frequency) and 249 times faster than when performed on a traditional ARM system (CortexTM-A8, 1 GHz). Shuaishuai Ding, ZhiLei Chai, Daojing He, Qiwei Peng 0001 |
ICPADS | 3 |
| 2016 | Implementing Dense Optical Flow Computation on a Heterogeneous FPGA SoC in CabstractHigh-quality optical flow computation algorithms are computationally intensive. The low computational speed of such algorithms causes difficulties for real-world applications. In this article, we propose an optimized implementation of the classical Combine-Brightness-Gradient (CBG) model on the Xilinx ZYNQ FPGA-SoC, by taking advantage of the inherent algorithmic parallelism and ZYNQ architecture. The execution time decreases to 0.82 second with a lower power consumption (1.881W). It is better than software implementation on PC (Intel i7-3520M, 2.9GHz), which costs 2.635 seconds and 35W. We use C rather than HDLs to describe the algorithm for rapid prototyping. Jiuzhen Liang, ZhiLei Chai |
ACM Trans. Archit. Code Optim. | 5 |
| 2015 | An Embedded FPGA Operating System Optimized for Vision Computing (Abstract Only)abstractAlthough FPGA's power and performance advantages were recognized widely, designing applications on FPGA-based systems is traditionally a task undertaken by hardware experts. It is significant to allow application-level programmers with less system-level but more algorithm knowledge to realize their applications conveniently on FPGAs. In this paper, an embedded FPGA operating system is proposed to facilitate application-level programmers to use FPGAs. Firstly, it builds specific I/Os and optimizes bus interconnection among I/Os, DDR memory, user IPs etc within the FPGA for vision computing. Secondly, it manages resources of the FPGA such as I/Os, DDR memory, communication etc, frees users from low-level details. Thirdly, it schedules tasks (IPs) executed on the FPGA dynamically in runtime, which makes the FPGA multiplexed when necessary. After porting the FPGA operating system to different FPGA platforms and implementing vision algorithms based on that, it shows the FPGA operating system is able to simplify algorithm development on FPGA platforms and improve portability of user applications. Furthermore, implementation results of several popular vision algorithms show the FPGA operating system is efficient and effective for vision computing. Finally, experimental results shows that for multiple algorithms requiring more FPGA resources, runtime task scheduling of multiple IPs is more efficient than a fixed IP when the SoC of FPGA is considered. ZhiLei Chai, Haojie Zhou |
FPGA | 1 |
| 2015 | Parallel Implementation of Dense Optical Flow Computation on Many-Core Processor
Linhua Jiang, ZhiLei Chai |
ICA3PP (1) | 6 |
| 2014 | Implementing FPGA-based energy-efficient dense optical flow computation with high portability in C (abstract only)abstractOptical flow computation is widely used in many video/image based applications such as motion detection, video compression etc. Dense optical flow field that provides more details of information is more useful in lots of applications. However, high-quality algorithms for dense optical flow computation are computationally expensive. For instance, on the ARM Cortex-A9 processor within ZYNQ, the popular linear variational method Combine-Brightness-Gradient (CBG), spends $26.68s per frame to compute optical flow when the image size is 640 x 480. It is difficult to be sped up especially when embedded systems with power constraints are considered. Poor portability is another factor to limit current implementations of optical flow computation to be used in more applications. In this paper, a high-performance, low-power FPGA-accelerated implementation of dense optical flow computation is presented. One high-quality dense optical flow method, the Combine-Brightness-Gradient model, is implemented. C code instead of VHDL/Verilog HDL is used to improve the productivity. Portability of the system is designed carefully for deploying it on different platforms conveniently. Experimental results show 12 fps and 0.38J per frame are achieved by this optical flow computing system when 640 x 480 image is used and optical flow for all pixels are computed. Furthermore, portability is demonstrated by implementing the optical flow algorithm on different heterogeneous platforms such as the ZYNQ-7000 SoC and the PC-FPGA platform with a Kintex-7 FPGA respectively. Wenmin Yang, ZhiLei Chai |
FPGA | 4 |
| 2014 | Using C to implement high-efficient computation of dense optical flow on FPGA-accelerated heterogeneous platformsabstractHigh-quality algorithms for dense optical flow computation are computationally intensive. To compute them with high speed and low power is vital to make optical flow computation applicable in real-world applications. In contrast to only the Horn-Schunck model being studied on FPGA-based systems today, one of the best linear variational methods for dense optical flow computation, Combine-Brightness-Gradient, is implemented on FPGA-accelerated heterogeneous platforms in this paper. C instead of HDLs is employed and optimizing techniques based on the algorithmic parallelism and hardware architecture are introduced. Experimental results show that 30-110x improvement of the computing efficiency over CPUs was achieved. The FPGA-accelerated version is able to process 640 × 480 image at 12 fps with 0.38 J per frame, while it is 0.8 fps and around 40 J on CPUs. Through demonstrating high performance and low power of dense optical flow algorithm on FPGA-based heterogeneous platforms implemented in C, this paper shows that the off-the-shelf commodity FPGAs coupled with High-Level-Synthesis (HLS) tools could provide an available option when computational efficiency together with development speed are required. ZhiLei Chai, Haojie Zhou |
FPT | 1 |
| 2014 | Different lighting processing and feature extraction methods for efficient face recognitionabstractThis study studies different lighting processing and feature extraction methods for efficient face recognition. The purpose is to find some robust face recognition methods by combining different feature extraction methods with illumination compensation. In this study, some typical illumination preprocessing approaches are reviewed including wavelet transformation, self‐quotient image, Retinex, smoothing, discrete cosine transform normalisation in logarithm domain, homomorphic filter and local contrast enhancement. As the main contribution, this study proposes two efficient feature extraction methods for face recognition. One is an adaptive feature extraction (AFE) based on curvelet transform. The other is a feature extraction technique named two‐dimensional principal component analysis (2DPCA) non‐parametric analysis of 2D subspace (2DPCA + 2DNSA). Two groups of experiments are designed to verify the proposed methods. The first group of experimental results show that the proposed AFE methods have better performance than conventional methods. In the second group of experiments, each feature extraction method is combined with nine different lighting processing methods. The results show that the proposed 2DPCA + 2DNSA method is more robust to lighting processing than other methods. Experimental results also show that lighting processing contribution to face recognition are quite different for different face databases. Jiuzhen Liang, ZhiLei Chai |
IET Image Process. | 3 |
| 2010 | Mechanism of method invocation and return in real-time embedded Java processorabstractIn the aspect of real-time Java, RTSJ makes a series of effective work and provides the criteria for researches. In order to provide efficient execution platform for real-time Java, a RTSJ-oriented 32-bit embedded processor that can directly execute Java bytecodes-JPOR-32 was designed. Among Java bytecodes, the method invocation and return instructions are extraordinary complex. Executing these instructions is time-consuming and the Worst Case Execution Time (WCET) of them is hard to predict. This paper analyzes detailedly the method invocation and return mechanism of JPOR-32. Through the preprocessor module, JPOR-32 accomplishes the unreal-time run-time operations in advance, and achieved WCET predictability. Besides, the method invocation and return instructions are optimized in JPOR-32 and the static conversion from bytecodes to microcodes is accomplished in the image files. Combined with the instruction refetching and buffering scheme as well as the optimized design of the run-time stack structure, JPOR-32 provides effective supports for the method invocation and return of Java. Guang Hu 0006, Xindong Ye, ZhiLei Chai, Shi-liang Tu |
CSCWD | 3 |