Junbo Tie

dblp:213/9266 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0002-8989-8931ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 8 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Vector Value Prediction with Element-wise Stride Compression
abstract
The increasing emphasis on vectorization and Single Instruction, Multiple Data (SIMD) processing reflects their central role in modern processors. However, as workloads in data processing, multimedia, and algorithmic operations grow in complexity, they introduce more pronounced data dependencies, leading to longer execution times compared to scalar instructions. To address these evolving challenges, we present the Vector Value TAGE predictor (VVTAGE), a novel value predictor specifically designed for vector instructions. Although value prediction has been proposed as a fundamental strategy to enhance processor performance, it has traditionally focused on predicting 64-bit scalar values to mitigate data dependencies and improve pipeline throughput. VVTAGE extends the prediction capabilities of existing scalar predictors to accommodate the wide vector registers used in contemporary processors. Our research demonstrates that VVTAGE can significantly improve processor performance, achieving performance gains of up to 20.1% and an average increase of 4.53% in the evaluated SIMD benchmarks. This innovative approach and surprising results represent a significant advancement in optimizing the performance of SIMD processors. Furthermore, to enhance the scalability of VVTAGE, we propose an element-wise stride compression method to reduce its storage overhead. Experimental results show that VVTAGE still retains 64% performance gain while reducing 15.4KB overhead.
Yanmeng Huang, Ling Yang 0008, Yuanhu Cheng, Quan Deng 0003, Junbo Tie, Yongwen Wang, Hai Zhong, Libo Huang 0002
ACM Great Lakes Symposium on VLSI6
2025 RISC-TAE: Instruction Set Extension for Transformer Model Acceleration
abstract
This paper proposes RISC-TAE(Transformer Acceleration Engine), a RISC-V instruction set extension with microarchitectural co-design, to meet the performance and energy-efficiency requirements of Transformer models in edge scenes.The design integrates operator-specialized computing units (GEMM/Softmax/GELU) with hardware-managed dataflow orchestration through custom RISC-V instructions, effectively resolving energy efficiency bottlenecks and memory access fragmentation in existing solutions. Experiments demonstrate that RISC-TAE achieves 23.35× and 8.01× speedups over scalar processor CVA6 and vector processor ARA respectively for BERT inference, while outperforming RISC-VTF by 1.4×, providing a scalable solution for edge Transformer deployment.
Yanping Shao, Zhouquan Liu, Junbo Tie, Gang Chen 0023, Libo Huang 0002
CASES5
2025 X-SA: An Efficient Configurable Systolic Array Computing Architecture for GPGPU
abstract
GPGPUs are pivotal for edge AI, but resource constraints demand efficient low-precision computation. Conventional GPGPUs face challenges in resource utilization, particularly with irregular matrices common in AI, and memory bandwidth limitations on edge devices. Traditional fixed-size systolic arrays often suffer from underutilization under varying workloads. This paper introduces X-SA, a configurable systolic array architecture tailored for INT8 matrix multiplication on GPGPUs in resource-constrained edge environments. X-SA distinctively employs a parameterized$2 \times N$processing element design enabling dynamic computational scaling, unlike fixed systolic arrays. It integrates an interleaved matrix buffer to alleviate memory bottlenecks and optimize dataflow. Experimental results demonstrate X-SA achieves a$2.83 \times$performance speedup over the Vortex baseline with minimal Look-Up Table overhead of 2.8% and Flip-Flops overhead of 1.4%.. It offers comparable performance to a standard$4 \times 4$systolic array but with significantly reduced area by 46.26% and power by 39.42%, and superior processing element utilization for irregular matrices. X-SA provides a approach to help improve the performance of some AI applications running on edge GPGPUs relatively in resource-constrained environments.
Yingsong Wang, Zhenzhen Jia, Ling Yang 0008, Hongbing Tan, Junsheng Chang, Junbo Tie, Libo Huang 0002
HPCC6
2025 SICA: A Multicore Neuromorphic Processor Featuring Sparse Integration and Communication-Aware Optimization
abstract
Neuromorphic computing has emerged as a promising paradigm due to its event-driven operation and energy efficiency, driving extensive research in neuromorphic processor development. When implementing spiking neural networks (SNNs) on such processors, two critical aspects must be addressed: neuron computation and spike communication. For neuron computation, previous work primarily relies on parallel accumulation via adder trees but fails to leverage the inherent sparsity in SNNs. For spike communication, conventional mesh topologies suffer from long-distance communication inefficiencies, while suboptimal mapping strategies further exacerbate latency issues. To address these challenges, we propose a low-overhead fast sparse detection mechanism that effectively exploits spike sparsity and optimizes the processor's workflow, thereby achieving efficient synaptic integration with minimal overhead. For spike communication, we employ an on-chip broadcast mechanism combined with a hybrid torus-mesh topology to significantly reduce communication latency, while systematically evaluating the impact of three distinct mapping strategies-random, sequential, and communication-aware mapping-on overall performance. Experimental results demonstrate significant improvements, with our solution delivering speedups of$1.26 \times$and$1.24 \times$compared to LSMCore on the N-MNIST and MNIST datasets, respectively. Furthermore, the communication-aware mapping strategy achieves a 24.87% reduction in communication latency, while the torus topology contributes an additional$\mathbf{1 4. 4 8} \boldsymbol{\%}$latency reduction.
Junbo Tie, Xun Xiao, Yuanfeng Luo, Yang Guo 0003, Lei Wang 0011
HPCC3
2025 STARTS: Simulation Traits Assisted Random Test Selection for Multiprocessor Verification
abstract
Test selection is vital for accelerating multiprocessor design verification. Current methods focus on the similarities among random test cases but overlook critical runtime characteristics that significantly impact outcomes. This limitation hinders their ability to capture patterns generated at runtime in multiprocessor systems. To address this, we propose Simulation Traits Assisted Random Test Selection (STARTS), which utilizes features obtained from fast model simulation of the target multiprocessor during verification. STARTS rapidly simulates test cases to predict hardware behavior and employs a GRU-based variational autoencoder with a self-attention mechanism to embed test cases, capturing both data and control dependencies between stimulus actions. The distance in latent space serves as a key selection criterion. By integrating hardware-related information with intrinsic multiprocessor characteristics, our method enhances test selection effectiveness. Experimental results show that STARTS reduces the time needed to achieve coverage goals compared to other unsupervised learning methods, demonstrating its superior performance in selecting optimal stimuli.
Li Zhou 0009, Menglong Lu, Junbo Tie
ITC5
2024 A Fast and Safe Neuromorphic Approach for Obstacle Avoidance of Unmanned Aerial Vehicle
abstract
Obstacle avoidance is a crucial task in unmanned aerial vehicles (UAV) motion planning. The accuracy and consistency of real-time visual information affect the gener-ation of obstacle avoidance commands, raising higher safety demands for obstacle avoidance. The neuromorphic computing-based obstacle avoidance solution can address these challenges. Dynamic vision sensors (DVS) exhibit low latency, low power consumption, and high dynamic range as novel neuromorphic sensors. Spiking neural networks (SNN) also leverage the same mechanism to efficiently process asynchronous and sparse event data generated by DVS, offering latency and energy efficiency advantages. Additionally, the optimal estimation method effectively mitigates the impact of noise and interference within the system, reducing the influence of errors on the algorithm and enhancing safety. Based on these considerations, this paper proposes a fast and safe obstacle avoidance framework. DVS is used to acquire event data from the environment, and a hardware-compatible lightweight SNN is employed to extract dynamic obstacle position information from the data. Compared to baseline methods, this approach reduces latency by 85%. Furthermore, two estimation methods are used to predict the movement of obstacles, ensuring flight safety by generating different UAV obstacle avoidance actions based on confidence intervals, even in the presence of obstacle information errors and omissions.
Zhong Wan, Xun Xiao, Jingyue Zhao, Junbo Tie, Renzhi Chen, Guangda Zhang, Huadong Dai
SMC5
2024 A security JPEG image system accelerated by NEON technology based on FT-2000/4
Junbo Tie, Lei Wang 0011
CCF Trans. High Perform. Comput.4
2024 Hierarchical Mapping of Large-Scale Spiking Convolutional Neural Networks Onto Resource-Constrained Neuromorphic Processor
abstract
Neuromorphic processors have been designed as non-von Neumann systems for energy-efficient spiking neural network (SNN) execution. Spiking convolutional neural networks (SCNNs), combining the advantage of convolutional neural network (CNN) and SNN, have been widely applied to vision tasks. However, as the scale of SCNNs increases, executing large-scale SCNNs on resource-constrained neuromorphic processor faces many challenges, including massive synapse pruning caused by resource competition, execution performance degradation, etc. Addressing these problems, we propose an efficient approach to map large-scale SCNNs onto resource-constrined neuromorphic processor. The approach consists of three steps: splitting, partitioning, and mapping. We explore three acyclic splitting strategies to divide large-scale SCNNs into subnetworks without cyclic dependency. Axon sharing is the guiding principle to partition subnetworks into multiple clusters. To obtain an optimal cluster-to-core mapping scheme, we use Non-dominated Sorting Genetic Algorithm to collaboratively optimize two metrics. We evaluate our approach with eight realistic SCNN applications. The results show that compared with existing state-of-the-art methods, our approach significantly reduces the synapse pruning and accuracy loss, and increases the execution performance.
Xun Xiao, Yao Wang 0002, Junbo Tie, Lei Wang 0011, Weixia Xu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2023 F-E Fusion: A Fast Detection Method of Moving UAV Based on Frame and Event Flow
Xun Xiao, Zhong Wan, Shasha Guo 0001, Junbo Tie
ICANN (8)5
2023 Dynamic Obstacle Avoidance for Unmanned Aerial Vehicle Using Dynamic Vision Sensor
Junbo Tie, Jingyue Zhao, Zhong Wan, Guangda Zhang, Lei Wang 0011
ICANN (10)2
2023 Back to Homogeneous Computing: A Tightly-Coupled Neuromorphic Processor With Neuromorphic ISA
abstract
In recent years, neuromorphic processors are widely used in many scenarios, showing extreme energy efficiency over traditional architectures. However, almost all existing neuromorphic hardware are following the heterogeneous computing methodology without Instruction Set Architecture (ISA), leading to inflexibility in programming. In this paper, we first propose a RISC-V Neuromorphic Extension (RVNE) to enable fine-grained and flexible homogeneous programming for neuromorphic algorithms while utilizing SNN sparsity from different levels of granularity and computing flows. Based on RVNE, we next implement a neuromorphic micro-architecture that is tightly coupled to the CPU pipeline to accelerate neuromorphic computing. To demonstrate the proposed homogeneous neuromorphic architecture, we implement a prototype processor called NeuroRVcore based on RISC-V ISA and an open-source RISC-V core. The evaluation results show that RVNE achieves a 2.8 × −4.3 × reduction in code density compared with the general-purpose ISAs. Compared with the state-of-the-art neuromorphic processor, the proposed homogeneous computing reduces energy consumption by 3.4%−22.5% while enabling fine-grained and flexible homogeneous programming.
Lei Wang 0011, Yao Wang 0002, Junbo Tie, Feng Wang 0050, LingHui Peng, Xun Xiao, Gan Zhou, Xuhu Yu, Xia Zhao 0004, Yuhua Tang, Weixia Xu 0001
IEEE Trans. Parallel Distributed Syst.5
2022 Multimodal Learning of Audio-Visual Speech Recognition with Liquid State Machine
Xuhu Yu, Changhao Chen, Junbo Tie, Shasha Guo 0001
ICONIP (6)4