Junyi Luo

dblp:199/4697 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 FLARE: Finetuning ReLU With FIRE for Efficient Long-Context Inference
abstract
Deploying large language models (LLMs) on resource-constrained edge devices, such as mobile phones or IoT devices, is highly desirable for enabling secure, personalized on-device AI. However, there are significant challenges due to these models’ high computational and memory demands. A key bottleneck lies in the Transformer’s attention block, especially when handling long contexts. Techniques like model architectures with Rectified Linear Unit (ReLU) activations for Softmax and FIRE positional encoding (a resource-efficient, automatic context-length-scaling alternative to Rotary Positional Embedding (RoPE)) have each independently shown promise in reducing the computational complexity of the attention block, but the proper alchemy for combining their benefits remains underexplored. In this paper, we show a method for combining FIRE and ReLU that maintains low-validation loss at long contexts. We also introduce FLARE, a new algorithm that further improves efficiency by removing operations from the learned relative position encoding in FIRE. Our approach leads to faster inference on long sequences, robust generalization to varying context lengths, and lower validation loss compared to baseline models. FLARE achieves a significant reduction in power and area consumption. On custom hardware, it achieves a 6× higher operating frequency than Softmax, while occupying 57× less silicon area (measured under different throughput settings) and consuming 600× less energy. Our results indicate that FLARE represents a significant step towards deploying powerful LLMs efficiently on resource-limited devices. Code and hardware designs are publicly available at: https://github.com/ReaLLMASIC/nanoGPT.
Michael Moffatt, Junyi Luo, Xinting Jiang, Guanchen Tao, Shiwei Liu 0002, Kauna Lei, Gregory Kielian, Mehdi Saligane
DATE2
2026 From Small to Large: A Heuristic Divide-and-Neural-Conquer Framework for Large-Scale Vehicle Routing Problems
Debing Wang, Junyi Luo, Zhanhong Fang, Yunfeng Xu, Zizhen Zhang
PPSN (1)2
2023 An Integer-Only and Group-Vector Systolic Accelerator for Efficiently Mapping Vision Transformer on Edge
abstract
Transformer-like network has shown remarkable high performance in both natural language processing and computer vision. However, the huge computational demands in non-linear floating-point arithmetic and the irregular memory access requirement in self-attention mechanism make it still a challenge to deploy Transformer on edge. To address the above issues, we propose integer-only quantization scheme for the simplification of non-linear operations (such as LayerNorm, Softmax and Gelu), meanwhile algorithm-hardware co-design strategy is applied to guarantee both the high accuracy and high efficiency. Besides, we construct general-purpose group vector systolic array to efficiently accelerate the matrix multiplication operations including both regular matrix-multiplication/convolution and the irregular multi-head self-attention mechanism. Unified data-package strategy and flexible on-/off-chip data storage management strategy are also proposed to further improve the performance. The design has been deployed on Xilinx ZCU102 FPGA platform, achieving an overall inference latency of 4.077ms and 11.15ms per image for ViT-tiny and ViT-s, respectively. The average throughput can reach as high as 762.7 GOPs, which shows significant improvement over the previous state-of-the-art FPGA Transformer accelerator.
Mingqiang Huang, Junyi Luo, Chenchen Ding, Zikun Wei, Sixiao Huang, Hao Yu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 A Precision-Scalable Energy-Efficient Bit-Split-and-Combination Vector Systolic Accelerator for NAS-Optimized DNNs on Edge
abstract
Optimized model and energy-efficient hardware are both required for deep neural networks (DNNs) in edge-computing area. Neural architecture search (NAS) methods are employed for DNN model optimization with resulted multi-precision networks. Previous works have proposed low-precision-combination (LPC) and high-precision-split (HPS) methods for multi-precision networks, which are not energy-efficient for precision-scalable vector implementation. In this paper, a bit-split-and-combination (BSC) based vector systolic accelerator is developed for a precision-scalable energy-efficient convolution on edge. The maximum energy efficiency of the proposed BSC vector processing element (PE) is up to 1.95× higher in 2-bit, 4-bit and 8-bit operations when compared with LPC and HPS PEs. Further with NAS optimized multi-precision CNN networks, the averaged energy efficiency of the proposed vector systolic BSC PE array achieves up to 2.18× higher in 2-bit, 4-bit and 8-bit operations than that of LPC and HPS PE arrays.
Kai Li 0024, Junzhuo Zhou, Junyi Luo, Zhengke Yang, Shuxin Yang, Wei Mao 0002, Mingqiang Huang, Hao Yu 0001
DATE4
2022 A High Throughput Multi-bit-width 3D Systolic Accelerator for NAS Optimized Deep Neural Networks on FPGA
abstract
Neural architecture search (NAS) optimized multi-bit-width convolutional neural network (CNN) maintains the balance between network performance and efficiency, thus enlightening a promising method for accurate yet energy-efficient edge computing. In this work, we propose a high throughput three-dimensional (3D) systolic accelerator for NAS optimized CNNs, in which the input feature matrix, weight matrix and output feature matrix are delivering vertically, horizontally and perpendicularly through the systolic array respectively. With 3D systolic data flow, the processing time and logic resources consumption can be both reduced compared to the classical non-stationary systolic array. Besides, Booth-based multi-bit-width (INT2/4/8) multiply-add-accumulation (MAC) unit is developed within the 3D systolic accelerator. Deployed on FPGA platform Xilinx ZCU102, peek performance of the convolutional layer can reach as high as 2775 GOPS for INT2, 1650 GOPS for INT4, and 816 GOPS for INT8 respectively. The average performance on accelerating full NAS VGG16 network is 647 GOPS.
Mingqiang Huang, Yucen Liu, Shuxin Yang, Kai Li 0024, Junyi Luo, Zhengke Yang, Qiufeng Li, Hao Yu 0001, Changhai Man
FPGA6
2020 MOOCCube: A Large-scale Data Repository for NLP Applications in MOOCs
abstract
Jifan Yu, Gan Luo, Tong Xiao, Qingyang Zhong, Yuquan Wang, Wenzheng Feng, Junyi Luo, Chenyu Wang, Lei Hou, Juanzi Li, Zhiyuan Liu, Jie Tang. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Jifan Yu, Gan Luo, Tong Xiao 0002, Qingyang Zhong, Yuquan Wang, Wenzheng Feng, Junyi Luo, Lei Hou 0001, Juan-Zi Li, Zhiyuan Liu 0001, Jie Tang 0001
ACL7