Jiajun Luo

dblp:212/7534 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SHMemora: Protective Key-Value Store on Distributed Shared Memory
Jiajun Luo, Yunpeng Xu, Shengwei Liu, Jin Xia, Huanchen Zhang, Shuwen Deng
ICDE1
2026 A Sub-50-nW Resistor-Less CMOS Current Reference Utilizing High-PSRR Voltage Biasing
Daokang Liu, Wenxin Zhao, Hainan Liu, Jiajun Luo
ISCAS4
2026 MergFS: Efficient Bridging of a 32-bit High-Speed Intra-Core Bus to a 64-bit Low-Speed AHB-Lite Bus
abstract
In the architecture of system-on-chip (SoC) design, the bus plays a critical role by facilitating inter-module connections and managing data transmission. Although the commercial bus solutions represented by the Cortex-M System Design Kit (CMSDK) are widely adopted in the industry, they exhibit lower communication efficiency in certain specialized requirements. This research focuses on optimizing the transition from a 32-bit high-speed intra-core bus to a 64-bit low-speed AHB-Lite bus. A new solution (MergFS) for this conversion process is proposed and implemented, which supports request merging and dynamic frequency switching. Experimental results show that, compared to the solution using CMSDK, MergFS reduces clock cycles by approximately 50% to 75% when processing multiple transactions. Additionally, synthesis results under the three different technologies show MergFS introduces ~4.5% area and ~6.02% power overhead on average.
Zewen Cao, Zhuo Peng, Yuying Dong, Hongrui Ruan, Chuanbin Zeng, Hualong Zhao, Jiajun Luo
IEEE Trans. Circuits Syst. I Regul. Pap.8
2025 CXL-INTERPLAY: Unraveling and Characterizing CXL Interference in Modern Computer Systems
abstract
Compute Express Link (CXL) is a promising technology that addresses memory and storage challenges. Despite its advantages, CXL faces performance threats from external interference when coexisting with current memory and storage systems. This interference is under-explored in existing research. To address this, we develop CXL-Interplay, systematically characterizing and analyzing interference from memory and storage systems. To the best of our knowledge, we are the first to characterize CXL interference on real CXL hardware. We also provide reverse-reasoning analysis with performance counters and kernel functions. In the end, we propose and evaluate mitigation solutions.
Shunyu Mao, Jiajun Luo, Jiapeng Zhou, Zheng Liu 0022, Teng Ma 0006, Shuwen Deng
DAC2
2025 Beyond A Single AI Cluster: A Survey of Decentralized LLM Training
abstract
The emergence of large language models (LLMs) has revolutionized AI development, yet their resource demands beyond a single cluster or even datacenter, limiting accessibility to well-resourced organizations.Decentralized training has emerged as a promising paradigm to leverage dispersed resources across clusters, datacenters and even regions, offering the potential to democratize LLM development for broader communities.As the first comprehensive exploration of this emerging field, we present decentralized LLM training as a resource-driven paradigm and categorize existing efforts into community-driven and organizational approaches.We further clarify this through: (1) a comparison with related paradigms, (2) characterization of decentralized resources, and (3) a taxonomy of recent advancements.We also provide up-to-date case studies and outline future directions to advance research in decentralized LLM training.
Haotian Dong, Jingyan Jiang, Rongwei Lu, Jiajun Luo, Jiajun Song, Zhi Wang 0001
EMNLP4
2025 DICE: Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
Jiajun Luo, Lizhuo Luo, Jianru Xu, Jiajun Song, Rongwei Lu, Zhi Wang 0001
ICCV1
2025 Accelerating Parallel Diffusion Model Serving with Residual Compression
abstract
Diffusion models produce realistic images and videos but require substantial computational resources, necessitating multi-accelerator parallelism for real-time deployment. However, parallel inference introduces significant communication overhead from exchanging large activations between devices, limiting efficiency and scalability. We present CompactFusion, a compression framework that significantly reduces communication while preserving generation quality. Our key observation is that diffusion activations exhibit strong temporal redundancy—adjacent steps produce highly similar activations, saturating bandwidth with near-duplicate data carrying little new information. To address this inefficiency, we seek a more compact representation that encodes only the essential information. CompactFusion achieves this via Residual Compression that transmits only compressed residuals (step-wise activation differences). Based on empirical analysis and theoretical justification, we show that it effectively removes redundant data, enabling substantial data reduction while maintaining high fidelity. We also integrate lightweight error feedback to prevent error accumulation. CompactFusion establishes a new paradigm for parallel diffusion inference, delivering lower latency and significantly higher generation quality than prior methods. On 4$\times$L20, it achieves $3.0\times$ speedup while greatly improving fidelity. It also uniquely supports communication-heavy strategies like sequence parallelism on slow networks, achieving $6.7\times$ speedup over prior overlap-based method. CompactFusion applies broadly across diffusion models and parallel settings, and integrates easily without requiring pipeline rework. Portable implementation demonstrated on xDiT is publicly available at https://github.com/Cobalt-27/CompactFusion
Jiajun Luo, Yicheng Xiao, Jianru Xu, Yangxiu You, Rongwei Lu, Jingyan Jiang, Zhi Wang 0001
NeurIPS1
2025 RIVL: A Low-Cost SoC Agile Development Platform for Multiple RISC-V Processors Design and Verification
abstract
Current processor chip designs are mainly oriented by performance, power and area (PPA), and developed using the waterfall model. However, there are two main challenges in this development model: 1) The end-to-end iteration cycle and cost of processor chip development are too high, and cannot flexibly respond to changes in chip fragmented design specifications. 2) Processor chip verification is less agile, and there is a lack of a full-chain processor agile design platform that can be easily ported to different development environments. To tackle both issues, we propose an object-oriented hardware agile design methodology, oriented by time, cost, and complexity, and have built the RIVL platform to support the agile development process for processors. RIVL integrates a highly automated design flow for processor RTL design, Integration, Verification, and Layout design to improve processor development efficiency. We achieved tape-out verification of more than 60 RISC-V processors through agile design methods, demonstrating the use and effectiveness of RIVL. We quantify the performance of CoreGen using CoreMark and demonstrate that CoreGen achieves industry-competitive performance.
Zewen Cao, Hualong Zhao, Zhuo Peng, Yuchi Miao, Chunan Zhuang, Hongrui Ruan, Yuying Dong, Chuanbin Zeng, Bo Li 0051, Jiajun Luo
IEEE Trans. Circuits Syst. I Regul. Pap.11
2024 A Joint Approach to Local Updating and Gradient Compression for Efficient Asynchronous Federated Learning
Jiajun Song, Jiajun Luo, Rongwei Lu, Shuzhao Xie, Bin Chen 0011, Zhi Wang 0001
Euro-Par (3)2
2023 FedLP: Layer-Wise Pruning Mechanism for Communication-Computation Efficient Federated Learning
abstract
Federated learning (FL) has prevailed as an efficient and privacy-preserved scheme for distributed learning. In this work, we mainly focus on the optimization of computation and communication in FL from a view of pruning. By adopting layer-wise pruning in local training and federated updating, we formulate an explicit FL pruning framework, FedLP (Federated Layer-wise Pruning), which is model-agnostic and universal for different types of deep learning models. Two specific schemes of FedLP are designed for scenarios with homogeneous local models and heterogeneous ones. Both theoretical and experimental evaluations are developed to verify that FedLP relieves the system bottlenecks of communication and computation with marginal performance decay. To the best of our knowledge, FedLP is the first framework that formally introduces the layer-wise pruning into FL. Within the scope of federated learning, more variants and combinations can be further designed based on FedLP.
Zheqi Zhu, Jiajun Luo, Fei Wang 0004, Chenghui Peng, Pingyi Fan, Khaled Ben Letaief
ICC3