Zikang Zhou

dblp:68/7727 · DBLP profile ↗
← Back
15ranked-venue papers
9as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 DSLA: An Energy-Efficient Dual-Sparsity LLM Accelerator With HiMix-BFP
abstract
Large language models (LLMs) have achieved great success in areas such as language understanding and text generation. However, their massive parameters incur high computational, storage, and energy costs, making deployment on resource- and power-constrained edge devices particularly challenging. Block Floating Point (BFP) reduces storage and computational overhead by grouping data into blocks and aligning them to the maximum exponent within each block, converting them into low-bit fixed-point numbers. Bidirectional Block Floating Point (BBFP) extends BFP by aligning data within a block to two different exponents, reducing quantization error for small values. However, at ultra-low bit widths, both methods suffer from severe accuracy degradation due to their sensitivity to outliers, which limits their ability to achieve aggressive energy savings. In addition, the exponent alignment procedure in these schemes inherently introduces bit-sliced sparsity, an opportunity that remains largely unexplored for further improving energy efficiency on edge accelerators. To address these challenges, we propose an energy-efficient dual-sparsity LLM accelerator (DSLA) that supports the HiMix-BFP data format. HiMix-BFP improves low-bit accuracy by preserving extra mantissa bits for the maximum value and adaptively selecting BFP or BBFP per block. The DSLA architecture efficiently exploits both value and bit-sliced sparsity across different bit widths to further enhance energy efficiency by employing low-bit computational units, load-balancing mechanisms, and a hierarchical bit-accumulation array. Experimental results demonstrate that HiMix-BFP reduces perplexity by up to 76%, while DSLA achieves up to$1.39\times $higher throughput and$1.59\times $greater energy efficiency compared to SOTA accelerators.
Zikang Zhou, Siyao Dai, Jun Han 0003
IEEE Trans. Circuits Syst. I Regul. Pap.1
2026 Area- and Utilization-Efficient LLM Accelerator With Fused Speculative Decoding for Edge-Side Inference
abstract
The inference of large language models (LLMs) on edge devices has always been a challenge, the autoregressive inference and enormous number of parameters result in long inference latency. Although speculative decoding is proposed for inference acceleration, it poses challenges for hardware accelerator design, as the differing computational characteristics of the drafting and verification stages make it difficult to optimize both chip area and hardware utilization. This brief introduces fused speculative decoding (FSD) to optimize LLM inference on edge devices by unifying all the operations to general matrix multiplications during inference. The proposed FSD-Infer algorithm fuses the drafting and verification phases of conventional speculative decoding and enables weight sharing, reducing off-chip memory access, and inference latency without requiring retraining or fine-tuning. For hardware, we introduce FSD-Acc, an area- and utilization-efficient hardware accelerator that efficiently executes the fused operations enabled by FSD-Infer. Experimental results show that compared with autoregressive inference, FSD-Infer reduces EMA by up to 26.11%, enhances arithmetic intensity by$10.26\times $, and speeds up GPU inference by up to$1.38\times $. When deployed on Xilinx ZCU102 FPGA board at 200 MHz working frequency, FSD-Acc achieves the best area efficiency and energy efficiency, outperforming the state-of-the-art LLM accelerator for edge inference by$2.53\times $(TPS/kDSP),$5.90\times $(TPS/kLUT) and$2.02\times $, respectively.
Zikang Zhou, Jun Han 0003
IEEE Trans. Very Large Scale Integr. Syst.2
2026 RAEnc: A Stall-Less Metadata Compression Framework for Return Address Integrity on High-Performance Embedded Processors
abstract
Protecting return address integrity (RAI) on high-performance embedded processors (HPEPs) is more challenging than on traditional microcontroller units (MCUs) or server processors. Isolation-based schemes such as shadow stacks fail against hardware memory threats, while crypto-based schemes remain vulnerable to forgery and replay attacks. To address the challenges of RAI protection in HPEPs, this article presentsRAEnc, a novel crypto-based scheme, which provides comprehensive RAI against both forgery and replay attacks with negligible performance overhead via hardware–software co-design. We first introduce the metadata compression during return address encryption (MCRAE), a cryptographic primitive compressing the call stack’s integrity state into a single, register-resident value using custom RISC-V instructions, thereby eliminating the memory attack surface. We then detail the stall-less encryption/decryption unit (EDU) that is tightly coupled with the processor pipeline. By employing speculative scheduling, a low-latency 128-bit QARMA unit, encryption/decryption cache (EDCache), and multichain parallelism, the EDU eliminates pipeline stalls common in decoupled accelerators. Implemented on the open-source BOOM processor,RAEncincurs a negligible 0.2% performance overhead on embedded and SPEC CPU2017 workloads with modest hardware cost, outperforming existing crypto-based RAI schemes in both security and performance.
Zikang Zhou, Kanheng Jiang, Yifan Zhao 0007, Jun Han 0003
IEEE Trans. Very Large Scale Integr. Syst.2
2025 ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling
abstract
Anticipating the multimodality of future events lays the foundation for safe autonomous driving. However, multimodal motion prediction for traffic agents has been clouded by the lack of multimodal ground truth. Existing works predominantly adopt the winner-take-all training strategy to tackle this challenge, yet still suffer from limited trajectory diversity and uncalibrated mode confidence. While some approaches address these limitations by generating excessive trajectory candidates, they necessitate a postprocessing stage to identify the most representative modes, a process lacking universal principles and compromising trajectory accuracy. We are thus motivated to introduce ModeSeq, a new multimodal prediction paradigm that models modes as sequences. Unlike the common practice of decoding multiple plausible trajectories in one shot, ModeSeq requires motion decoders to infer the next mode step by step, thereby more explicitly capturing the correlation between modes and significantly enhancing the ability to reason about multimodality. Leveraging the inductive bias of sequential mode prediction, we also propose the EarlyMatch-Take-All (EMTA) training strategy to diversify the trajectories further. Without relying on dense mode prediction or heuristic post-processing, ModeSeq considerably improves the diversity of multimodal output while attaining satisfactory trajectory accuracy, resulting in balanced performance on motion prediction benchmarks. Moreover, ModeSeq naturally emerges with the capability of mode extrapolation, which supports forecasting more behavior modes when the future is highly uncertain.
Zikang Zhou, Hengjian Zhou, Jianping Wang 0001, Yung-Hui Li, Yu-Kai Huang 0001
CVPR1
2025 A Memory-Efficient LLM Accelerator with Q-K Correlation Prediction using Cluster-Based Associative Array for Selective KV Accessing
abstract
Attention-based LLMs excel in text generation but face redundant computations in autoregressive token generation. While KV cache mitigates this, it introduces increased memory access overhead as sequences grow. We propose Sella, a hardware-software co-design using cluster-based associative arrays to predict Q-K correlations, enabling selective KV cache access and reducing memory access without retraining. Sella includes a specialized accelerator featuring a prediction engine to improve performance and energy efficiency. Experiments show Sella achieves $2.1 \times$, $93.8 \times$, $31.4 \times$, and $53.5 \times$ speedup over SpAtten, Sanger, TITAN RTX GPU, and Xeon CPU, respectively, reducing off-chip memory access by up to $66 \%$ with negligible accuracy loss. -Large Language Models, KV Cache, Accelerator
Zikang Zhou, Xuyang Duan, Jun Han 0003
DAC1
2025 Multi-DOF Fusion: A Flexible Fusion Strategy for Reducing Redundancy in CNN Workloads
Zikang Zhou, Siyao Dai, Xuyang Duan, Jun Han 0003
ACM Great Lakes Symposium on VLSI2
2025 VLSUMaP: A High-Performance Matrix Processor with Virtually Expanded LSU Boosting HBM Bandwidth Utilization
Xinjie Kong, Zikang Zhou, Zhuoyuan Yang, Zengshi Wang, Jun Han 0003
ACM Great Lakes Symposium on VLSI5
2025 High-Fidelity Face Swapping via Fine-grained Attribute Control with Diffusion Models
abstract
With the emergence and development of Generative Adversarial Networks (GANs) and diffusion models, facial swapping methods have seen significant advancements, particularly excelling in generating faces that maintain consistent identities. However, existing approaches often struggle to preserve key attributes of the target face under conditions such as large pose variations, differing lighting, and occlusions. To address these challenges, we propose FACSwap, a fine-grained attribute control framework based on diffusion models, designed to enhance facial swapping. Firstly, we introduce the Attribute-Preserving Attention Module (APAM), which leverages attention mechanisms for identity disentanglement and adversarial learning to extract fine-grained attribute features. FACSwap also incorporates a 3D landmark projector operation that considers characteristics from both the source and target faces, effectively preserving subtle facial attributes. Additionally, we introduce the Compound Augmentation Identity Module (CAIM) to further enhance identity similarity. Extensive quantitative and qualitative experiments conducted on the FFHQ and CelebA datasets demonstrate that FACSwap can generate high-fidelity swapped faces with superior facial attribute preservation and identity consistency, outperforming benchmarks established by traditional methods.
Zikang Zhou, Keze Wang
IJCNN1
2025 RALAD: Bridging the Real-to-Sim Domain Gap in Autonomous Driving with Retrieval-Augmented Learning
abstract
As end-to-end autonomous driving advances toward real-world deployment, ensuring the safety of autonomous vehicles (AVs) has become a critical requirement for their commercial viability. While rule-based AVs have traditionally undergone rigorous testing in both real-world and simulated environments before deployment, data-driven autonomous models are typically trained on real-world datasets, limiting their generalization to simulation environments. This poses a significant challenge for the development and testing of end-to-end autonomous driving. To address this issue, we propose Retrieval-Augmented Learning for Autonomous Driving (RALAD), a novel framework designed to bridge the real-to-sim gap in a cost-effective manner. RALAD consists of three key components: (1) domain adaptation via an enhanced Optimal Transport (OT) method, which retrieves the most similar scenarios between real and simulated environments; (2) feature fusion across similar scenarios, enabling the construction of a feature mapping between real-world and simulated domains; and (3) feature extraction freezing with fine-tuning on the fused features, allowing the model to learn simulation-specific characteristics through feature mapping. We evaluate RALAD on three monocular 3D object detection models, and the results demonstrate that our approach significantly improves model accuracy in simulation. Additionally, we use real autonomous vehicle for testing in real-world scenarios, and have established simulated scenes similar to reality for further testing, which illustrate the effectiveness of our method.
Jiacheng Zuo, Zikang Zhou, Yufei Cui, Ziquan Liu, Jianping Wang 0001, Nan Guan, Jin Wang 0009, Chun Jason Xue
IROS3
2024 ML-Fusion: Determining Memory Levels for Data Reuse Between DNN Layers
abstract
With the increasing complexity of applications and the improvement of computational power, modern neural networks (DNNs) have become more memory-intensive. To address the bandwidth problem, modern hardware architectures often incorporate multi-level memory to efficiently reuse data with different reuse distances. Additionally, a promising technique to reduce DNN bandwidth requirements is layer fusion which reuses inter-layer data at on-chip memory. However, previous studies on inter-layer data reuse scheduling have focused primarily on reusing at the outermost on-chip memory, neglecting the exploration of multi-level architecture, which represents a significant optimization space.
Zikang Zhou, Xuyang Duan, Jun Han 0003
ACM Great Lakes Symposium on VLSI1
2024 BehaviorGPT: Smart Agent Simulation for Autonomous Driving with Next-Patch Prediction
abstract
Simulating realistic behaviors of traffic agents is pivotal for efficiently validating the safety of autonomous driving systems. Existing data-driven simulators primarily use an encoder-decoder architecture to encode the historical trajectories before decoding the future. However, the heterogeneity between encoders and decoders complicates the models, and the manual separation of historical and future trajectories leads to low data utilization. Given these limitations, we propose BehaviorGPT, a homogeneous and fully autoregressive Transformer designed to simulate the sequential behavior of multiple agents. Crucially, our approach discards the traditional separation between "history" and "future" by modeling each time step as the "current" one for motion generation, leading to a simpler, more parameter- and data-efficient agent simulator. We further introduce the Next-Patch Prediction Paradigm (NP3) to mitigate the negative effects of autoregressive modeling, in which models are trained to reason at the patch level of trajectories and capture long-range spatial-temporal interactions. Despite having merely 3M model parameters, BehaviorGPT won first place in the 2024 Waymo Open Sim Agents Challenge with a realism score of 0.7473 and a minADE score of 1.4147, demonstrating its exceptional performance in traffic agent simulation.
Zikang Zhou, Xinhong Chen 0003, Jianping Wang 0001, Nan Guan, Kui Wu 0001, Yung-Hui Li, Yu-Kai Huang 0001, Chun Jason Xue
NeurIPS1
2024 A Design Framework for Generating Energy-Efficient Accelerator on FPGA Toward Low-Level Vision
abstract
Low-level vision algorithms play an increasingly crucial role in a wide range of applications, such as biomedical, security, and autopilot. The low-level vision accelerators have also been extensively researched. As low-level vision is often deployed in embedded devices, its accelerators need to achieve high energy efficiency. Meanwhile, the broad application scenarios of low-level vision contribute to its rapid iteration. Designing energy-efficient accelerators for quickly evolving low-level vision algorithms demands substantial effort. Therefore, a design framework specifically tailored for the generation of low-level vision accelerators is urgently needed. In this article, we propose an end-to-end algorithm-hardware generation framework, EffiVision, on field-programmable gate array (FPGA), aimed at generating highly energy-efficient dedicated accelerators for low-level vision neural networks. EffiVision proposes a hardware template that features multiple parallelisms and large architecture exploration spaces specifically designed to accommodate the characteristics of low-level vision networks. Then, it employs activation-weight aware mixed-precision quantization and FPGA-aware NNLUTs to search the suitable hardware parameters within the hardware template, generating highly energy-efficient accelerators tailored for low-level vision networks. We used EffiVision to perform hardware generation for three low-level vision neural networks fast super-resolution convolutional neural network (FSRCNN), denoising convolutional neural network (DnCNN), and demosaicing convolutional neural network (DMCNN) on Xilinx FPGA development boards, achieving the best energy efficiencies of 174.9, 97.8, and 92.7 GOPS/W, respectively. The generated accelerators of FSRCNN and DnCNN are$1.11\times $and$3.37\times $more efficient than previous works.
Zikang Zhou, Xuyang Duan, Jun Han 0003
IEEE Trans. Very Large Scale Integr. Syst.1
2023 Query-Centric Trajectory Prediction
abstract
Predicting the future trajectories of surrounding agents is essential for autonomous vehicles to operate safely. This paper presents QCNet, a modeling framework toward pushing the boundaries of trajectory prediction. First, we identify that the agent-centric modeling scheme used by existing approaches requires re-normalizing and re-encoding the input whenever the observation window slides forward, leading to redundant computations during online prediction. To overcome this limitation and achieve faster inference, we introduce a query-centric paradigm for scene encoding, which enables the reuse of past computations by learning representations independent of the global spacetime coordinate system. Sharing the invariant scene features among all target agents further allows the parallelism of multi-agent trajectory decoding. Second, even given rich encodings of the scene, existing decoding strategies struggle to capture the multimodality inherent in agents' future behavior, especially when the prediction horizon is long. To tackle this challenge, we first employ anchor-free queries to generate trajectory proposals in a recurrent fashion, which allows the model to utilize different scene contexts when decoding waypoints at different horizons. A refinement module then takes the trajectory proposals as anchors and leverages anchor-based queries to refine the trajectories further. By supplying adaptive and high-quality anchors to the refinement module, our query-based decoder can better deal with the multimodality in the output of trajectory prediction. Our approach ranks 1ston Argoverse 1 and Argoverse 2 motion forecasting benchmarks, outperforming all methods on all main metrics by a large margin. Meanwhile, our model can achieve streaming scene encoding and parallel multi-agent decoding thanks to the query-centric design ethos.
Zikang Zhou, Jianping Wang 0001, Yung-Hui Li, Yu-Kai Huang 0001
CVPR1
2023 Improving the Generalizability of Trajectory Prediction Models with Frenét-Based Domain Normalization
abstract
Predicting the future trajectories of robots' nearby objects plays a pivotal role in applications such as autonomous driving. While learning-based trajectory prediction methods have achieved remarkable performance on public benchmarks, the generalization ability of these approaches remains questionable. The poor generalizability on unseen domains, a well-recognized defect of data-driven approaches, can potentially harm the real-world performance of trajectory prediction models. We are thus motivated to improve models' generalization ability instead of merely pursuing high accuracy on average. Due to the lack of benchmarks for quantifying the generalization ability of trajectory predictors, we first construct a new benchmark called argoverse-shift, where the data distributions of domains are significantly different. Using this benchmark for evaluation, we identify that the domain shift problem seriously hinders the generalization of trajectory predictors since state-of-the-art approaches suffer from severe performance degradation when facing those out-of-distribution scenes. To enhance the robustness of models against domain shift problem, we propose a plug-and-play strategy for domain normalization in trajectory prediction. Our strategy utilizes the Frenét coordinate frame for modeling and can effectively narrow the domain gap of different scenes caused by the variety of road geometry and topology. Experiments show that our strategy noticeably boosts the prediction performance of the state-of-the-art in domains that were previously unseen to the models, thereby improving the generalization ability of data-driven trajectory prediction methods.
Luyao Ye, Zikang Zhou, Jianping Wang 0001
ICRA2
2022 HiVT: Hierarchical Vector Transformer for Multi-Agent Motion Prediction
abstract
Accurately predicting the future motions of surrounding traffic agents is critical for the safety of autonomous ve-hicles. Recently, vectorized approaches have dominated the motion prediction community due to their capability of capturing complex interactions in traffic scenes. How-ever, existing methods neglect the symmetries of the prob-lem and suffer from the expensive computational cost, facing the challenge of making real-time multi-agent motion prediction without sacrificing the prediction performance. To tackle this challenge, we propose Hierarchical Vector Transformer (HiVT) for fast and accurate multi-agent motion prediction. By decomposing the problem into local con-text extraction and global interaction modeling, our method can effectively and efficiently model a large number of agents in the scene. Meanwhile, we propose a translation-invariant scene representation and rotation-invariant spa-tial learning modules, which extract features robust to the geometric transformations of the scene and enable the model to make accurate predictions for multiple agents in a single forward pass. Experiments show that HiVT achieves the state-of-the-art performance on the Argoverse motion forecasting benchmark with a small model size and can make fast multi-agent motion prediction.
Zikang Zhou, Luyao Ye, Jianping Wang 0001, Kui Wu 0001, Kejie Lu
CVPR1