Xuyang Duan

dblp:279/9599 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-2609-846XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 A Memory-Efficient LLM Accelerator with Q-K Correlation Prediction using Cluster-Based Associative Array for Selective KV Accessing
abstract
Attention-based LLMs excel in text generation but face redundant computations in autoregressive token generation. While KV cache mitigates this, it introduces increased memory access overhead as sequences grow. We propose Sella, a hardware-software co-design using cluster-based associative arrays to predict Q-K correlations, enabling selective KV cache access and reducing memory access without retraining. Sella includes a specialized accelerator featuring a prediction engine to improve performance and energy efficiency. Experiments show Sella achieves $2.1 \times$, $93.8 \times$, $31.4 \times$, and $53.5 \times$ speedup over SpAtten, Sanger, TITAN RTX GPU, and Xeon CPU, respectively, reducing off-chip memory access by up to $66 \%$ with negligible accuracy loss. -Large Language Models, KV Cache, Accelerator
Zikang Zhou, Xuyang Duan, Jun Han 0003
DAC3
2025 Multi-DOF Fusion: A Flexible Fusion Strategy for Reducing Redundancy in CNN Workloads
Zikang Zhou, Siyao Dai, Xuyang Duan, Jun Han 0003
ACM Great Lakes Symposium on VLSI4
2025 HCTSR: A Hybrid CNN/Transformer Super-Resolution Processor with Depth-Scalable Non-Overlapping Window Attention
Xinhua Shi, Xuyang Duan, Jun Han 0003
ACM Great Lakes Symposium on VLSI3
2025 A Redundant Parallel Continuum Manipulator With Stiffness-Varying and Force-Sensing Capability
abstract
This paper presents the design, analysis, and validation of a novel redundant planar parallel continuum manipulator (PCM) consisting of four flexible links coupled at the rigid mid-and end-platform. To address complex geometry/static hybrid constraints at the rigid platforms, a general framework is developed for the kinetostatics modeling and analysis. Benefiting from the redundant actuation design, the Cartesian stiffness of the studied PCM can be further adjusted as a secondary task to positioning. Utilizing the proposed deflection-based force sensing method, the external load exerted on the end-effector can also be identified by measuring the pose of rigid platforms. To validate the proposed design, a prototype is built, on which a series of experiments have been conducted. The results show that, with the proposed kinetostatics models, the prototype exhibits a mean position and orientation error of 0.52 mm and 0.41$^{\circ}$. Finally, the capability of stiffness-varying and force-sensing is demonstrated to verify the feasibility and potential applications of the designed redundant PCM.Note to Practitioners—This paper is motivated by the need for more adaptable robotic systems in industrial and interactive settings. In this work, a novel redundant parallel continuum manipulator with variable passive compliance and force-sensing capability is presented. This planar manipulator comprises four flexible limbs, each driven by independent linear modules, thus the redundancy in the actuation can be leveraged for varying Cartesian stiffness. By means of measuring the pose of the end-effector, the interaction force exerted on the end-effector can be identified according to the force sensing model. The structural compliance of the manipulator enhances safety in human-robot collaboration, while variable stiffness accommodates the specific requirements of different interaction tasks. Furthermore, the force-sensing capability allows for precise feedback control, significantly improving safety during interactions with humans. Therefore, the presented PCM has the potential for interactive manipulation applications, including delicate and fragile object handling, as well as human-robot collaboration in assembly tasks.
Shujie Tang, Zhenkun Liang, Xuyang Duan, Hao Wang 0015, Genliang Chen
IEEE Trans Autom. Sci. Eng.4
2025 Design of a Novel Force-Controlled End-Effector With Passive Structural Compliance and Intrinsic Contact Sensing
abstract
This paper presents a novel active force-controlled end-effector with passive structural compliance and intrinsic contact sensing capabilities. Slender elastic beams are introduced to provide the end-effector with structural compliance to accommodate inevitable fluctuations through deformation during operation. Meanwhile, flexible strain gauges are embedded into the slender elastic beams to endow the end-effector with intrinsic sensing capability. A kinetostatic model is established in closed-form to map the local deflections of strain gauges to the contact force with the environment. On this basis, model-based feedback control can be readily implemented to actively adjust the contact force between the end-effector and the workpiece. A prototype is developed, on which a variety of experiments are conducted to validate the effectiveness of the proposed end-effector. The results demonstrate that the end-effector can control the contact force accurately, with a maximum error of around 0.65N. Furthermore, a demonstration of robotic grinding using the developed prototype showcases its potential for practical applications.
Tianyi Yan, Jianhuan Chen, Siyue Yao, Xuyang Duan, Yanjun Wang 0011, Hao Wang 0015, Genliang Chen
IEEE Trans Autom. Sci. Eng.4
2024 ML-Fusion: Determining Memory Levels for Data Reuse Between DNN Layers
abstract
With the increasing complexity of applications and the improvement of computational power, modern neural networks (DNNs) have become more memory-intensive. To address the bandwidth problem, modern hardware architectures often incorporate multi-level memory to efficiently reuse data with different reuse distances. Additionally, a promising technique to reduce DNN bandwidth requirements is layer fusion which reuses inter-layer data at on-chip memory. However, previous studies on inter-layer data reuse scheduling have focused primarily on reusing at the outermost on-chip memory, neglecting the exploration of multi-level architecture, which represents a significant optimization space.
Zikang Zhou, Xuyang Duan, Jun Han 0003
ACM Great Lakes Symposium on VLSI2
2024 An Energy-Efficient BNN Accelerator With Two-Stage Value Prediction for Sparse-Edge Gesture Recognition
abstract
In recent years, natural, flexible, and contactless vision-based gesture recognition has received significant attention in human-computer interaction. However, employing convolutional neural networks (CNNs) for RGB or RGB-D gestures can result in excessive power consumption and poor energy efficiency, making them unsuitable for embedded systems. In this paper, we propose a lightweight sparse binarized neural network (sBNN) model for edge gesture recognition that achieves an accuracy of 89.43%-99.92% on four open-source gesture datasets with$\leq 20.26$million operations (MOP) and$\leq 15.83$-Kilobytes (KB) parameters. We find high channel-level sparsity in the activation maps of sBNN when edge gestures are used as inputs. The sparse activation maps have multiple identical activation vectors called sparse activation vectors (SAV), which lead to highly repeated calculations. In order to avoid this issue, we propose a two-stage value prediction approach to skip these calculations, achieving a speedup of 1.03x-1.83x. Moreover, to reduce on- chip memory, the compression technique is applied to the sparse activation maps, providing a compression rate of 1.72x-3.45x. Finally, we implement an energy-efficient sparse BNN accelerator (SBA) on an embedded field-programmable gate array (FPGA). The experimental results show that our SBA has a latency of 26.3-46.8-$\mu \text{s}$, a power consumption of 0.807 W, and an energy efficiency of 536.22-952.70-GOPS/W at 50-MHz frequency. Our SBA offers lower latency, lower power consumption, and higher energy efficiency than previous state-of-the-art gesture recognition accelerators.
Yitong Rong, Xuyang Duan, Xu Cheng 0002, Xiaoyang Zeng, Jun Han 0003
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 A Design Framework for Generating Energy-Efficient Accelerator on FPGA Toward Low-Level Vision
abstract
Low-level vision algorithms play an increasingly crucial role in a wide range of applications, such as biomedical, security, and autopilot. The low-level vision accelerators have also been extensively researched. As low-level vision is often deployed in embedded devices, its accelerators need to achieve high energy efficiency. Meanwhile, the broad application scenarios of low-level vision contribute to its rapid iteration. Designing energy-efficient accelerators for quickly evolving low-level vision algorithms demands substantial effort. Therefore, a design framework specifically tailored for the generation of low-level vision accelerators is urgently needed. In this article, we propose an end-to-end algorithm-hardware generation framework, EffiVision, on field-programmable gate array (FPGA), aimed at generating highly energy-efficient dedicated accelerators for low-level vision neural networks. EffiVision proposes a hardware template that features multiple parallelisms and large architecture exploration spaces specifically designed to accommodate the characteristics of low-level vision networks. Then, it employs activation-weight aware mixed-precision quantization and FPGA-aware NNLUTs to search the suitable hardware parameters within the hardware template, generating highly energy-efficient accelerators tailored for low-level vision networks. We used EffiVision to perform hardware generation for three low-level vision neural networks fast super-resolution convolutional neural network (FSRCNN), denoising convolutional neural network (DnCNN), and demosaicing convolutional neural network (DMCNN) on Xilinx FPGA development boards, achieving the best energy efficiencies of 174.9, 97.8, and 92.7 GOPS/W, respectively. The generated accelerators of FSRCNN and DnCNN are$1.11\times $and$3.37\times $more efficient than previous works.
Zikang Zhou, Xuyang Duan, Jun Han 0003
IEEE Trans. Very Large Scale Integr. Syst.2