EDBT 2026 Demo / reviewers in the wild / expert
Yuxin Yang 0002
dblp:146/9561-2
· DBLP profile ↗
7ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0002-3007-3705ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dadu-Corki: Algorithm-Architecture Co-Design for Embodied AI-powered Robotic ManipulationabstractEmbodied AI robots have the potential to fundamentally improve the way human beings live and manufacture.Continued progress in the burgeoning field of using large language models to control robots depends critically on an efficient computing substrate, and this trend is strongly evident in manipulation tasks.In particular, today's computing systems for embodied AI robots for manipulation tasks are designed purely based on the interest of algorithm developers, where robot actions are divided into a discrete frame basis.Such an execution pipeline creates high latency and energy consumption.This paper proposes Corki, an algorithm-architecture co-design framework for real-time embodied AI-powered robotic manipulation applications.We aim to decouple LLM inference, robotic control, and data communication in the embodied AI robots' compute pipeline.Instead of predicting action for one single frame, * equal contribution. Yiyang Huang 0002, Yuhui Hao, Bo Yu 0014, Yuxin Yang 0002, Feng Min, Yinhe Han 0001, Lin Ma 0002, Shaoshan Liu, Qiang Liu 0011, Yiming Gan |
ISCA | 5 |
| 2025 | A Data-Centric Software-Hardware Co-Designed Architecture for Large-Scale Graph ProcessingabstractGraph processing plays an important role in many practical applications. However, the inherent characteristics of graph processing, including random memory access and the low computation-to-communication ratio, make it difficult to efficiently execute on traditional computing architectures, such as CPUs and GPUs. Near-memory computing has the characteristics of low latency and high bandwidth. It is widely regarded as a promising direction for designing graph processing accelerators. However, the storage space of a single device cannot meet the demand of large-scale graph processing. Using multiple devices will bring lots of inter-device data transmission, which may counteract the benefits of near-memory computing. To fundamentally reduce the data transmission overhead, we propose a data-centric graph processing framework for systems with multiple near-memory computing devices. The framework uses a data-centric programming model as the software hardware interface. For software, we propose an optimized data flow and a heuristic multi-step weighted maximum matching algorithm to achieve efficient inter-device communication and ensure load balancing. For hardware, we design a data reuse driven task controller and a data type-aware on-chip memory, which can effectively improve the utilization of the on-chip memory. Compared with the two most recent near-memory graph accelerators, our framework significantly reduces energy consumption and inter-device communication. Zerun Li, Xiaoming Chen 0003, Yuxin Yang 0002, Feng Min, Xiaoyu Zhang 0009, Yinhe Han 0001 |
IEEE Trans. Computers | 3 |
| 2023 | Dadu-RBD: Robot Rigid Body Dynamics Accelerator with Multifunctional PipelinesabstractRigid body dynamics is a core technology in the robotics field. In trajectory optimization and model predictive control algorithms, there are usually a large number of rigid body dynamics computing tasks. Using CPUs to process these tasks consumes a lot of time, which will affect the real-time performance of robots. To this end, we propose a multifunctional robot rigid body dynamics accelerator, named Dadu-RBD, to address the performance bottleneck. By analyzing different functions commonly used in robot dynamics calculations, we summarize their relationships and characteristics, then optimize them according to the hardware. Based on this, Dadu-RBD can fully reuse common hardware modules when processing different computing tasks. By dynamically switching the dataflow path, Dadu-RBD can accelerate various dynamics functions without reconfiguring the hardware. We design the Round-Trip Pipeline and Structure-Adaptive Pipelines for Dadu-RBD, which can greatly improve the throughput of the accelerator. Robots with different structures and parameters can be optimized specifically. Compared with the state-of-the-art CPU, GPU dynamics libraries and FPGA accelerator, Dadu-RBD can significantly improve the performance. Yuxin Yang 0002, Xiaoming Chen 0003, Yinhe Han 0001 |
MICRO | 1 |
| 2022 | Re-FeMAT: A Reconfigurable Multifunctional FeFET-Based Memory ArchitectureabstractMost of current processing-in-memory (PIM) architectures are application specific, that is, they can only accelerate particular functions, e.g., matrix-vector dot product for neural network acceleration. However, practical applications usually involve various functions. In order to accelerate different functions, various accelerators, and dedicated circuits have been proposed. In this work, by exploring the similarities among some commonly used dedicated circuits, we adopt ferroelectric field-effect transistors (FeFETs) to build a reconfigurable multifunctional memory architecture named Re-FeMAT. Re-FeMAT is composed of multiple processing elements (PEs). Each PE is not only a nonvolatile memory array, but also can perform logic operations (i.e., the PIM mode), convolutions (i.e., the binary convolutional neural network and the convolutional neural network (CNN) acceleration mode) and content search (i.e., the ternary content-addressable memory (TCAM) mode) without changing the circuit structure. Re-FeMAT can support applications that require multiple functions. As an example, by configuring different PEs to different working modes and using a simulated annealing algorithm or a tabu search algorithm to optimize the task-PE assignment, Re-FeMAT can completely accelerate few-shot learning applications. Our simulation results based on a calibrated FeFET model show that the proposed Re-FeMAT architecture achieves better performance and power efficiency than the previous FeMAT architecture. Compared with FeFET-based single-functional circuits, though the power dissipation of Re-FeMAT is higher in some modes, the power-delay product is still smaller. Compared with a state-of-the-art FeFET-based multifunctional accelerator named attention-in-memory, Re-FeMAT achieves lower power, latency, and energy when accelerating a complete few-shot learning task. Xiaoyu Zhang 0009, Rui Liu 0045, Yuxin Yang 0002, Yinhe Han 0001, Xiaoming Chen 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Dadu-CD: Fast and Efficient Processing-in-Memory Accelerator for Collision DetectionabstractCollision detection is a fundamental task in motion planning of robotics. Typically, the performance of collision detection is the bottleneck of an entire motion planning, and so does the energy consumption. Several hardware accelerators have been proposed for collision detection, which achieves higher performance and energy efficiency than general-purpose CPUs and GPUs. However, existing accelerators are still facing the limited memory bandwidth bottleneck, due to the large data volume required by the parallel processing cores and the limited DRAM bandwidth. In this work, we propose a novel collision detection accelerator by employing the processing-in-memory technique. We elaborate the in-memory processing architecture to fully utilize the internal bandwidth of DRAM banks. To make the algorithm and hardware suitable for in-memory processing to be highly efficient, a set of innovative software and hardware techniques are also proposed. Compared with a state-of-the-art ASIC-based collision detection accelerator, both performance and energy efficiency of our accelerator are significantly improved. Yuxin Yang 0002, Xiaoming Chen 0003, Yinhe Han 0001 |
DAC | 1 |
| 2020 | Accelerating RRT Motion Planning Using TCAMabstractReal-time motion planning is important for robot movement. In motion planning, path search and collision detection are two performance bottlenecks. In this paper, we adopt a range-based matching scheme with ternary content-addressable memories (TCAMs) to accelerate the processes of both nearest neighbor search and collision detection. In our approach, the nearest node search and collision detection can be both processed in a few TCAM lookup cycles. The evaluation shows that the TCAM-based accelerator is 236× faster than CPU for motion planning tasks. It is 5.4× faster and at least 8.8× more energy-efficient than a state-of-the-art dedicated ASIC-based accelerator. Yuxin Yang 0002, Shiqi Lian, Xiaoming Chen 0003, Yinhe Han 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2020 | DaDu Series - Fast and Efficient Robot AcceleratorsabstractResearch on accelerators for robotics is increasing. This article introduces the kinematics, motion planning and collision detection algorithms and our accelerators in robotics, and then analyzes their advantages, disadvantages and bottlenecks. In view of the shortcomings of the existing accelerators, this paper will show a series accelerators named "DaDu" that we have proposed. For kinematics, we have proposed Dadu [1] to accelerate the inverse kinematics algorithm, which achieves 1700x speedup than the CPU implementation, 30x speedup than the GPU implementation, and 776x higher energy efficiency than the GPU implementation. For motion planning, we have proposed Dadu-P [2] to accelerate the PRM algorithm. It can get 26.5x speedup than an existing CPU-based approach for collision detection. Furthermore, with an incremental approach, the performance of motion planning can further be improved by 10x while the solution quality is degraded by 10% only. For the collision detection algorithm in motion planning, the proposed accelerator Dadu-CD [3] elaborates the in-memory processing architecture, achieving at least 5x speedup than Dadu-P in the total planning time and 9.55x lower energy consumption than Dadu-P. Yinhe Han 0001, Yuxin Yang 0002, Xiaoming Chen 0003, Shiqi Lian |
ICCAD | 2 |