VLDB 2026 Research / reviewers in the wild / expert
Bo Yu 0014
dblp:75/2868-14
· DBLP profile ↗
38ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0002-0139-3622ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 5 first-author · 20 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AIRSTONE: Open sourced hardware accelerators and tools for efficient and safe embodied AI computing
Bo Yu 0014, Yuhui Hao, Yiming Gan, Shaoshan Liu |
Future Gener. Comput. Syst. | 1 |
| 2026 | Corrigendum: Unified and Efficient Factor Graph Accelerator Design for Robotic OptimizationabstractThis is a corrigendum for the article "Unified and Efficient Factor Graph Accelerator Design for Robotic Optimization" published in ACM Trans. Arch. Code Optim. 22, 4, Article 153 (December 2025), 23 pages. Qiang Liu 0011, Yihao Hua, Yuhui Hao, Bo Yu 0014, Shaoshan Liu, Yiming Gan |
ACM Trans. Archit. Code Optim. | 4 |
| 2026 | R2MOAG: Robust Roadside Monocular 3D Object Detection with Adaptive Token and Ground EmbeddingabstractRoadside cameras effectively enhance the perception capabilities of embodied artificial intelligence systems such as vehicles by compensating for the limitations of vehicle-mounted cameras, which are prone to occlusion and have a limited sensing range, thereby improving the safety of autonomous vehicles. However, existing object detection systems often encounter perception errors when handling comprehensive viewpoint noise in roadside scenes, as well as variations in traffic flow, lighting conditions, and camera poses. This makes it challenging for them to perform robustly in complex road environments. To address these issues, we propose \(\mathrm{R^{2}MOAG}\) , a highly robust monocular 3D object detection method for roadside systems, based on ground perception embedding and heterogeneous visual tokens. The proposed method extracts detailed road information through ground plane equations and utilizes heterogeneous visual tokens to focus on foreground features. By integrating low-dimensional ground information with high-dimensional visual features, the model is provided with clear and rich cues for object detection, significantly enhancing its stability. We conducted extensive experiments on the widely recognized roadside datasets DAIR-V2X-I and Rope3D. The results show that, in terms of overall performance, the proposed model achieved a 4.65% and 4.26% improvement in the \(AP_{3D}|_{R40}\) metric for the vehicle category on these two datasets, respectively. Moreover, the model maintained stable recognition performance across various road scenarios and camera poses, demonstrating exceptional robustness. Jie Tang 0003, Haoran Pan, Bo Yu 0014, Shaoshan Liu |
ACM Trans. Cyber Phys. Syst. | 3 |
| 2026 | AIRSPEED: An Open Source Data Production Platform for Embodied Artificial IntelligenceabstractThe development of embodied AI (EAI) critically depends on efficient data acquisition, yet faces persistent challenges including high costs, limited training scenarios, and lack of standardized datasets. We present AIRSPEED, an open source data production platform designed to address these bottlenecks through three core innovations. First, AIRSPEED achieves hardware–software decoupling via unified robot and simulation interfaces, enabling seamless integration with diverse data collection devices and simulation platforms. Second, it supports comprehensive data production methods spanning teleoperation and teaching approaches, as well as synthetic data generation through data synthesis and virtual teleoperation. Third, AIRSPEED automates pyramid-structured dataset construction compatible with both HDF5 and LeRobot formats, significantly reducing manual overhead. Experimental validation demonstrates substantial efficiency gains, achieving up to 35.6× acceleration in dataset construction and 6.0× overall speedup compared to manual workflows. With end-to-end latency as low as 3 ms and compression throughput exceeding 296 MB/s, AIRSPEED establishes a scalable foundation for EAI data production. AISPEED is open sourced on this website: URL . Xuan Xia, Xianqiao Tong, Bo Yu 0014, Jialin Jiao, Xinmin Ding, Hongjun Zhou, Haoran Tong, Tongyi Shen, Ning Ding 0003, Shaoshan Liu |
ACM Trans. Cyber Phys. Syst. | 3 |
| 2026 | PRTF: Polar Space Represented Multi-View 3D Object Detection With Temporal Fusion EnhancementabstractAutonomous driving technology is becoming a significant trend in the development of public transportation. A critical task in autonomous driving perception is 3D object detection, which provides essential data support for downstream applications. Most mainstream 3D object detection methods rely on the Cartesian coordinate system, where they construct object queries to interact with image features and position embedding. However, these methods have the following problems: 1) Sensor-captured detail information diminishes with increasing distance, while pixels represent the same space in Cartesian coordinates, preventing the model from fully leveraging details in closer regions. 2) Multi-view images suffer from spatial misalignment due to overlapping fields of view. 3) The performance of existing single-branch depth prediction networks lacks the necessary accuracy. These issues hinder the feature interaction and affect detection performance. We propose an innovative framework PRTF. Based on Polar space, we design the Two-Stage Transformation Encoder: in the first stage, Dual-DepthNet is used to improve the accuracy of depth prediction. In the second stage, Polar points are generated to address spatial misalignment, enabling effective encoding of details at close distance. In the Temporal Decoder, object queries are leveraged to integrate temporal information, effectively compensating for ambiguous information. By enhancing spatial information at both near and far distances in Polar space, the overall performance of multi-view 3D object detection is significantly improved. PRTF achieves state-of-the-art performance on nuScenes Test with 56.1% mAP and 63.9% NDS, exceeding multi-modal frameworks that combine image and radar data. Jie Tang 0003, Yefei Hou, Bo Yu 0014 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | KARMA: Augmenting Embodied AI Agents with Long-and-Short Term Memory SystemsabstractEmbodied AI agents responsible for executing interconnected, long-sequence household tasks often face difficulties with in-context memory, leading to inefficiencies and errors in task execution. To address this issue, we introduce KARMA, an innovative memory system that integrates longterm and short-term memory modules, enhancing large language models (LLMs) for planning in embodied agents through memory-augmented prompting. Karma distinguishes between long-term and short-term memory, with long-term memory capturing comprehensive 3D scene graphs as representations of the environment, while short-term memory dynamically records changes in objects' positions and states. This dualmemory structure allows agents to retrieve relevant past scene experiences, thereby improving the accuracy and efficiency of task planning. Short-term memory employs strategies for effective and adaptive memory replacement, ensuring the retention of critical information while discarding less pertinent data. Compared to state-of-the-art embodied agents enhanced with memory, our memory-augmented embodied AI agent improves success rates by$1.3 \times$and$2.3 \times$in Composite Tasks and Complex Tasks within the AI2-THOR simulator, respectively, and enhances task execution efficiency by$3.4 \times$and$62.7 \times$. Furthermore, we demonstrate that KARMA's plug-and-play capability allows for seamless deployment on real-world robotic systems, such as mobile manipulation platforms. Through this plug-and-play memory system, KARMA significantly enhances the ability of embodied agents to generate coherent and contextually appropriate plans, making the execution of complex household tasks more efficient. Our code is available at https://github.com/WZX0Swarm0Robotics/KARMA/tree/master. Bo Yu 0014, Junzhe Zhao, Sai Hou, Xing Hu 0001, Yinhe Han 0001, Yiming Gan |
ICRA | 2 |
| 2025 | Dadu-Corki: Algorithm-Architecture Co-Design for Embodied AI-powered Robotic ManipulationabstractEmbodied AI robots have the potential to fundamentally improve the way human beings live and manufacture.Continued progress in the burgeoning field of using large language models to control robots depends critically on an efficient computing substrate, and this trend is strongly evident in manipulation tasks.In particular, today's computing systems for embodied AI robots for manipulation tasks are designed purely based on the interest of algorithm developers, where robot actions are divided into a discrete frame basis.Such an execution pipeline creates high latency and energy consumption.This paper proposes Corki, an algorithm-architecture co-design framework for real-time embodied AI-powered robotic manipulation applications.We aim to decouple LLM inference, robotic control, and data communication in the embodied AI robots' compute pipeline.Instead of predicting action for one single frame, * equal contribution. Yiyang Huang 0002, Yuhui Hao, Bo Yu 0014, Yuxin Yang 0002, Feng Min, Yinhe Han 0001, Lin Ma 0002, Shaoshan Liu, Qiang Liu 0011, Yiming Gan |
ISCA | 3 |
| 2025 | Towards Large-Scale In-Context Reinforcement Learning by Meta-Training in Randomized WorldsabstractIn-Context Reinforcement Learning (ICRL) enables agents to learn automatically and on-the-fly from their interactive experiences. However, a major challenge in scaling up ICRL is the lack of scalable task collections. To address this, we propose the procedurally generated tabular Markov Decision Processes, named AnyMDP. Through a carefully designed randomization process, AnyMDP is capable of generating high-quality tasks on a large scale while maintaining relatively low structural biases. To facilitate efficient meta-training at scale, we further introduce decoupled policy distillation and induce prior information in the ICRL framework. Our results demonstrate that, with a sufficiently large scale of AnyMDP tasks, the proposed model can generalize to tasks that were not considered in the training set through versatile in-context learning paradigms. The scalable task set provided by AnyMDP also enables a more thorough empirical investigation of the relationship between data distribution and ICRL performance. We further show that the generalization of ICRL potentially comes at the cost of increased task diversity and longer adaptation periods. This finding carries critical implications for scaling robust ICRL capabilities, highlighting the necessity of diverse and extensive task design, and prioritizing asymptotic performance over few-shot adaptation. Fan Wang 0021, Pengtao Shao, Bo Yu 0014, Shaoshan Liu, Ning Ding 0003, Yang Cao 0010, Yu Kang 0001, Haifeng Wang 0001 |
NeurIPS | 4 |
| 2025 | EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and RetrievalabstractObject-goal navigation (ObjNav) tasks an agent with navigating to the location of a specific object in an unseen environment.
Embodied agents equipped with large language models (LLMs) and online constructed navigation maps can perform ObjNav in a zero-shot manner. However, existing agents heavily rely on giant LLMs on the cloud, e.g., GPT-4, while directly switching to small LLMs, e.g., LLaMA3.2-11b, suffer from significant success rate drops due to limited model capacity for understanding complex navigation maps, which prevents deploying ObjNav on local devices.
At the same time, the long prompt introduced by the navigation map description will cause high planning latency on local devices.
In this paper, we propose EfficientNav to enable on-device efficient LLM-based zero-shot ObjNav. To help the smaller LLMs better understand the environment, we propose semantics-aware memory retrieval to prune redundant information in navigation maps.
To reduce planning latency, we propose discrete memory caching and attention-based memory clustering to efficiently save and re-use the KV cache.
Extensive experimental results demonstrate that EfficientNav
achieves 11.1\% improvement in success rate on HM3D benchmark over GPT-4-based baselines,
and demonstrates 6.7$\times$ real-time latency reduction and 4.7$\times$ end-to-end latency reduction over GPT-4 planner. Our code is available on https://github.com/PKU-SEC-Lab/EfficientNav. Sunjian Zheng, Tong Xie, Tianshi Xu, Bo Yu 0014, Fan Wang 0021, Jie Tang 0003, Shaoshan Liu |
NeurIPS | 5 |
| 2025 | KINDRED: Heterogeneous Split-Lock Architecture for Safe Autonomous MachinesabstractWith the increasing practicality of autonomous vehicles and drones, the importance of reliability requirements has escalated substantially. In many instances, traditional system designs tend to overlook reliability issues, emphasizing primarily on performance constraints. However, certain designers may opt for a lock-step (redundant) system design, duplicating every component, which in turn can result in significant performance, energy, and cost overheads. In software for autonomous machines, such as self-driving vehicles, performance degradation can increase reaction time, posing safety risks and reducing mission success rates. This article introduces a novel multi-domain lock-step system design, Kindred , which places a strong emphasis on maximizing reliability while minimizing performance overhead. The proposed approach capitalizes on the inherent diversity in fault tolerance among various tasks within autonomous machine software, intelligently scheduling only the vulnerable nodes in the lock-domain. The primary challenge addressed in this study involves the intelligent task scheduling across different domains, complemented by efficient error detection and correction in the lock-domain. In a real system demonstration, we illustrate the effectiveness of Kindred , showcasing its ability to attain the same level of reliability as a full lock-step system while incurring only a mere 2.8% overhead, as opposed to a fully split system, indicating the advantages and potential of our multi-domain lock-step system design in achieving high reliability without compromising performance. Yiming Gan, Jingwen Leng, Bo Yu 0014, Yuhao Zhu 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow OptimizationabstractPath planning is a critical task for autonomous driving, aiming to generate smooth, collision-free, and feasible paths based on input perception and localization information. The planning task is both highly time-sensitive and computationally intensive, posing significant challenges to resource-constrained autonomous driving hardware. In this article, we propose an end-to-end framework for accelerating path planning on FPGA platforms. This framework focuses on accelerating quadratic programming (QP) solving, which is the core of optimization-based path planning and has the most computationally-intensive workloads. Our method leverages a hardware-friendly alternating direction method of multipliers (ADMM) to solve QP problems while employing a highly parallelizable preconditioned conjugate gradient (PCG) method for solving the associated linear systems. We analyze the sparse patterns of matrix operations in QP and design customized storage schemes along with efficient sparse matrix multiplication and sparse matrix-vector multiplication units. Our customized design significantly reduces resource consumption for data storage and computation while dramatically speeding up matrix operations. Additionally, we propose a multi-level dataflow optimization strategy. Within individual operators, we achieve acceleration through parallelization and pipelining. For different operators in an algorithm, we analyze inter-operator data dependencies to enable fine-grained pipelining. At the system level, we map different steps of the planning process to the CPU and FPGA and pipeline these steps to enhance end-to-end throughput. We implement and validate our design on the AMD ZCU102 platform. Our implementation achieves state-of-the-art performance in both latency and energy efficiency compared with existing works, including an average 1.48× speedup over the best FPGA-based design, a 2.89× speedup compared with the state-of-the-art QP solver on an Intel i7-11800H CPU, a 5.62× speedup over an ARM Cortex-A57 embedded CPU, and a 1.56× speedup over state-of-the-art GPU-based work. Furthermore, our design delivers a 2.05× improvement in throughput compared with the state-of-the-art FPGA-based design. Hongzheng Tian, Bo Yu 0014, Shaoshan Liu, Sitao Huang |
ACM Trans. Archit. Code Optim. | 5 |
| 2025 | Elevation-Aware Map Matching Model Leveraging Transfer Learning in Sparse Data ConditionsabstractMap matching is a pivotal component of intelligent urban transportation, offering foundational data for technologies such as path planning, traffic analysis, and trajectory analysis. Diverging from conventional rule-based and topological map matching algorithms, we approach the map matching task from a data-driven perspective, presenting an Elevation-Aware Map Matching Model under conditions of sparse data. This paper initiates from the vehicular standpoint, constructing an Elevation-Aware Unit utilizing imagery and sensor data to acquire elevation information for diverse urban roads. Subsequently, this unit is integrated into the map matching model, enhancing the model’s resilience to noise. Concurrently, employing a Fine-tuning transfer learning approach, we formulate a cross-domain map matching model to maximize the reduction of model development costs. The model undergoes testing on real-world datasets, employing four metrics for evaluation. The results indicate the superiority of this map matching model over existing counterparts, particularly in intricate urban road scenarios where the model exhibits outstanding performance. Additionally, we validate the effectiveness of the Elevation-Aware Unit, underscoring the significance of height information for map matching models. Jie Tang 0003, Sunjian Zheng, Bo Yu 0014, Xue (Steve) Liu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | ORIANNA: An Accelerator Generation Framework for Optimization-based Robotic ApplicationsabstractDespite extensive efforts, existing approaches to design accelerators for optimization-based robotic applications have limitations. Some approaches focus on accelerating general matrix operations, but they fail to fully exploit the specific sparse structure commonly found in many robotic algorithms. On the other hand, certain methods require manual design of dedicated accelerators, resulting in inefficiencies and significant non-recurring engineering (NRE) costs. Yuhui Hao, Yiming Gan, Bo Yu 0014, Qiang Liu 0011, Yinhe Han 0001, Zishen Wan, Shaoshan Liu |
ASPLOS (2) | 3 |
| 2024 | Accelerating Autonomous Path Planning on FPGAs with Sparsity-Aware HW/SW Co-OptimizationsabstractPath planning is a critical task in autonomous driving systems, with quadratic programming being the most time-consuming component. Solving quadratic programming problems using a CPU not only takes a long time but can also lead to high power consumption and costs. In this work, we propose an FPGA-based acceleration method for quadratic programming based path planning problems. Our approach leverages an operator splitting solver for quadratic programs (OSQP) and employs the preconditioned conjugate gradient (PCG) method for solving linear equations, which proves to be more scalable and hardware-friendly than the original direct method. We propose optimizations for better memory management, and boost processing throughput and reduce execution time by task level and operator level parallelism with hardware pipelining. Our FPGA-based implementation achieves up to 1.8× speedup and 3.2× power reduction compared with the Intel i5 CPU, 3.1× speedup compared with ARM Cortex-A57. Hongzheng Tian, Bo Yu 0014, Shaoshan Liu, Sitao Huang |
FPGA | 5 |
| 2024 | Dataflow Accelerator Architecture for Autonomous Machine ComputingabstractCommercial autonomous machines is a thriving sector, one that is likely the next ubiquitous computing platform, after Personal Computers (PC), cloud computing, and mobile computing. Nevertheless, a suitable computing substrate for autonomous machines is missing, and many companies are forced to develop ad hoc computing solutions that are neither principled nor extensible. By analyzing the demands of autonomous machine computing, this article proposes Dataflow Accelerator Architecture (DAA), a modern instantiation of the classic dataflow principle, that matches the characteristics of autonomous machine software. Shaoshan Liu, Yuhao Zhu 0001, Bo Yu 0014, Jean-Luc Gaudiot, Guangrong Gao |
ICCAD | 3 |
| 2024 | A Sparsity-Aware Autonomous Path Planning Accelerator with Algorithm-Architecture Co-DesignabstractPath planning is a critical task in autonomous driving systems that is most susceptible to real-time constraints but often demands computationally intensive mathematical solvers, two contradictory goals. This conflict makes the computing of path planning a paramount challenge. At the heart of most path planners is the quadratic programming (QP) solver, which places excessive demands on the CPU in real-world autonomous driving applications. In this paper, we present an FPGA-based acceleration framework for path planning problems. Our approach leverages an operator splitting solver for quadratic programs (OSQP) and employs the preconditioned conjugate gradient (PCG) method for solving linear systems, which are customized to be more hardware-friendly than prior works. Specific memory management and parallel processing were tailored to the matrix pattern, and the incorporation of pipelining was executed to enhance throughput and execution speed. Our FPGA-based implementation achieves state-of-the-art performance against existing works, including an average 1.98× speedup compared with the state-of-the-art QP solver on Intel i7-11800H CPU, 3.90× speedup over an ARM Cortex-A57 embedded CPU, and 12.3× speedup over an NVIDIA RTX 3090 GPU. Hongzheng Tian, Bo Yu 0014, Shaoshan Liu, Sitao Huang |
ICCAD | 5 |
| 2024 | 3-D Line Matching Network Based on Matching Existence Guidance and Knowledge DistillationabstractIn applications, such as scene reconstruction and odometry, accurate matching associations for 3-D lines are crucial. Real-world scenes introduce inconsistencies due to variations in perspective, leading to nonoverlapping data acting as noise. Accurately matching partially overlapping sets of 3-D lines becomes challenging, potentially resulting in failed scene reconstruction and erroneous positioning. Prior approaches relied on the traditional iterative closest line (ICL) methods, involving iterative calculations and sensitivity to initial poses, and were prone to matching failures in low-overlap rate data and singular pattern scenes. Existing 3-D line matching networks either did not consider the noise in 3-D line collections or failed to retain more valid matching pairs, while these models often require a larger number of parameters and inference time. To address these issues, this article proposes matching existence guidance module (MEG)-Net, a Plücker line matching network guided by the existence of matches. It leverages the rich geometric characteristics of 3-D lines represented as Plücker lines, enhancing feature robustness. By guiding the model to handle the noisy data through match existence guidance, it improves the model’s performance on partially overlapping 3-D line data. Experiments on the indoor and outdoor data sets and the Out of Distribution (OOD) data sets demonstrate that the MEG-Net outperforms traditional methods and baseline models in 3-D line matching, with better scalability and noise robustness, achieving state-of-the-art results. Additionally, we propose an innovative knowledge distillation method based on the matching matrices, training a more efficient MEG-Net mini student model with approximately 70% fewer parameters and multiply accumulate operations (MACs), while maintaining superior performance and faster inference speeds on the indoor data sets. Jie Tang 0003, Bo Yu 0014, Xue (Steve) Liu |
IEEE Internet Things J. | 3 |
| 2023 | BLITZCRANK: Factor Graph Accelerator for Motion PlanningabstractFactor graph is a graph representing the factorization of a probability distribution function and serves as a perfect abstraction in many autonomous machine computing stacks, such as planning, localization, tracking and control, which are challenging tasks for autonomous systems with real-time and energy constraints.In this paper, we present BLITZCRANK, an accelerator for motion planning algorithms using the abstraction of a factor graph. By formulating motion planning as a factor graph inference, we successfully reduce the scale of the problem and utilize the inherent matrix sparsity. BLITZCRANK is able to realize the user-defined optimal design by finding the optimal order of the factor graph inference. With a domain specific balancing order, BLITZCRANK achieves up to 7.4× speed up and 29.7× energy reduction compared to the software implementation on Intel CPU. Yuhui Hao, Yiming Gan, Bo Yu 0014, Qiang Liu 0011, Shaoshan Liu, Yuhao Zhu 0001 |
DAC | 3 |
| 2023 | Invited: Autonomous Driving Digital Twin Empowered Design Automation: An Industry PerspectiveabstractDesigning reliable computing systems for autonomous driving is extremely challenging, as the performance and reliability of the systems have to be thoroughly evaluated under an extremely large amount of driving scenarios. Physically constructing scenarios and conducting testing for autonomous driving systems is time-consuming and expensive. To minimize the need for physical testing and improve development efficiency, we developed a digital-twin-based simulation, which can generate an integral, precise, and comprehensive representation of physical scenarios. In this paper, we share our experiences with the digital-twin-based simulation for autonomous driving, particularly the design requirements and components demanded to facilitate virtual environment construction and design verification, which could greatly improve development efficiency. Bo Yu 0014, Jie Tang 0003, Shaoshan Liu |
DAC | 1 |
| 2023 | HAU$\mathbf {M^3}$: A Height Aware Urban Map Matching Mechanism
Jie Tang 0003, Sunjian Zheng, Bo Yu 0014, Shaoshan Liu |
MobiQuitous (1) | 3 |
| 2023 | Data Fusion in Infrastructure-Augmented Autonomous Driving System: Why? Where? and How?abstractThis article is the first to provide a thorough system design overview along with the fusion methods selection criteria of a real-world cooperative autonomous driving system enabled by the Internet of Things (IoT), named infrastructure-augmented autonomous driving (IAAD). We present an in-depth introduction to the IAAD hardware and software on both road side and vehicle side. We extensively characterize the IAAD system and observe that the network condition fluctuation along the road is the main roadblock for cooperative autonomous driving. To address this challenge, we propose new fusion methods, dubbed “interframe fusion” and “planning fusion” to complement the state-of-the-art “intraframe fusion.” We demonstrate that each fusion method has its own benefit and constraint. In order to select the best fusion method under varying network conditions, we propose “fusion criteria” to instruct the IAAD system to intelligently make the selection and implement a system framework named adaptive spatial-temporal (S–T) choice to realize the adaptive fusion guided by the “fusion criteria.” Our real-world field data verifies that S–T choice has significantly improved autonomous driving’s safety and reliability by decreasing the fusion miss ratio from 30% to 7% and remain the planning displacement error within the 1.7 m instead of 4 m when the network condition exacerbates. Bo Yu 0014, Jie Tang 0003, Shuaiwen Song, Cong Liu 0005, Yang Hu 0001 |
IEEE Internet Things J. | 3 |
| 2023 | An Energy Efficient and Runtime Reconfigurable Accelerator for Robotic LocalizationabstractAccurate and efficient localization of robots under limited on-board resources has fueled specialized localization accelerators. Despite many recent efforts, accelerating robotic localization is still fundamentally challenging. To tackle the challenges, the paper proposes a configurable hardware architecture and a design space optimization method to automatically generate an optimal accelerator design under the design constraints. Data locality, sparsity, and fixed-point arithmetic optimization techniques that are specific to the localization algorithm are exploited to customize the accelerator. In addition, a low-cost runtime configuration mechanism is proposed to enable the accelerator to continuously optimize itself at runtime according to the operating environment to save power while sustaining performance and accuracy. The evaluation on FPGA demonstrates that the proposed accelerator achieves orders of magnitude performance improvement and/or energy savings compared to the software implementation on Intel and Arm CPUs; and substantially outperforms existing FPGA accelerators in terms of performance and energy. Qiang Liu 0011, Yuhui Hao, Weizhuang Liu, Bo Yu 0014, Yiming Gan, Jie Tang 0003, Shaoshan Liu, Yuhao Zhu 0001 |
IEEE Trans. Computers | 4 |
| 2022 | Factor Graph Accelerator for LiDAR-Inertial Odometry (Invited Paper)abstractFactor graph is a graph representing the factorization of a probability distribution function, and has been utilized in many autonomous machine computing tasks, such as localization, tracking, planning and control etc. We are developing an architecture with the goal of using factor graph as a common abstraction for most, if not, all autonomous machine computing tasks. If successful, the architecture would provide a very simple interface of mapping autonomous machine functions to the underlying compute hardware. As a first step of such an attempt, this paper presents our most recent work of developing a factor graph accelerator for LiDAR-Inertial Odometry (LIO), an essential task in many autonomous machines, such as autonomous vehicles and mobile robots. By modeling LIO as a factor graph, the proposed accelerator not only supports multi-sensor fusion such as LiDAR, inertial measurement unit (IMU), GPS, etc., but solves the global optimization problem of robot navigation in batch or incremental modes. Our evaluation demonstrates that the proposed design significantly improves the real-time performance and energy efficiency of autonomous machine navigation systems. The initial success suggests the potential of generalizing the factor graph architecture as a common abstraction for autonomous machine computing, including tracking, planning, and control etc. Yuhui Hao, Bo Yu 0014, Qiang Liu 0011, Shaoshan Liu, Yuhao Zhu 0001 |
ICCAD | 2 |
| 2022 | Braum: Analyzing and Protecting Autonomous Machine Software StackabstractAutonomous machines, such as Autonomous Vehicles (AV), are vulnerable to a variety of different faults such as radiation-induced soft/transient errors, adversarial attacks, and software bugs, which all jeopardize the reliability of autonomous machines. How vulnerable the AV software stack is to different error sources, however, remains an open question. This paper performs comprehensively fault injections to study how the AV software stack behaves under different error sources. We show that algorithms in an AV software stack inherently possess different forms of masking mechanisms. Based on the characteristic of the inherent fault tolerance mechanisms, we formalize the notion of Fault Tolerance Level (FTL), which quantifies how faults in an algorithm can be masked and/or attenuated without affecting the actuator commands, providing opportunities to relax fault protection. Leveraging the FTL formulation, we propose a dynamic protection system, which, at the high level, spends the limited protection budget (e.g., spatial/temporal redundancy) on the most vulnerable parts of the AV software (i.e., with the lowest FTL). Using Autoware as a case study, we show that our system reduces the error rate of AV software stack by more than 90% with negliaible performance overhead. Yiming Gan, Paul N. Whatmough, Jingwen Leng, Bo Yu 0014, Shaoshan Liu, Yuhao Zhu 0001 |
ISSRE | 4 |
| 2022 | Brief Industry Paper: The Necessity of Adaptive Data Fusion in Infrastructure-Augmented Autonomous Driving SystemabstractThis paper is the first to provide a thorough system design overview along with the fusion methods selection criteria of a real-world cooperative autonomous driving system, named Infrastructure-Augmented Autonomous Driving or IAAD. We present an in-depth introduction of the IAAD hardware and software on both road-side and vehicle-side computing/communication platforms. We extensively characterize the IAAD system in the context of real-world deployment scenarios and observe that the network condition fluctuates along the road is currently the main technical roadblock for cooperative autonomous driving. To address this challenge, we propose new fusion methods, dubbed “inter-frame fusion” and “planning fusion” to complement the current state-of-the-art “intra-frame fusion”. We demonstrate that each fusion method has its own benefit and constraint. Adaptively choosing the fusion method according to the real-world condition will benefit the SoV without the violation of the SoV's safety requirements. Shaoshan Liu, Bo Yu 0014, Jie Tang 0003, Shuaiwen Song, Cong Liu 0005, Yang Hu 0001 |
RTAS | 4 |
| 2021 | On Designing Computing Systems for Autonomous Vehicles: a PerceptIn Case StudyabstractPerceptIn develops and commercializes autonomous vehicles for micromobility around the globe. This paper makes a holistic summary of PerceptIn's development and operating experiences. It provides the business tale behind our product, and presents the development of the computing system for our vehicles. We illustrate the design decision made for the computing system, and show the advantage of offloading localization workloads onto an FPGA platform. Bo Yu 0014, Jie Tang 0003, Shaoshan Liu |
ASP-DAC | 1 |
| 2021 | Invited: Towards Fully Intelligent Transportation through Infrastructure-Vehicle Cooperative Autonomous Driving: Challenges and OpportunitiesabstractThe infrastructure-vehicle cooperative autonomous driving approach relies on the cooperation between intelligent roads and intelligent vehicles. This approach is not only safer but also more economical compared to the traditional on-vehicle-only autonomous driving. In this paper, we introduce the real-world deployment experiences of infrastructure-vehicle cooperative autonomous driving by PerceptIn, where a three-stage development roadmap is taken: infrastructure-augmented autonomous driving (IAAD), infrastructure-guided autonomous driving (IGAD), and infrastructure-planned autonomous driving (IPAD). We then discuss the future research challenges and opportunities for such approach. Shaoshan Liu, Bo Yu 0014, Jie Tang 0003, Qi Zhu 0002 |
DAC | 2 |
| 2021 | Eudoxus: Characterizing and Accelerating Localization in Autonomous Machines Industry Track PaperabstractWe develop and commercialize autonomous machines, such as logistic robots and self-driving cars, around the globe. A critical challenge to our—and any—autonomous machine is accurate and efficient localization under resource constraints, which has fueled specialized localization accelerators recently. Prior acceleration efforts are point solutions in that they each specialize for a specific localization algorithm. In real-world commercial deployments, however, autonomous machines routinely operate under different environments and no single localization algorithm fits all the environments. Simply stacking together point solutions not only leads to cost and power budget overrun, but also results in an overly complicated software stack. This paper demonstrates our new software-hardware co-designed framework for autonomous machine localization, which adapts to different operating scenarios by fusing fundamental algorithmic primitives. Through characterizing the software framework, we identify ideal acceleration candidates that contribute significantly to the end-to-end latency and/or latency variation. We show how to co-design a hardware accelerator to systematically exploit the parallelisms, locality, and common building blocks inherent in the localization framework. We build, deploy, and evaluate an FPGA prototype on our next-generation self-driving cars. To demonstrate the flexibility of our framework, we also instantiate another FPGA prototype targeting drones, which represent mobile autonomous machines. We achieve about $2 \times$ speedup and $4 \times$ energy reduction compared to widely-deployed, optimized implementations on general-purpose platforms. Yiming Gan, Bo Yu 0014, Boyuan Tian, Leimeng Xu, Shaoshan Liu, Qiang Liu 0011, Jie Tang 0003, Yuhao Zhu 0001 |
HPCA | 2 |
| 2021 | Archytas: A Framework for Synthesizing and Dynamically Optimizing Accelerators for Robotic LocalizationabstractDespite many recent efforts, accelerating robotic computing is still fundamentally challenging for two reasons. First, robotics software stack is extremely complicated. Manually designing an accelerator while meeting the latency, power, and resource specifications is unscalable. Second, the environment in which an autonomous machine operates constantly changes; a static accelerator design leads to wasteful computation. Weizhuang Liu, Bo Yu 0014, Yiming Gan, Qiang Liu 0011, Jie Tang 0003, Shaoshan Liu, Yuhao Zhu 0001 |
MICRO | 2 |
| 2021 | Brief Industry Paper: The Matter of Time - A General and Efficient System for Precise Sensor Synchronization in Robotic ComputingabstractTime synchronization is a critical task in robotic computing such as autonomous driving. In the past few years, as we developed advanced robotic applications, our synchronization system has evolved as well. In this paper, we first introduce the time synchronization problem and explain the challenges of time synchronization, especially in robotic workloads. Summarizing these challenges, we then present a general hardware synchronization system for robotic computing, which delivers high synchronization accuracy while maintaining low energy and resource consumption. The proposed hardware synchronization system is a key building block in our future robotic products. Shaoshan Liu, Bo Yu 0014, Kunai Zhang, Yisong Qiao, Thomas Yuang Li, Jie Tang 0003, Yuhao Zhu 0001 |
RTAS | 2 |
| 2020 | π-Map: A Decision-Based Sensor Fusion with Global Optimization for Indoor MappingabstractIn this paper, we propose π-map, a tightly coupled fusion mechanism that dynamically consumes LiDAR and sonar data to generate reliable and scalable indoor maps for autonomous robot navigation. The key novelty of π-map over previous attempts is the utilization of a fusion mechanism that works in three stages: the first LiDAR scan matching stage efficiently generates initial key localization poses; the second optimization stage is used to eliminate errors accumulated from the previous stage and guarantees that accurate large-scale maps can be generated; then the final revisit scan fusion stage effectively fuses the LiDAR map and the sonar map to generate a highly accurate representation of the indoor environment. We evaluate π-map on both large and small environments and verify its superiority over existing fusion methods. Zhiliu Yang, Bo Yu 0014, Jie Tang 0003, Shaoshan Liu, Chen Liu 0001 |
IROS | 2 |
| 2020 | Building the Computing System for Autonomous Micromobility Vehicles: Design Constraints and Architectural OptimizationsabstractThis paper presents the computing system design in our commercial autonomous vehicles, and provides a detailed performance, energy, and cost analyses. Drawing from our commercial deployment experience, this paper has two objectives. First, we highlight design constraints unique to autonomous vehicles that might change the way we approach existing architecture problems. Second, we identify new architecture and systems problems that are perhaps less studied before but are critical to autonomous vehicles. Bo Yu 0014, Leimeng Xu, Jie Tang 0003, Shaoshan Liu, Yuhao Zhu 0001 |
MICRO | 1 |
| 2020 | $\pi$π-BA: Bundle Adjustment Hardware Accelerator Based on Distribution of 3D-Point ObservationsabstractBundle adjustment (BA) is a fundamental optimization technique used in many crucial applications, including 3D scene reconstruction, robotic localization, camera calibration, autonomous driving, street view map generation, and even space exploration etc. Essentially, BA is a joint non-linear optimization problem, and one which can consume a significant amount of time and power, especially for large optimization problems. Previous approaches of optimizing BA performance heavily rely on parallel processing or distributed computing, which trade higher power consumption for higher performance. In this article we propose p-BA, the first hardware-software co-designed BA hardware accelerator that exploits custom hardware to simultaneously achieve higher performance and power efficiency. Specifically, based on our key observation that not all 3D points appear on all images in a BA problem, we designed a Co-Observation Optimization technique to accelerate BA operations with optimized usage of memory and computation resources. In addition, we developed a hardware-friendly differentiation method, which combines the analytic and forward automatic differentiation to calculate derivatives of projection function in the BA problem. We have implemented the proposed design on an embedded FPGA SoC, and experimental results confirm that p-BA outperforms the existing software implementations in terms of performance and power consumption. Qiang Liu 0011, Shuzhen Qin, Bo Yu 0014, Jie Tang 0003, Shaoshan Liu |
IEEE Trans. Computers | 3 |
| 2019 | π-BA: Bundle Adjustment Acceleration on Embedded FPGAs with Co-observation OptimizationabstractBundle adjustment (BA) is a fundamental optimization technique used in many crucial applications, including 3D scene reconstruction, robotic localization, camera calibration, autonomous driving, space exploration, street view map generation etc. Essentially, BA is a joint non-linear optimization problem, and one which can consume a significant amount of time and power, especially for large optimization problems. Previous approaches of optimizing BA performance heavily rely on parallel processing or distributed computing, which trade higher power consumption for higher performance. In this paper we propose π-BA, the first hardware-software co-designed BA engine on an embedded FPGA-SoC that exploits custom hardware for higher performance and power efficiency. Specifically, based on our key observation that not all points appear on all images in a BA problem, we designed and implemented a Co-Observation Optimization technique to accelerate BA operations with optimized usage of memory and computation resources. Experimental results confirm that π-BA outperforms the existing software implementations in terms of performance and power consumption. Shuzhen Qin, Qiang Liu 0011, Bo Yu 0014, Shaoshan Liu |
FCCM | 3 |
| 2019 | Edge Computing for Autonomous Driving: Opportunities and ChallengesabstractSafety is the most important requirement for autonomous vehicles; hence, the ultimate challenge of designing an edge computing ecosystem for autonomous vehicles is to deliver enough computing power, redundancy, and security so as to guarantee the safety of autonomous vehicles. Specifically, autonomous driving systems are extremely complex; they tightly integrate many technologies, including sensing, localization, perception, decision making, as well as the smooth interactions with cloud platforms for high-definition (HD) map generation and data storage. These complexities impose numerous challenges for the design of autonomous driving edge computing systems. First, edge computing systems for autonomous driving need to process an enormous amount of data in real time, and often the incoming data from different sensors are highly heterogeneous. Since autonomous driving edge computing systems are mobile, they often have very strict energy consumption restrictions. Thus, it is imperative to deliver sufficient computing power with reasonable energy consumption, to guarantee the safety of autonomous vehicles, even at high speed. Second, in addition to the edge system design, vehicle-to-everything (V2X) provides redundancy for autonomous driving workloads and alleviates stringent performance and energy constraints on the edge side. With V2X, more research is required to define how vehicles cooperate with each other and the infrastructure. Last, safety cannot be guaranteed when security is compromised. Thus, protecting autonomous driving edge computing systems against attacks at different layers of the sensing and computing stack is of paramount concern. In this paper, we review state-of-the-art approaches in these areas as well as explore potential solutions to address these challenges. Shaoshan Liu, Liangkai Liu, Jie Tang 0003, Bo Yu 0014, Yifan Wang 0005, Weisong Shi |
Proc. IEEE | 4 |
| 2018 | π-SoC: Heterogeneous SoC Architecture for Visual Inertial SLAM ApplicationsabstractIn recent years, we have observed a clear trend in the rapid rise of autonomous vehicles and robotics. One of the core technologies enabling these applications, Simultaneous Localization And Mapping (SLAM), imposes two main challenges: first, these workloads are computationally intensive and they often have real-time requirements; second, these workloads run on battery-powered mobile devices with limited energy budget. Hence, performance should be improved while simultaneously reducing energy consumption, two rather contradicting goals by conventional wisdom. Previous attempts to optimize SLAM performance and energy efficiency usually involve optimizing one function and fail to approach the problem systematically. In this paper, we first study the characteristics of visual inertial SLAM workloads on existing heterogeneous SoCs. Then based on the initial findings, we propose π-SoC, a heterogeneous SoC design that systematically optimize the IO interface, the memory hierarchy, as well as the the hardware accelerator. We implemented this system on a Xilinx Zynq UltraScale MPSoC and was able to deliver over 60 FPS performance with average power less than 5 W. Jie Tang 0003, Bo Yu 0014, Shaoshan Liu, Zhe Zhang 0006, Weikang Fang |
IROS | 2 |
| 2017 | FPGA-based ORB feature extraction for real-time visual SLAMabstractSimultaneous Localization And Mapping (SLAM) is the problem of constructing or updating a map of an unknown environment while simultaneously keeping track of an agent's location within it. How to enable SLAM robustly and durably on mobile, or even IoT grade devices, is the main challenge faced by the industry today. The main problems we need to address are: 1.) how to accelerate the SLAM pipeline to meet real-time requirements; and 2.) how to reduce SLAM energy consumption to extend battery life. After delving into the problem, we found out that feature extraction is indeed the bottleneck of performance and energy consumption. Hence, in this paper, we design, implement, and evaluate a hardware ORB feature extractor and prove that our design is a great balance between performance and energy consumption compared with ARM Krait and Intel Core i5. Weikang Fang, Bo Yu 0014, Shaoshan Liu |
FPT | 3 |
| 2010 | A Reconfigurable Hebbian Eigenfilter for Neurophysiological Spike Train AnalysisabstractThe emergence of multi-electrode array enables the study of real-time neurophysiological activities across multiple regions of the brain. However, the real-time extracellular action potentials recorded on any electrode represent the simultaneous electrical activity of an unknown number of neurons which present a critical challenge to the accuracy of interpretation and identification of the neural circuitry in the subsequent analysis. In this paper, we present a principal component analysis approach utilizing Hebbian eigenfilter to identify the corresponding electrical activities of each neuron, namely spike sorting. The Hebbian eigenfilter greatly simplifies the computational complexity of eigen-projection. An efficient FPGA-based Hebbian eigenfilter is proposed. The performance, accuracy and power consumption of our Hebbian eigenfilter are thoroughly evaluated through synthetic spike trains. The proposal enables real-time spike sorting and analysis, and leads the way towards future motor and cognitive neuroprosthetics. Bo Yu 0014, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev, Yihe Sun, Chi-Sang Poon |
FPL | 1 |