EDBT 2026 Demo / reviewers in the wild / expert
Shaoshan Liu
dblp:32/4500
· DBLP profile ↗
67ranked-venue papers
17as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 50 · 14 first-author · 24 since 2021Software engineering, systems software and programming languages · 9 · 6 since 2021Artificial intelligence and machine learning · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IEEE Computer Society in the Age of AI: A Strategic Position Paper for the Next Decade
Jean-Luc Gaudiot, Shaoshan Liu |
COMPSAC | 2 |
| 2026 | AIRSTONE: Open sourced hardware accelerators and tools for efficient and safe embodied AI computing
Bo Yu 0014, Yuhui Hao, Yiming Gan, Shaoshan Liu |
Future Gener. Comput. Syst. | 4 |
| 2026 | Corrigendum: Unified and Efficient Factor Graph Accelerator Design for Robotic OptimizationabstractThis is a corrigendum for the article "Unified and Efficient Factor Graph Accelerator Design for Robotic Optimization" published in ACM Trans. Arch. Code Optim. 22, 4, Article 153 (December 2025), 23 pages. Qiang Liu 0011, Yihao Hua, Yuhui Hao, Bo Yu 0014, Shaoshan Liu, Yiming Gan |
ACM Trans. Archit. Code Optim. | 5 |
| 2026 | R2MOAG: Robust Roadside Monocular 3D Object Detection with Adaptive Token and Ground EmbeddingabstractRoadside cameras effectively enhance the perception capabilities of embodied artificial intelligence systems such as vehicles by compensating for the limitations of vehicle-mounted cameras, which are prone to occlusion and have a limited sensing range, thereby improving the safety of autonomous vehicles. However, existing object detection systems often encounter perception errors when handling comprehensive viewpoint noise in roadside scenes, as well as variations in traffic flow, lighting conditions, and camera poses. This makes it challenging for them to perform robustly in complex road environments. To address these issues, we propose \(\mathrm{R^{2}MOAG}\) , a highly robust monocular 3D object detection method for roadside systems, based on ground perception embedding and heterogeneous visual tokens. The proposed method extracts detailed road information through ground plane equations and utilizes heterogeneous visual tokens to focus on foreground features. By integrating low-dimensional ground information with high-dimensional visual features, the model is provided with clear and rich cues for object detection, significantly enhancing its stability. We conducted extensive experiments on the widely recognized roadside datasets DAIR-V2X-I and Rope3D. The results show that, in terms of overall performance, the proposed model achieved a 4.65% and 4.26% improvement in the \(AP_{3D}|_{R40}\) metric for the vehicle category on these two datasets, respectively. Moreover, the model maintained stable recognition performance across various road scenarios and camera poses, demonstrating exceptional robustness. Jie Tang 0003, Haoran Pan, Bo Yu 0014, Shaoshan Liu |
ACM Trans. Cyber Phys. Syst. | 4 |
| 2026 | AIRSPEED: An Open Source Data Production Platform for Embodied Artificial IntelligenceabstractThe development of embodied AI (EAI) critically depends on efficient data acquisition, yet faces persistent challenges including high costs, limited training scenarios, and lack of standardized datasets. We present AIRSPEED, an open source data production platform designed to address these bottlenecks through three core innovations. First, AIRSPEED achieves hardware–software decoupling via unified robot and simulation interfaces, enabling seamless integration with diverse data collection devices and simulation platforms. Second, it supports comprehensive data production methods spanning teleoperation and teaching approaches, as well as synthetic data generation through data synthesis and virtual teleoperation. Third, AIRSPEED automates pyramid-structured dataset construction compatible with both HDF5 and LeRobot formats, significantly reducing manual overhead. Experimental validation demonstrates substantial efficiency gains, achieving up to 35.6× acceleration in dataset construction and 6.0× overall speedup compared to manual workflows. With end-to-end latency as low as 3 ms and compression throughput exceeding 296 MB/s, AIRSPEED establishes a scalable foundation for EAI data production. AISPEED is open sourced on this website: URL . Xuan Xia, Xianqiao Tong, Bo Yu 0014, Jialin Jiao, Xinmin Ding, Hongjun Zhou, Haoran Tong, Tongyi Shen, Ning Ding 0003, Shaoshan Liu |
ACM Trans. Cyber Phys. Syst. | 12 |
| 2025 | Dadu-Corki: Algorithm-Architecture Co-Design for Embodied AI-powered Robotic ManipulationabstractEmbodied AI robots have the potential to fundamentally improve the way human beings live and manufacture.Continued progress in the burgeoning field of using large language models to control robots depends critically on an efficient computing substrate, and this trend is strongly evident in manipulation tasks.In particular, today's computing systems for embodied AI robots for manipulation tasks are designed purely based on the interest of algorithm developers, where robot actions are divided into a discrete frame basis.Such an execution pipeline creates high latency and energy consumption.This paper proposes Corki, an algorithm-architecture co-design framework for real-time embodied AI-powered robotic manipulation applications.We aim to decouple LLM inference, robotic control, and data communication in the embodied AI robots' compute pipeline.Instead of predicting action for one single frame, * equal contribution. Yiyang Huang 0002, Yuhui Hao, Bo Yu 0014, Yuxin Yang 0002, Feng Min, Yinhe Han 0001, Lin Ma 0002, Shaoshan Liu, Qiang Liu 0011, Yiming Gan |
ISCA | 9 |
| 2025 | Towards Large-Scale In-Context Reinforcement Learning by Meta-Training in Randomized WorldsabstractIn-Context Reinforcement Learning (ICRL) enables agents to learn automatically and on-the-fly from their interactive experiences. However, a major challenge in scaling up ICRL is the lack of scalable task collections. To address this, we propose the procedurally generated tabular Markov Decision Processes, named AnyMDP. Through a carefully designed randomization process, AnyMDP is capable of generating high-quality tasks on a large scale while maintaining relatively low structural biases. To facilitate efficient meta-training at scale, we further introduce decoupled policy distillation and induce prior information in the ICRL framework. Our results demonstrate that, with a sufficiently large scale of AnyMDP tasks, the proposed model can generalize to tasks that were not considered in the training set through versatile in-context learning paradigms. The scalable task set provided by AnyMDP also enables a more thorough empirical investigation of the relationship between data distribution and ICRL performance. We further show that the generalization of ICRL potentially comes at the cost of increased task diversity and longer adaptation periods. This finding carries critical implications for scaling robust ICRL capabilities, highlighting the necessity of diverse and extensive task design, and prioritizing asymptotic performance over few-shot adaptation. Fan Wang 0021, Pengtao Shao, Bo Yu 0014, Shaoshan Liu, Ning Ding 0003, Yang Cao 0010, Yu Kang 0001, Haifeng Wang 0001 |
NeurIPS | 5 |
| 2025 | EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and RetrievalabstractObject-goal navigation (ObjNav) tasks an agent with navigating to the location of a specific object in an unseen environment.
Embodied agents equipped with large language models (LLMs) and online constructed navigation maps can perform ObjNav in a zero-shot manner. However, existing agents heavily rely on giant LLMs on the cloud, e.g., GPT-4, while directly switching to small LLMs, e.g., LLaMA3.2-11b, suffer from significant success rate drops due to limited model capacity for understanding complex navigation maps, which prevents deploying ObjNav on local devices.
At the same time, the long prompt introduced by the navigation map description will cause high planning latency on local devices.
In this paper, we propose EfficientNav to enable on-device efficient LLM-based zero-shot ObjNav. To help the smaller LLMs better understand the environment, we propose semantics-aware memory retrieval to prune redundant information in navigation maps.
To reduce planning latency, we propose discrete memory caching and attention-based memory clustering to efficiently save and re-use the KV cache.
Extensive experimental results demonstrate that EfficientNav
achieves 11.1\% improvement in success rate on HM3D benchmark over GPT-4-based baselines,
and demonstrates 6.7$\times$ real-time latency reduction and 4.7$\times$ end-to-end latency reduction over GPT-4 planner. Our code is available on https://github.com/PKU-SEC-Lab/EfficientNav. Sunjian Zheng, Tong Xie, Tianshi Xu, Bo Yu 0014, Fan Wang 0021, Jie Tang 0003, Shaoshan Liu |
NeurIPS | 8 |
| 2025 | A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow OptimizationabstractPath planning is a critical task for autonomous driving, aiming to generate smooth, collision-free, and feasible paths based on input perception and localization information. The planning task is both highly time-sensitive and computationally intensive, posing significant challenges to resource-constrained autonomous driving hardware. In this article, we propose an end-to-end framework for accelerating path planning on FPGA platforms. This framework focuses on accelerating quadratic programming (QP) solving, which is the core of optimization-based path planning and has the most computationally-intensive workloads. Our method leverages a hardware-friendly alternating direction method of multipliers (ADMM) to solve QP problems while employing a highly parallelizable preconditioned conjugate gradient (PCG) method for solving the associated linear systems. We analyze the sparse patterns of matrix operations in QP and design customized storage schemes along with efficient sparse matrix multiplication and sparse matrix-vector multiplication units. Our customized design significantly reduces resource consumption for data storage and computation while dramatically speeding up matrix operations. Additionally, we propose a multi-level dataflow optimization strategy. Within individual operators, we achieve acceleration through parallelization and pipelining. For different operators in an algorithm, we analyze inter-operator data dependencies to enable fine-grained pipelining. At the system level, we map different steps of the planning process to the CPU and FPGA and pipeline these steps to enhance end-to-end throughput. We implement and validate our design on the AMD ZCU102 platform. Our implementation achieves state-of-the-art performance in both latency and energy efficiency compared with existing works, including an average 1.48× speedup over the best FPGA-based design, a 2.89× speedup compared with the state-of-the-art QP solver on an Intel i7-11800H CPU, a 5.62× speedup over an ARM Cortex-A57 embedded CPU, and a 1.56× speedup over state-of-the-art GPU-based work. Furthermore, our design delivers a 2.05× improvement in throughput compared with the state-of-the-art FPGA-based design. Hongzheng Tian, Bo Yu 0014, Shaoshan Liu, Sitao Huang |
ACM Trans. Archit. Code Optim. | 6 |
| 2024 | ORIANNA: An Accelerator Generation Framework for Optimization-based Robotic ApplicationsabstractDespite extensive efforts, existing approaches to design accelerators for optimization-based robotic applications have limitations. Some approaches focus on accelerating general matrix operations, but they fail to fully exploit the specific sparse structure commonly found in many robotic algorithms. On the other hand, certain methods require manual design of dedicated accelerators, resulting in inefficiencies and significant non-recurring engineering (NRE) costs. Yuhui Hao, Yiming Gan, Bo Yu 0014, Qiang Liu 0011, Yinhe Han 0001, Zishen Wan, Shaoshan Liu |
ASPLOS (2) | 7 |
| 2024 | Accelerating Autonomous Path Planning on FPGAs with Sparsity-Aware HW/SW Co-OptimizationsabstractPath planning is a critical task in autonomous driving systems, with quadratic programming being the most time-consuming component. Solving quadratic programming problems using a CPU not only takes a long time but can also lead to high power consumption and costs. In this work, we propose an FPGA-based acceleration method for quadratic programming based path planning problems. Our approach leverages an operator splitting solver for quadratic programs (OSQP) and employs the preconditioned conjugate gradient (PCG) method for solving linear equations, which proves to be more scalable and hardware-friendly than the original direct method. We propose optimizations for better memory management, and boost processing throughput and reduce execution time by task level and operator level parallelism with hardware pipelining. Our FPGA-based implementation achieves up to 1.8× speedup and 3.2× power reduction compared with the Intel i5 CPU, 3.1× speedup compared with ARM Cortex-A57. Hongzheng Tian, Bo Yu 0014, Shaoshan Liu, Sitao Huang |
FPGA | 6 |
| 2024 | Dataflow Accelerator Architecture for Autonomous Machine ComputingabstractCommercial autonomous machines is a thriving sector, one that is likely the next ubiquitous computing platform, after Personal Computers (PC), cloud computing, and mobile computing. Nevertheless, a suitable computing substrate for autonomous machines is missing, and many companies are forced to develop ad hoc computing solutions that are neither principled nor extensible. By analyzing the demands of autonomous machine computing, this article proposes Dataflow Accelerator Architecture (DAA), a modern instantiation of the classic dataflow principle, that matches the characteristics of autonomous machine software. Shaoshan Liu, Yuhao Zhu 0001, Bo Yu 0014, Jean-Luc Gaudiot, Guangrong Gao |
ICCAD | 1 |
| 2024 | A Sparsity-Aware Autonomous Path Planning Accelerator with Algorithm-Architecture Co-DesignabstractPath planning is a critical task in autonomous driving systems that is most susceptible to real-time constraints but often demands computationally intensive mathematical solvers, two contradictory goals. This conflict makes the computing of path planning a paramount challenge. At the heart of most path planners is the quadratic programming (QP) solver, which places excessive demands on the CPU in real-world autonomous driving applications. In this paper, we present an FPGA-based acceleration framework for path planning problems. Our approach leverages an operator splitting solver for quadratic programs (OSQP) and employs the preconditioned conjugate gradient (PCG) method for solving linear systems, which are customized to be more hardware-friendly than prior works. Specific memory management and parallel processing were tailored to the matrix pattern, and the incorporation of pipelining was executed to enhance throughput and execution speed. Our FPGA-based implementation achieves state-of-the-art performance against existing works, including an average 1.98× speedup compared with the state-of-the-art QP solver on Intel i7-11800H CPU, 3.90× speedup over an ARM Cortex-A57 embedded CPU, and 12.3× speedup over an NVIDIA RTX 3090 GPU. Hongzheng Tian, Bo Yu 0014, Shaoshan Liu, Sitao Huang |
ICCAD | 6 |
| 2024 | AICOM-MP: an AI-based monkeypox detector for resource-constrained environmentsabstractUnder the Autonomous Mobile Clinics (AMCs) initiative, the AI Clinics on Mobile (AICOM) project is developing, open sourcing, and standardising health AI technologies on low-end mobile devices to enable health-care access in least-developed countries (LDCs).As the first step, we introduce AICOM-MP, an AI-based monkeypox detector specially aiming for handling images taken from resourceconstrained devices.We have developed AICOM-MP with the following principles: minimisation of gender, racial, and age bias; ability to conduct binary classification without over-relying on computing power; capacity to produce accurate results irrespective of images' background, resolution, and quality.AICOM-MP has achieved stateof-the-art (SOTA) performance.We have hosted AICOM-MP as a web service to allow universal access to monkeypox screening technology, and open-sourced both the source code and the dataset of AICOM-MP to allow health AI professionals to integrate AICOM-MP into their services. Tim Tianyi Yang, Tom Tianze Yang, Shaoshan Liu, Xue (Steve) Liu |
Connect. Sci. | 5 |
| 2023 | BLITZCRANK: Factor Graph Accelerator for Motion PlanningabstractFactor graph is a graph representing the factorization of a probability distribution function and serves as a perfect abstraction in many autonomous machine computing stacks, such as planning, localization, tracking and control, which are challenging tasks for autonomous systems with real-time and energy constraints.In this paper, we present BLITZCRANK, an accelerator for motion planning algorithms using the abstraction of a factor graph. By formulating motion planning as a factor graph inference, we successfully reduce the scale of the problem and utilize the inherent matrix sparsity. BLITZCRANK is able to realize the user-defined optimal design by finding the optimal order of the factor graph inference. With a domain specific balancing order, BLITZCRANK achieves up to 7.4× speed up and 29.7× energy reduction compared to the software implementation on Intel CPU. Yuhui Hao, Yiming Gan, Bo Yu 0014, Qiang Liu 0011, Shaoshan Liu, Yuhao Zhu 0001 |
DAC | 5 |
| 2023 | Invited: Autonomous Driving Digital Twin Empowered Design Automation: An Industry PerspectiveabstractDesigning reliable computing systems for autonomous driving is extremely challenging, as the performance and reliability of the systems have to be thoroughly evaluated under an extremely large amount of driving scenarios. Physically constructing scenarios and conducting testing for autonomous driving systems is time-consuming and expensive. To minimize the need for physical testing and improve development efficiency, we developed a digital-twin-based simulation, which can generate an integral, precise, and comprehensive representation of physical scenarios. In this paper, we share our experiences with the digital-twin-based simulation for autonomous driving, particularly the design requirements and components demanded to facilitate virtual environment construction and design verification, which could greatly improve development efficiency. Bo Yu 0014, Jie Tang 0003, Shaoshan Liu |
DAC | 3 |
| 2023 | Analysis and Optimization of Worst-Case Time Disparity in Cause-Effect ChainsabstractIn automotive systems, an important timing requirement is that the time disparity (the maximum difference among the timestamps of all raw data produced by sensors that an output originates from) must be bounded in a certain range, so that information from different sensors can be correctly synchronized and fused. In this paper, we study the problem of analyzing the worst-case time disparity in cause-effect chains. In particular, we present two bounds, where the first one assumes all chains are independent from each other and the second one takes the fork-join structures into consideration to perform more precise analysis. Moreover, we propose a solution to cut down the worst-case time disparity for a task by designing buffers with proper sizes. Experiments are conducted to show the correctness and effectiveness of both our analysis and optimization methods. Xu Jiang 0004, Xiantong Luo, Nan Guan, Zheng Dong 0002, Shaoshan Liu, Wang Yi 0001 |
DATE | 5 |
| 2023 | HAU$\mathbf {M^3}$: A Height Aware Urban Map Matching Mechanism
Jie Tang 0003, Sunjian Zheng, Bo Yu 0014, Shaoshan Liu |
MobiQuitous (1) | 4 |
| 2023 | An Energy Efficient and Runtime Reconfigurable Accelerator for Robotic LocalizationabstractAccurate and efficient localization of robots under limited on-board resources has fueled specialized localization accelerators. Despite many recent efforts, accelerating robotic localization is still fundamentally challenging. To tackle the challenges, the paper proposes a configurable hardware architecture and a design space optimization method to automatically generate an optimal accelerator design under the design constraints. Data locality, sparsity, and fixed-point arithmetic optimization techniques that are specific to the localization algorithm are exploited to customize the accelerator. In addition, a low-cost runtime configuration mechanism is proposed to enable the accelerator to continuously optimize itself at runtime according to the operating environment to save power while sustaining performance and accuracy. The evaluation on FPGA demonstrates that the proposed accelerator achieves orders of magnitude performance improvement and/or energy savings compared to the software implementation on Intel and Arm CPUs; and substantially outperforms existing FPGA accelerators in terms of performance and energy. Qiang Liu 0011, Yuhui Hao, Weizhuang Liu, Bo Yu 0014, Yiming Gan, Jie Tang 0003, Shaoshan Liu, Yuhao Zhu 0001 |
IEEE Trans. Computers | 7 |
| 2022 | Programming Autonomous Machines : Special Session PaperabstractOne key technical challenge in the age of autonomous machines is the programming of autonomous machines, which demands the synergy across multiple domains, including fundamental computer science, computer architecture, and robotics, and requires expertise from both academia and industry. This paper discusses the programming theory and practices tied to producing real-life autonomous machines, and covers aspects from high-level concepts down to low-level code generation in the context of specific functional requirements, performance expectation, and implementation constraints of autonomous machines. Shaoshan Liu, Xiaoming Li 0010, Tongsheng Geng, Stéphane Zuckerman, Jean-Luc Gaudiot |
EMSOFT | 1 |
| 2022 | Factor Graph Accelerator for LiDAR-Inertial Odometry (Invited Paper)abstractFactor graph is a graph representing the factorization of a probability distribution function, and has been utilized in many autonomous machine computing tasks, such as localization, tracking, planning and control etc. We are developing an architecture with the goal of using factor graph as a common abstraction for most, if not, all autonomous machine computing tasks. If successful, the architecture would provide a very simple interface of mapping autonomous machine functions to the underlying compute hardware. As a first step of such an attempt, this paper presents our most recent work of developing a factor graph accelerator for LiDAR-Inertial Odometry (LIO), an essential task in many autonomous machines, such as autonomous vehicles and mobile robots. By modeling LIO as a factor graph, the proposed accelerator not only supports multi-sensor fusion such as LiDAR, inertial measurement unit (IMU), GPS, etc., but solves the global optimization problem of robot navigation in batch or incremental modes. Our evaluation demonstrates that the proposed design significantly improves the real-time performance and energy efficiency of autonomous machine navigation systems. The initial success suggests the potential of generalizing the factor graph architecture as a common abstraction for autonomous machine computing, including tracking, planning, and control etc. Yuhui Hao, Bo Yu 0014, Qiang Liu 0011, Shaoshan Liu, Yuhao Zhu 0001 |
ICCAD | 4 |
| 2022 | Braum: Analyzing and Protecting Autonomous Machine Software StackabstractAutonomous machines, such as Autonomous Vehicles (AV), are vulnerable to a variety of different faults such as radiation-induced soft/transient errors, adversarial attacks, and software bugs, which all jeopardize the reliability of autonomous machines. How vulnerable the AV software stack is to different error sources, however, remains an open question. This paper performs comprehensively fault injections to study how the AV software stack behaves under different error sources. We show that algorithms in an AV software stack inherently possess different forms of masking mechanisms. Based on the characteristic of the inherent fault tolerance mechanisms, we formalize the notion of Fault Tolerance Level (FTL), which quantifies how faults in an algorithm can be masked and/or attenuated without affecting the actuator commands, providing opportunities to relax fault protection. Leveraging the FTL formulation, we propose a dynamic protection system, which, at the high level, spends the limited protection budget (e.g., spatial/temporal redundancy) on the most vulnerable parts of the AV software (i.e., with the lowest FTL). Using Autoware as a case study, we show that our system reduces the error rate of AV software stack by more than 90% with negliaible performance overhead. Yiming Gan, Paul N. Whatmough, Jingwen Leng, Bo Yu 0014, Shaoshan Liu, Yuhao Zhu 0001 |
ISSRE | 5 |
| 2022 | Brief Industry Paper: The Necessity of Adaptive Data Fusion in Infrastructure-Augmented Autonomous Driving SystemabstractThis paper is the first to provide a thorough system design overview along with the fusion methods selection criteria of a real-world cooperative autonomous driving system, named Infrastructure-Augmented Autonomous Driving or IAAD. We present an in-depth introduction of the IAAD hardware and software on both road-side and vehicle-side computing/communication platforms. We extensively characterize the IAAD system in the context of real-world deployment scenarios and observe that the network condition fluctuates along the road is currently the main technical roadblock for cooperative autonomous driving. To address this challenge, we propose new fusion methods, dubbed “inter-frame fusion” and “planning fusion” to complement the current state-of-the-art “intra-frame fusion”. We demonstrate that each fusion method has its own benefit and constraint. Adaptively choosing the fusion method according to the real-world condition will benefit the SoV without the violation of the SoV's safety requirements. Shaoshan Liu, Bo Yu 0014, Jie Tang 0003, Shuaiwen Song, Cong Liu 0005, Yang Hu 0001 |
RTAS | 1 |
| 2022 | Brief Industry Paper: Enabling Level-4 Autonomous Driving on a Single $1k Off-the-Shelf CardabstractIn the past few years we have developed hardware computing systems for commercial autonomous vehicles, but inevitably the high development cost and long turn-around time have been major roadblocks for commercial deployment. Hence we also explored the potential of software optimization. This paper, for the first-time, shows that it is feasible to enable full leve1-4 autonomous driving workloads on a single off-the-shelf card (Jetson AGX Xavier) for less than ${\$}1\mathrm{k}$, an order of magnitude less than the state-of-the-art systems, while meeting all the requirements of latency. The success comes from the resolution of some important issues shared by existing practices through a series of measures and innovations. Hsin-Hsuan Sung, Yuanchao Xu 0001, Jiexiong Guan, Wei Niu 0002, Bin Ren 0002, Yanzhi Wang 0001, Shaoshan Liu, Xipeng Shen |
RTAS | 7 |
| 2022 | Rise of the Automotive Health-Domain Controllers: Empowering Healthcare Services in Intelligent VehiclesabstractWe are facing a global healthcare crisis today as the healthcare cost is ever climbing, but with the aging population, government fiscal revenue is ever dropping. To address this imminent problem, we can start by enabling affordable anywhere anytime healthcare access through delivering healthcare services on intelligent vehicles. The foundation upon which healthcare services can be provided on intelligent vehicles is an automotive health-domain controller (AHDC), which is missing today. In this article, for the first time, we explain the necessity and define the functionalities of AHDCs. In addition, we delve into the technical challenges, requirements, and feasibility of integrating multiple key features into AHDCs. It is hoped that this will help the community standardize on-vehicle healthcare services provisioning, and lead to the universal adoption of the revolutionary mobile healthcare system. Shaoshan Liu, Yuzhang Huang, Ao Kong, Jie Tang 0003, Xue (Steve) Liu |
IEEE Internet Things J. | 1 |
| 2021 | On Designing Computing Systems for Autonomous Vehicles: a PerceptIn Case StudyabstractPerceptIn develops and commercializes autonomous vehicles for micromobility around the globe. This paper makes a holistic summary of PerceptIn's development and operating experiences. It provides the business tale behind our product, and presents the development of the computing system for our vehicles. We illustrate the design decision made for the computing system, and show the advantage of offloading localization workloads onto an FPGA platform. Bo Yu 0014, Jie Tang 0003, Shaoshan Liu |
ASP-DAC | 3 |
| 2021 | Streaming Data Priority Scheduling Framework for Autonomous Driving by EdgeabstractIn recent years, intelligent vehicles like autonomous vehicles generate a huge amount of sensing data continuously. The computations on those data streams are far beyond the processing capacity of on-board computing. To deal with the streaming data process in real-time, the deployment of streaming data processing system by edge turns to the first choice in terms of performance. However, the existing frameworks cannot satisfy the complicated demands from autonomous driving tasks and lack the ability in supporting the task priority scheduling. In this paper, we propose a streaming data priority scheduling framework for autonomous driving by edge on Spark Streaming and make an implementation on Spark 2.3.0. The proposed framework can identify the priorities among different data processing tasks and implement the task scheduling based on non-preemptive priority queuing theory. To meet differentiated service level requirements, the proposed non-preemptive priority queuing scheduling mechanism considers the priority category of tasks, the distance between vehicles and edge nodes, and the priority weight of vehicles. Experiments show that this mechanism can effectively identify the priority information of different tasks from different vehicles and reduce the end-to-end latency of high-priority tasks by up to 46% than low-priority tasks. Lingbing Yao, Hang Zhao 0016, Jie Tang 0003, Shaoshan Liu, Jean-Luc Gaudiot |
COMPSAC | 4 |
| 2021 | Invited: Towards Fully Intelligent Transportation through Infrastructure-Vehicle Cooperative Autonomous Driving: Challenges and OpportunitiesabstractThe infrastructure-vehicle cooperative autonomous driving approach relies on the cooperation between intelligent roads and intelligent vehicles. This approach is not only safer but also more economical compared to the traditional on-vehicle-only autonomous driving. In this paper, we introduce the real-world deployment experiences of infrastructure-vehicle cooperative autonomous driving by PerceptIn, where a three-stage development roadmap is taken: infrastructure-augmented autonomous driving (IAAD), infrastructure-guided autonomous driving (IGAD), and infrastructure-planned autonomous driving (IPAD). We then discuss the future research challenges and opportunities for such approach. Shaoshan Liu, Bo Yu 0014, Jie Tang 0003, Qi Zhu 0002 |
DAC | 1 |
| 2021 | Eudoxus: Characterizing and Accelerating Localization in Autonomous Machines Industry Track PaperabstractWe develop and commercialize autonomous machines, such as logistic robots and self-driving cars, around the globe. A critical challenge to our—and any—autonomous machine is accurate and efficient localization under resource constraints, which has fueled specialized localization accelerators recently. Prior acceleration efforts are point solutions in that they each specialize for a specific localization algorithm. In real-world commercial deployments, however, autonomous machines routinely operate under different environments and no single localization algorithm fits all the environments. Simply stacking together point solutions not only leads to cost and power budget overrun, but also results in an overly complicated software stack. This paper demonstrates our new software-hardware co-designed framework for autonomous machine localization, which adapts to different operating scenarios by fusing fundamental algorithmic primitives. Through characterizing the software framework, we identify ideal acceleration candidates that contribute significantly to the end-to-end latency and/or latency variation. We show how to co-design a hardware accelerator to systematically exploit the parallelisms, locality, and common building blocks inherent in the localization framework. We build, deploy, and evaluate an FPGA prototype on our next-generation self-driving cars. To demonstrate the flexibility of our framework, we also instantiate another FPGA prototype targeting drones, which represent mobile autonomous machines. We achieve about $2 \times$ speedup and $4 \times$ energy reduction compared to widely-deployed, optimized implementations on general-purpose platforms. Yiming Gan, Bo Yu 0014, Boyuan Tian, Leimeng Xu, Shaoshan Liu, Qiang Liu 0011, Jie Tang 0003, Yuhao Zhu 0001 |
HPCA | 6 |
| 2021 | Archytas: A Framework for Synthesizing and Dynamically Optimizing Accelerators for Robotic LocalizationabstractDespite many recent efforts, accelerating robotic computing is still fundamentally challenging for two reasons. First, robotics software stack is extremely complicated. Manually designing an accelerator while meeting the latency, power, and resource specifications is unscalable. Second, the environment in which an autonomous machine operates constantly changes; a static accelerator design leads to wasteful computation. Weizhuang Liu, Bo Yu 0014, Yiming Gan, Qiang Liu 0011, Jie Tang 0003, Shaoshan Liu, Yuhao Zhu 0001 |
MICRO | 6 |
| 2021 | A KNN Query Method for Autonomous Driving Sensor Data
Jie Tang 0003, Jiehui Zhang, Zhixin Zeng, Shaoshan Liu |
NPC | 4 |
| 2021 | Brief Industry Paper: An Edge-Based High-Definition Map Crowdsourcing Task Distribution Framework for Autonomous DrivingabstractFacing the difficulty and inefficiency of creating and maintaining High-Definition (HD) maps in our commercial deployments, we have developed an edge-based crowdsourcing task distribution framework for HD Map in autonomous driving. Our key observation is that: HD map data crowdsourcing exhibits the diminishing marginal utility thus there exists an inflection point for maximum utility, meanwhile its premature convergence of utility will leave some map updates not notified in time. Based on this observation, we develop a periodic crowdsourcing task distribution framework. It discretizes the demands for collecting source data into different periods and uses an optimal stopping rule to terminate the data collection for the maximum crowdsourcing utility. The experimental results verify that our crowdsourcing framework can achieve high time coverage and high efficiency with lower cost. Donghua Li, Jie Tang 0003, Shaoshan Liu |
RTAS | 3 |
| 2021 | Brief Industry Paper: The Matter of Time - A General and Efficient System for Precise Sensor Synchronization in Robotic ComputingabstractTime synchronization is a critical task in robotic computing such as autonomous driving. In the past few years, as we developed advanced robotic applications, our synchronization system has evolved as well. In this paper, we first introduce the time synchronization problem and explain the challenges of time synchronization, especially in robotic workloads. Summarizing these challenges, we then present a general hardware synchronization system for robotic computing, which delivers high synchronization accuracy while maintaining low energy and resource consumption. The proposed hardware synchronization system is a key building block in our future robotic products. Shaoshan Liu, Bo Yu 0014, Kunai Zhang, Yisong Qiao, Thomas Yuang Li, Jie Tang 0003, Yuhao Zhu 0001 |
RTAS | 1 |
| 2021 | Brief Industry Paper: An Infrastructure-Aided High Definition Map Data Provisioning Service for Autonomous DrivingabstractAs a fundamental component in the autonomous driving technology stack, High Definition Maps (HD map) provide high-precision descriptions of the environment. It enables extremely accurate perception and localization while improving the efficiency of path planning. However, the HD map's extremely large data volume poses great challenges for the real-time and safety requirements of autonomous driving. Based on our real-world deployment experiences, we first demonstrate how the existing data transmission mechanism is weak in supporting HD map services. To address this problem, we propose an HD map data service mechanism on top of Vehicle-to-Infrastructure (V2I) data transmission under a tight time and energy budget. By this mechanism, the selected road side unit (RSU) nodes cooperate on map provisioning tasks and transmit HD map data proportionately. Furthermore, we model the real-time map data service into a partial knapsack problem and develop a greedy data transmission algorithm. Experimental results confirm that the proposed mechanism can ensure the real-time HD map data service meanwhile meeting the energy limits. Jinliang Xie, Jie Tang 0003, Yanzhi Wang 0001, Qi Zhu 0002, Shaoshan Liu |
RTAS | 5 |
| 2021 | Brief Industry Paper: Towards Real-Time 3D Object Detection for Autonomous Vehicles with Pruning SearchabstractIn autonomous driving, 3D object detection is es-sential as it provides basic knowledge about the environment. However, as deep learning based 3D detection methods are usually computation intensive, it is challenging to support realtime 3D object detection on edge-computing devices in selfdriving cars with limited computation and memory resources. To facilitate this, we propose a compiler-aware pruning search framework, to achieve real-time inference of 3D object detection on the resource-limited mobile devices. Specifically, a generator is applied to sample better pruning proposals in the search space based on current proposals with their performance, and an evaluator is adopted to evaluate the sampled pruning proposal performance. To accelerate the search, the evaluator employs Bayesian optimization with an ensemble of neural predictors. We demonstrate in experiments that for the first time, the pruning search framework can achieve real-time 3D object detection on mobile (Samsung Galaxy S20 phone) with state-of-the-art detection performance. Pu Zhao 0001, Wei Niu 0002, Geng Yuan, Yuxuan Cai 0001, Hsin-Hsuan Sung, Shaoshan Liu, Sijia Liu 0001, Xipeng Shen, Bin Ren 0002, Yanzhi Wang 0001, Xue Lin 0001 |
RTAS | 6 |
| 2020 | Real-Time Spatio-Temporal LiDAR Point Cloud CompressionabstractCompressing massive LiDAR point clouds in real-time is critical to autonomous machines such as drones and self-driving cars. While most of the recent prior work has focused on compressing individual point cloud frames, this paper proposes a novel system that effectively compresses a sequence of point clouds. The idea to exploit both the spatial and temporal redundancies in a sequence of point cloud frames. We first identify a key frame in a point cloud sequence and spatially encode the key frame by iterative plane fitting. We then exploit the fact that consecutive point clouds have large overlaps in the physical space, and thus spatially encoded data can be (re-)used to encode the temporal stream. Temporal encoding by reusing spatial encoding data not only improves the compression rate, but also avoids redundant computations, which significantly improves the compression speed. Experiments show that our compression system achieves 40× to 90× compression rate, significantly higher than the MPEG's LiDAR point cloud compression standard, while retaining high end-to-end application accuracies. Meanwhile, our compression system has a compression speed that matches the point cloud generation rate by today LiDARs and out-performs existing compression systems, enabling real-time point cloud transmission. Yu Feng 0007, Shaoshan Liu, Yuhao Zhu 0001 |
IROS | 2 |
| 2020 | π-Map: A Decision-Based Sensor Fusion with Global Optimization for Indoor MappingabstractIn this paper, we propose π-map, a tightly coupled fusion mechanism that dynamically consumes LiDAR and sonar data to generate reliable and scalable indoor maps for autonomous robot navigation. The key novelty of π-map over previous attempts is the utilization of a fusion mechanism that works in three stages: the first LiDAR scan matching stage efficiently generates initial key localization poses; the second optimization stage is used to eliminate errors accumulated from the previous stage and guarantees that accurate large-scale maps can be generated; then the final revisit scan fusion stage effectively fuses the LiDAR map and the sonar map to generate a highly accurate representation of the indoor environment. We evaluate π-map on both large and small environments and verify its superiority over existing fusion methods. Zhiliu Yang, Bo Yu 0014, Jie Tang 0003, Shaoshan Liu, Chen Liu 0001 |
IROS | 5 |
| 2020 | Building the Computing System for Autonomous Micromobility Vehicles: Design Constraints and Architectural OptimizationsabstractThis paper presents the computing system design in our commercial autonomous vehicles, and provides a detailed performance, energy, and cost analyses. Drawing from our commercial deployment experience, this paper has two objectives. First, we highlight design constraints unique to autonomous vehicles that might change the way we approach existing architecture problems. Second, we identify new architecture and systems problems that are perhaps less studied before but are critical to autonomous vehicles. Bo Yu 0014, Leimeng Xu, Jie Tang 0003, Shaoshan Liu, Yuhao Zhu 0001 |
MICRO | 5 |
| 2020 | π-Hub: Large-scale video learning, storage, and retrieval on heterogeneous hardware platforms
Jie Tang 0003, Shaoshan Liu, Jie Cao 0003, Bolin Ding, Jean-Luc Gaudiot, Weisong Shi |
Future Gener. Comput. Syst. | 2 |
| 2020 | $\pi$π-BA: Bundle Adjustment Hardware Accelerator Based on Distribution of 3D-Point ObservationsabstractBundle adjustment (BA) is a fundamental optimization technique used in many crucial applications, including 3D scene reconstruction, robotic localization, camera calibration, autonomous driving, street view map generation, and even space exploration etc. Essentially, BA is a joint non-linear optimization problem, and one which can consume a significant amount of time and power, especially for large optimization problems. Previous approaches of optimizing BA performance heavily rely on parallel processing or distributed computing, which trade higher power consumption for higher performance. In this article we propose p-BA, the first hardware-software co-designed BA hardware accelerator that exploits custom hardware to simultaneously achieve higher performance and power efficiency. Specifically, based on our key observation that not all 3D points appear on all images in a BA problem, we designed a Co-Observation Optimization technique to accelerate BA operations with optimized usage of memory and computation resources. In addition, we developed a hardware-friendly differentiation method, which combines the analytic and forward automatic differentiation to calculate derivatives of projection function in the BA problem. We have implemented the proposed design on an embedded FPGA SoC, and experimental results confirm that p-BA outperforms the existing software implementations in terms of performance and power consumption. Qiang Liu 0011, Shuzhen Qin, Bo Yu 0014, Jie Tang 0003, Shaoshan Liu |
IEEE Trans. Computers | 5 |
| 2019 | π-BA: Bundle Adjustment Acceleration on Embedded FPGAs with Co-observation OptimizationabstractBundle adjustment (BA) is a fundamental optimization technique used in many crucial applications, including 3D scene reconstruction, robotic localization, camera calibration, autonomous driving, space exploration, street view map generation etc. Essentially, BA is a joint non-linear optimization problem, and one which can consume a significant amount of time and power, especially for large optimization problems. Previous approaches of optimizing BA performance heavily rely on parallel processing or distributed computing, which trade higher power consumption for higher performance. In this paper we propose π-BA, the first hardware-software co-designed BA engine on an embedded FPGA-SoC that exploits custom hardware for higher performance and power efficiency. Specifically, based on our key observation that not all points appear on all images in a BA problem, we designed and implemented a Co-Observation Optimization technique to accelerate BA operations with optimized usage of memory and computation resources. Experimental results confirm that π-BA outperforms the existing software implementations in terms of performance and power consumption. Shuzhen Qin, Qiang Liu 0011, Bo Yu 0014, Shaoshan Liu |
FCCM | 4 |
| 2019 | A DAG Refactor Based Automatic Execution Optimization Mechanism for Spark
Hang Zhao 0016, Donghua Li, Jie Tang 0003, Shaoshan Liu |
NPC | 5 |
| 2019 | Edge Computing for Autonomous Driving: Opportunities and ChallengesabstractSafety is the most important requirement for autonomous vehicles; hence, the ultimate challenge of designing an edge computing ecosystem for autonomous vehicles is to deliver enough computing power, redundancy, and security so as to guarantee the safety of autonomous vehicles. Specifically, autonomous driving systems are extremely complex; they tightly integrate many technologies, including sensing, localization, perception, decision making, as well as the smooth interactions with cloud platforms for high-definition (HD) map generation and data storage. These complexities impose numerous challenges for the design of autonomous driving edge computing systems. First, edge computing systems for autonomous driving need to process an enormous amount of data in real time, and often the incoming data from different sensors are highly heterogeneous. Since autonomous driving edge computing systems are mobile, they often have very strict energy consumption restrictions. Thus, it is imperative to deliver sufficient computing power with reasonable energy consumption, to guarantee the safety of autonomous vehicles, even at high speed. Second, in addition to the edge system design, vehicle-to-everything (V2X) provides redundancy for autonomous driving workloads and alleviates stringent performance and energy constraints on the edge side. With V2X, more research is required to define how vehicles cooperate with each other and the infrastructure. Last, safety cannot be guaranteed when security is compromised. Thus, protecting autonomous driving edge computing systems against attacks at different layers of the sensing and computing stack is of paramount concern. In this paper, we review state-of-the-art approaches in these areas as well as explore potential solutions to address these challenges. Shaoshan Liu, Liangkai Liu, Jie Tang 0003, Bo Yu 0014, Yifan Wang 0005, Weisong Shi |
Proc. IEEE | 1 |
| 2018 | Teaching Autonomous Driving Using a Modular and Integrated ApproachabstractIntroduction: Teaching autonomous driving is a challenging task. Indeed, most existing autonomous driving teaching activities focus on a few of the technologies involved. This not only fails to provide a comprehensive coverage, but also sets a high entry barrier for students with different backgrounds. Objective: The primary objective of this study is to present a modular, integrated approach towards teaching autonomous driving. Methods: We organize the technologies used in autonomous driving into modules. This is described in the textbook we have developed as well as a series of multimedia online lectures designed to provide technical overview for each module. Once the students have understood these modules, the experimental platforms for integration we have developed allow the students to fully understand how the modules interact with each other. Results: To verify this teaching approach, we present three case studies: an introductory class on autonomous driving for students with only a basic technology background; a new session in an existing embedded systems class to demonstrate how embedded system technologies can be applied towards autonomous driving; and an industry professional training session to quickly bring up experienced engineers to work in autonomous driving. The results show that students can maintain a high interest level and make great progress by starting with familiar concepts before moving onto other modules. Conclusions: Autonomous driving is not one single technology, but rather a complex system integrating many technologies. Our modular and integrated approach is an effective method in teaching autonomous driving. Jie Tang 0003, Shaoshan Liu, Songwen Pei, Stéphane Zuckerman, Chen Liu 0001, Weisong Shi, Jean-Luc Gaudiot |
COMPSAC (1) | 2 |
| 2018 | PIRVS: An Advanced Visual-Inertial SLAM System with Flexible Sensor Fusion and Hardware Co-DesignabstractIn this paper, we present the PerceptIn Robotics Vision System (PIRVS), a visual-inertial computing hardware with embedded simultaneous localization and mapping (SLAM) algorithm. The PIRVS hardware is equipped with a multi-core processor, a global-shutter stereo camera, and an IMU with precise hardware synchronization. The PIRVS software features a flexible sensor fusion approach to not only tightly integrate visual measurements with inertial measurements and also to loosely couple with additional sensor modalities. It runs in real-time on both PC and the PIRVS hardware. We perform a thorough evaluation of the proposed system using multiple public visual-inertial datasets. Experimental results demonstrate that our system reaches comparable accuracy of state-of-the-art visual-inertial algorithms on PC, while being more efficient on the PIRVS hardware. Zhe Zhang 0006, Shaoshan Liu, Grace Tsai, Hongbing Hu, Chen-Chi Chu |
ICRA | 2 |
| 2018 | π-SoC: Heterogeneous SoC Architecture for Visual Inertial SLAM ApplicationsabstractIn recent years, we have observed a clear trend in the rapid rise of autonomous vehicles and robotics. One of the core technologies enabling these applications, Simultaneous Localization And Mapping (SLAM), imposes two main challenges: first, these workloads are computationally intensive and they often have real-time requirements; second, these workloads run on battery-powered mobile devices with limited energy budget. Hence, performance should be improved while simultaneously reducing energy consumption, two rather contradicting goals by conventional wisdom. Previous attempts to optimize SLAM performance and energy efficiency usually involve optimizing one function and fail to approach the problem systematically. In this paper, we first study the characteristics of visual inertial SLAM workloads on existing heterogeneous SoCs. Then based on the initial findings, we propose π-SoC, a heterogeneous SoC design that systematically optimize the IO interface, the memory hierarchy, as well as the the hardware accelerator. We implemented this system on a Xilinx Zynq UltraScale MPSoC and was able to deliver over 60 FPS performance with average power less than 5 W. Jie Tang 0003, Bo Yu 0014, Shaoshan Liu, Zhe Zhang 0006, Weikang Fang |
IROS | 3 |
| 2018 | Trifo-VIO: Robust and Efficient Stereo Visual Inertial Odometry Using Points and LinesabstractIn this paper, we present the Trifo Visual Inertial Odometry (Trifo-VIO), a tightly-coupled filtering-based stereo VIO system using both points and lines. Line features help improve system robustness in challenging scenarios when point features cannot be reliably detected or tracked, e.g. low-texture environment or lighting change. In addition, we propose a novel lightweight filtering-based loop closing technique to reduce accumulated drift without global bundle adjustment or pose graph optimization. We formulate loop closure as EKF updates to optimally relocate the current sliding window maintained by the filter to past keyframes. We also present the Trifo Ironsides dataset, a new visual-inertial dataset, featuring high-quality synchronized stereo camera and IMU data from the Ironsides sensor [3] with various motion types and textures and millimeter-accuracy groundtruth. To validate the performance of the proposed system, we conduct extensive comparison with state-of-the-art approaches (OKVIS, VINS-MONO and S-MSCKF) using both the public EuRoC dataset and the Trifo Ironsides dataset. Grace Tsai, Zhe Zhang 0006, Shaoshan Liu, Chen-Chi Chu, Hongbing Hu |
IROS | 4 |
| 2017 | FPGA-based ORB feature extraction for real-time visual SLAMabstractSimultaneous Localization And Mapping (SLAM) is the problem of constructing or updating a map of an unknown environment while simultaneously keeping track of an agent's location within it. How to enable SLAM robustly and durably on mobile, or even IoT grade devices, is the main challenge faced by the industry today. The main problems we need to address are: 1.) how to accelerate the SLAM pipeline to meet real-time requirements; and 2.) how to reduce SLAM energy consumption to extend battery life. After delving into the problem, we found out that feature extraction is indeed the bottleneck of performance and energy consumption. Hence, in this paper, we design, implement, and evaluate a hardware ORB feature extractor and prove that our design is a great balance between performance and energy consumption compared with ARM Krait and Intel Core i5. Weikang Fang, Bo Yu 0014, Shaoshan Liu |
FPT | 4 |
| 2013 | OCP: Offload Co-Processor for energy efficiency in embedded mobile systemsabstractIn current embedded mobile systems design, the application processor (AP) is often woken up to service interrupts and user requests. However, this kind of wakeups from sleep is very expensive in terms of battery usage. In the observation that the operating system/driver workloads are very light-weight, in this paper we propose the Offload Co-Processor (OCP) SoC architecture. In the OCP SoC design, when the device is idle, we offload the operating system workloads (mainly interrupt handling workloads) to an ultra-low-power coprocessor. This way, the co-processor would be able to handle most wake-up requests without awakening the heavy-weight AP, thus avoiding the overhead of AP spin-up/down. Using GPS continuous sampling workload as a case study, we show that the proposed OCP SoC design would extend battery life by 3.5 folds. Jie Tang 0003, Chen Liu 0001, Yu-Liang Chou, Shaoshan Liu |
ASAP | 4 |
| 2013 | Pinned OS/Services: A Case Study of XML Parsing on Intel SCC
Jie Tang 0003, Pollawat Thanarungroj, Chen Liu 0001, Shaoshan Liu, Zhimin Gu, Jean-Luc Gaudiot |
J. Comput. Sci. Technol. | 4 |
| 2013 | Acceleration of XML Parsing through PrefetchingabstractExtensible Markup Language (XML) has become a widely adopted standard for data representation and exchange. However, its features also introduce significant overhead threatening the performance of modern applications. In this paper, we present a study of XML parsing and determine that memory-side data loading in the parsing stage incurs a significant performance overhead, as much as the computation does. Hence, we propose memory-side acceleration which incorporates of data prefetching techniques, and can be applied on top of computation-side acceleration to speed up the XML data parsing. To this end, we study here the impact of our proposed scheme on the performance and energy consumption and demonstrated how it is capable of improving performance by up to 20 percent as well as produce up to 12.77 percent of energy saving when implemented in 32-nm technology. In addition, we implement a prefetcher on an platform in an effort to evaluate its implementation feasibility in terms of area and energy overhead. Jie Tang 0003, Shaoshan Liu, Chen Liu 0001, Zhimin Gu, Jean-Luc Gaudiot |
IEEE Trans. Computers | 2 |
| 2013 | Achieving energy efficiency through runtime partial reconfiguration on reconfigurable systemsabstractOne major advantage of reconfigurable computing systems is their ability to reconfigure hardware at runtime. In this paper, we study the feasibility of achieving energy efficiency in reconfigurable computing systems (e.g., FPGAs) through runtime partial reconfiguration (PR) techniques. In the ideal scenario, we use a hardware accelerator to accelerate certain parts of the program execution; when the accelerator is not active, we use partial reconfiguration to unload it to reduce power consumption. Since the reconfiguration process may introduce a high energy overhead, it is unclear whether this approach is efficient. To approach this problem, we first analytically identify the conditions under which partial reconfiguration can reduce energy consumption. Our results indicate that the key to reduce partial reconfiguration energy overhead is to minimize the time overhead of the reconfiguration process. Based on this analysis, we design and implement a fast reconfiguration engine that achieves close-to-ideal throughput on Xilinx Virtex-4 FPGAs. Our fast reconfiguration engine utilizes a master-slave DMA pair to stream data between the SRAM and the Internal Configuration Access Port (ICAP). We experimentally verify our proposed solutions and compare our design to existing energy reduction techniques, such as clock gating. The results of our study show that by using partial reconfiguration to eliminate the power consumption of the accelerator when it is inactive, we can accelerate program execution and at the same time reduce the overall energy consumption by half. Shaoshan Liu, Richard Neil Pittman, Alessandro Forin, Jean-Luc Gaudiot |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2012 | Packer: Parallel Garbage Collection Based on Virtual SpacesabstractThe fundamental challenge of garbage collector (GC) design is to maximize the recycled space with minimal time overhead. For efficient memory management, in many GC designs the heap is divided into large object space (LOS) and normal object space (non-LOS). When either space is full, garbage collection is triggered even though the other space may still have plenty of room, thus leading to inefficient space utilization. Also, space partitioning in existing GC designs implies different GC algorithms for different spaces. This not only prolongs the pause time of garbage collection, but also makes collection inefficient on multiple spaces. To address these problems, we propose Packer, a parallel garbage collection algorithm based on the novel concept of virtual spaces. Instead of physically dividing the heap into multiple spaces, Packer manages multiple virtual spaces in one physical space. With multiple virtual spaces, Packer offers efficient memory management. With one physical space, Packer avoids the problem of an inefficient space utilization. To reduce the garbage collection pause time, we also propose a novel parallelization method that is applicable to multiple virtual spaces. Specifically, we reduce the compacting GC parallelization problem into a discreted acyclic graph (DAG) traversal parallelization problem, and apply it to both normal and large object compaction. Shaoshan Liu, Jie Tang 0003, Ligang Wang 0001, Xiao-Feng Li, Jean-Luc Gaudiot |
IEEE Trans. Computers | 1 |
| 2012 | Minimizing the runtime partial reconfiguration overheads in reconfigurable systems
Shaoshan Liu, Richard Neil Pittman, Alessandro Forin, Jean-Luc Gaudiot |
J. Supercomput. | 1 |
| 2012 | Achieving middleware execution efficiency: hardware-assisted garbage collection operationsabstractAlthough virtualization technologies bring many benefits to cloud computing environments, as the virtual machines provide more features, the middleware layer has become bloated, introducing a high overhead. Our ultimate goal is to provide hardware-assisted solutions to improve the middleware performance in cloud computing environments. As a starting point, in this paper, we design, implement, and evaluate specialized hardware instructions to accelerate GC operations. We select GC because it is a common component in virtual machine designs and it incurs high performance and energy consumption overheads. We performed a profiling study on various GC algorithms to identify the GC performance hotspots, which contribute to more than 50% of the total GC execution time. By moving these hotspot functions into hardware, we achieved an order of magnitude speedup and significant improvement on energy efficiency. In addition, the results of our performance estimation study indicate that the hardware-assisted GC instructions can reduce the GC execution time by half and lead to a 7% improvement on the overall execution time. Jie Tang 0003, Shaoshan Liu, Zhimin Gu, Xiao-Feng Li, Jean-Luc Gaudiot |
J. Supercomput. | 2 |
| 2011 | Memory-Side Acceleration for XML Parsing
Jie Tang 0003, Shaoshan Liu, Zhimin Gu, Chen Liu 0001, Jean-Luc Gaudiot |
NPC | 2 |
| 2011 | Workload Characterization of Cryptography Algorithms for Hardware AccelerationabstractData encryption/decryption has become an essential component for modern information exchange. However, executing these cryptographic algorithms is often associated with huge overhead and the need to reduce this overhead arises correspondingly. In this paper, we select nine widely adopted cryptography algorithms and study their workload characteristics. Different from many previous works, we consider the overhead not only from the perspective of computation but also focusing on the memory access pattern. We break down the function execution time to identify the software bottleneck suitable for hardware acceleration. Then we categorize the operations needed by these algorithms. In particular, we introduce a concept called 'Load-Store Block' (LSB) and perform LSB identification of various algorithms. Our results illustrate that for cryptographic algorithms, the execution rate of most hotspot functions is more than 60%; memory access instruction ratio is mostly more than 60%; and LSB instructions account for more than 30% for selected benchmarks. Based on our findings, we suggest future directions in designing either the hardware accelerator associated with microprocessor or specific microprocessor for cryptography applications. Jed Kao-Tung Chang, Chen Liu 0001, Shaoshan Liu, Jean-Luc Gaudiot |
ICPE | 3 |
| 2010 | On energy efficiency of reconfigurable systems with run-time partial reconfigurationabstractIn this paper we study whether partial reconfiguration can be used to reduce FPGA energy consumption. In an ideal scenario, we will have a hardware accelerator to assist with certain parts of program execution. When the accelerator is not active, we use partial reconfiguration to unload it to reduce both static and dynamic power. However, the reconfiguration process may introduce a high energy overhead, thus it is unclear whether this approach is feasible. To approach this problem, we identify the conditions under which partial reconfiguration can be used to reduce energy consumption, and we propose solutions to minimize the configuration energy overhead. The results of our study show that by using partial reconfiguration to reduce the power consumption of the accelerator when it is inactive, we can accelerate program execution and at the same time halve the overall energy consumption. Shaoshan Liu, Richard Neil Pittman, Alessandro Form, Jean-Luc Gaudiot |
ASAP | 1 |
| 2010 | Hardware-assisted middleware: Acceleration of garbage collection operationsabstractAlthough the virtualization technology brings many benefits to cloud computing environments, as the virtual machines provide more features, the middleware layer has become bloated, introducing a high overhead. Our ultimate goal is to provide hardware-assisted solutions to improve the middleware performance in cloud computing environments. As a starting point, in this paper, we design, implement, and evaluate specialized hardware instructions to accelerate GC operations. We select GC because it is a common component in virtual machine designs and it incurs high performance and energy consumption overheads. We performed a profiling study on various GC algorithms to identify the GC performance hotspots, which contribute to more than 50% of the total GC execution time. By moving these hotspot functions into hardware, we managed to achieve an order of magnitude speedup. Jie Tang 0003, Shaoshan Liu, Zhimin Gu, Xiao-Feng Li, Jean-Luc Gaudiot |
ASAP | 2 |
| 2010 | Energy reduction with run-time partial reconfiguration (abstract only)abstractWe study whether partial reconfiguration can be used to reduce FPGA energy consumption. In the ideal scenario, we have a hardware accelerator to accelerate certain parts of the program execution. And when the accelerator is not active, we use partial reconfiguration to unload it to reduce both static and dynamic power. However, the reconfiguration process may introduce a high energy overhead, thus it is unclear whether this approach is feasible. To approach this problem, we identify the conditions under which partial reconfiguration can be used to reduce energy consumption, and we propose solutions to minimize the configuration energy overhead. The results of our study show that by using partial reconfiguration to eliminate the power consumption of the accelerator when it is inactive, we can accelerate program execution and at the same time reduce the overall energy consumption by half. Shaoshan Liu, Richard Neil Pittman, Alessandro Forin |
FPGA | 1 |
| 2010 | Minimizing partial reconfiguration overhead with fully streaming DMA engines and intelligent ICAP controller (abstract only)abstractConfiguration overhead seriously limits the usefulness of FPGA partial reconfiguration. In this paper, we propose a combination of two techniques to minimize the partial reconfiguration performance overhead. First, we design and implement fully streaming DMA engines to nearly saturate configuration throughput. Second, we exploit a simple form of configuration data redundancy to compress the configuration bitstreams, and we implement an intelligent ICAP controller to perform decompression at runtime. The results show that our design achieves an effective configuration data transfer throughput of up to 1.2 Gbytes/s, which actually well surpasses the theoretical upper bound of the data transfer throughput, 400 Mbytes/s. Specifically, our fully streaming DMA engines reduce the configuration time from the range of seconds to the range of milliseconds, a more than 1000-fold improvement. In addition, our simple compression scheme achieves up to 75% reduction of bitstream size and results in a decompression circuit with negligible hardware overhead. Shaoshan Liu, Richard Neil Pittman, Alessandro Forin |
FPGA | 1 |
| 2010 | A Theoretical Framework for Value Prediction in Parallel SystemsabstractWe present here a theoretical framework towards a fundamental understanding of the effects of value prediction. Our framework consists of two parts: first, an identification of the theoretical limit of value prediction and an indication of the potential to improve parallelism through the exploitation of value predictability; second, a demonstration of the feasibility of data prediction and a theoretical support to verify this feasibility. The experiment results demonstrate the immense potential of value prediction in enhancing the performance of many-core architectures. Shaoshan Liu, Christine Eisenbeis, Jean-Luc Gaudiot |
ICPP | 1 |
| 2010 | Speculative Execution on GPU: An Exploratory StudyabstractWe explore the possibility of using GPUs for speculative execution: we implement software value prediction techniques to accelerate programs with limited parallelism, and software speculation techniques to accelerate programs that contain runtime parallelism, which are hard to parallelize statically. Our experiment results show that due to the relatively high overhead, mapping software value prediction techniques on existing GPUs may not bring any immediate performance gain. On the other hand, although software speculation techniques introduce some overhead as well, mapping these techniques to existing GPUs can already bring some performance gain over CPU. Shaoshan Liu, Christine Eisenbeis, Jean-Luc Gaudiot |
ICPP | 1 |
| 2010 | Hardware-assisted security mechanism: The acceleration of cryptographic operations with low hardware costabstractThis paper presents generic cryptographic accelerator. Certain "hotspot function" are found in the cryptographic algorithm which consume a substantial amount of execution time of the specific algorithm. INTEL performance analyzer VTune was used which analyzes the software performance on IA-32 and Intel64-based machines to examne the hotsport function. By moving the operations to hardware, we can reduce the overheads introduced by the crypto-computation so that the computing resource can focus on the useful work. Jed Kao-Tung Chang, Shaoshan Liu, Jean-Luc Gaudiot, Chen Liu 0001 |
IPCCC | 2 |
| 2010 | The Performance Analysis and Hardware Acceleration of Crypto-computations for Enhanced SecurityabstractSecurity is very important in modern life due to most information is now stored in digital format. A good security mechanism will keep information secrecy and integrity, hence, plays an important role in modern information exchange. However, cryptography algorithms are extremely expensive in terms of execution time. To make data not easily being cracked, many arithmetic and logical operations will be executed in the encryption/decryption process with many data movement. This means the cryptographic applications are both computation and memory intensive. Using a general-purpose processor for this scenario would not be very cost-effective. This study addresses this problem. Compared to the previous designs, we used a performance analyzer to identify “hotspot” functions across a set of benchmarks. The hotspot function consumes a substantial amount of the execution time of the specific algorithm. Then we translate these hotspot functions into hardware accelerators to improve the performance. Overall we achieve 34 - 83 folds of speedup. Jed Kao-Tung Chang, Shaoshan Liu, Jean-Luc Gaudiot, Chen Liu 0001 |
PRDC | 2 |
| 2009 | Packer: An innovative space-time-efficient parallel garbage collection algorithm based on virtual spacesabstractThe fundamental challenge of garbage collector (GC) design is to maximize the recycled space with minimal time overhead. For efficient memory management, in many GC designs the heap is divided into large object space (LOS) and non-large object space (non-LOS). When one of the spaces is full, garbage collection is triggered even though the other space may still have a lot of free room, thus leading to inefficient space utilization. Also, space partitioning in existing GC designs implies different GC algorithms for different spaces. This not only prolongs the pause time of garbage collection, but also makes collection not efficient on multiple spaces. To address these problems, we propose Packer, a space-and-time-efficient parallel garbage collection algorithm based on the novel concept of virtual spaces. Instead of physically dividing the heap into multiple spaces, Packer manages multiple virtual spaces in one physically shared space. With multiple virtual spaces, Packer offers the advantage of efficient memory management. At the same time, with one physically shared space, Packer avoids the problem of inefficient space utilization. To reduce the garbage collection pause time of Packer, we also propose a novel parallelization method that is applicable to multiple virtual spaces. We reduce the compacting GC parallelization problem into a tree traversal parallelization problem, and apply it to both normal and large object compaction. Shaoshan Liu, Ligang Wang 0001, Xiao-Feng Li, Jean-Luc Gaudiot |
IPDPS | 1 |
| 2009 | Potential Impact of Value Prediction on Communication in Many-Core ArchitecturesabstractThe newly emerging many-core-on-a-chip designs have renewed an intense interest in parallel processing. By applying Amdahl's formulation to the programs in the PARSEC and SPLASH-2 benchmark suites, we find that most applications may not have sufficient parallelism to efficiently utilize modern parallel machines. The long sequential portions in these application programs are caused by computation as well as communication latency. However, value prediction techniques may allow the ldquoparallelizationrdquo of the sequential portion by predicting values before they are produced. In conventional superscalar architectures, the computation latency dominates the sequential sections. Thus, value prediction techniques may be used to predict the computation result before it is produced. In many-core architectures, since the communication latency increases with the number of cores, value prediction techniques may be used to reduce both the communication and computation latency. In this paper, we extend Amdahl's formulation to model the data redundancy inherent to each benchmark, thereby identifying the potential of value prediction techniques. Our analysis shows that the performance of PARSEC benchmarks may improve by a factor of 180 and 230 percent for the SPLASH-2 suite, compared to when only the intrinsic parallelism is considered. This demonstrates the immense potential of fine-grained value prediction in reducing the communication latency in many-core architectures. Shaoshan Liu, Jean-Luc Gaudiot |
IEEE Trans. Computers | 1 |