VLDB 2026 Research / reviewers in the wild / expert
Jie Tang 0003
dblp:181/2702-3
· DBLP profile ↗
40ranked-venue papers
16as first author
23since 2021 · last 2026
0000-0001-8602-7754ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 9 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | R2MOAG: Robust Roadside Monocular 3D Object Detection with Adaptive Token and Ground EmbeddingabstractRoadside cameras effectively enhance the perception capabilities of embodied artificial intelligence systems such as vehicles by compensating for the limitations of vehicle-mounted cameras, which are prone to occlusion and have a limited sensing range, thereby improving the safety of autonomous vehicles. However, existing object detection systems often encounter perception errors when handling comprehensive viewpoint noise in roadside scenes, as well as variations in traffic flow, lighting conditions, and camera poses. This makes it challenging for them to perform robustly in complex road environments. To address these issues, we propose \(\mathrm{R^{2}MOAG}\) , a highly robust monocular 3D object detection method for roadside systems, based on ground perception embedding and heterogeneous visual tokens. The proposed method extracts detailed road information through ground plane equations and utilizes heterogeneous visual tokens to focus on foreground features. By integrating low-dimensional ground information with high-dimensional visual features, the model is provided with clear and rich cues for object detection, significantly enhancing its stability. We conducted extensive experiments on the widely recognized roadside datasets DAIR-V2X-I and Rope3D. The results show that, in terms of overall performance, the proposed model achieved a 4.65% and 4.26% improvement in the \(AP_{3D}|_{R40}\) metric for the vehicle category on these two datasets, respectively. Moreover, the model maintained stable recognition performance across various road scenarios and camera poses, demonstrating exceptional robustness. Jie Tang 0003, Haoran Pan, Bo Yu 0014, Shaoshan Liu |
ACM Trans. Cyber Phys. Syst. | 1 |
| 2026 | PRTF: Polar Space Represented Multi-View 3D Object Detection With Temporal Fusion EnhancementabstractAutonomous driving technology is becoming a significant trend in the development of public transportation. A critical task in autonomous driving perception is 3D object detection, which provides essential data support for downstream applications. Most mainstream 3D object detection methods rely on the Cartesian coordinate system, where they construct object queries to interact with image features and position embedding. However, these methods have the following problems: 1) Sensor-captured detail information diminishes with increasing distance, while pixels represent the same space in Cartesian coordinates, preventing the model from fully leveraging details in closer regions. 2) Multi-view images suffer from spatial misalignment due to overlapping fields of view. 3) The performance of existing single-branch depth prediction networks lacks the necessary accuracy. These issues hinder the feature interaction and affect detection performance. We propose an innovative framework PRTF. Based on Polar space, we design the Two-Stage Transformation Encoder: in the first stage, Dual-DepthNet is used to improve the accuracy of depth prediction. In the second stage, Polar points are generated to address spatial misalignment, enabling effective encoding of details at close distance. In the Temporal Decoder, object queries are leveraged to integrate temporal information, effectively compensating for ambiguous information. By enhancing spatial information at both near and far distances in Polar space, the overall performance of multi-view 3D object detection is significantly improved. PRTF achieves state-of-the-art performance on nuScenes Test with 56.1% mAP and 63.9% NDS, exceeding multi-modal frameworks that combine image and radar data. Jie Tang 0003, Yefei Hou, Bo Yu 0014 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | True Match: Leveraging 2D-Assisted Queries for Multi-view 3D Detection in Polar SpaceabstractSparse Query-Based paradigms for multi-view 3D object detection achieved remarkable success, but the redundant predictions still affects the detection accuracy and object localization because of the following shortcomings: 1) 3D position embedding exhibits weak spatial perception ability and cannot capture subtle differences between similar objects. 2) Randomly initialized object queries lack prior knowledge of reference points and object information, which hinders their accuracy in matching with corresponding objects. 3) The temporal fusion process fails to effectively capture object details, making it challenging to localize object accurately. To address these issues, we propose an innovative framework T-Match, which leverages prior knowledge from a 2D detector to initialize object queries. Based on Polar space, both 2D and 3D information are integrated to comprehensively capture object details, object queries iteratively updates through Cross-Domain Spatio-Temporal Attention, which incorporates cross-domain object information, and Polar-Aware Cross Attention, which aggregates image features fused with Polar position embedding, refining matching results while reducing redundant predictions. T-Match achieves state-of-the-art performance on nuScenes Test with 57.5% mAP and 64.9% NDS, exceeding multi-modal frameworks that combine image and radar data. Yefei Hou, Jie Tang 0003 |
ICME | 2 |
| 2025 | EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and RetrievalabstractObject-goal navigation (ObjNav) tasks an agent with navigating to the location of a specific object in an unseen environment.
Embodied agents equipped with large language models (LLMs) and online constructed navigation maps can perform ObjNav in a zero-shot manner. However, existing agents heavily rely on giant LLMs on the cloud, e.g., GPT-4, while directly switching to small LLMs, e.g., LLaMA3.2-11b, suffer from significant success rate drops due to limited model capacity for understanding complex navigation maps, which prevents deploying ObjNav on local devices.
At the same time, the long prompt introduced by the navigation map description will cause high planning latency on local devices.
In this paper, we propose EfficientNav to enable on-device efficient LLM-based zero-shot ObjNav. To help the smaller LLMs better understand the environment, we propose semantics-aware memory retrieval to prune redundant information in navigation maps.
To reduce planning latency, we propose discrete memory caching and attention-based memory clustering to efficiently save and re-use the KV cache.
Extensive experimental results demonstrate that EfficientNav
achieves 11.1\% improvement in success rate on HM3D benchmark over GPT-4-based baselines,
and demonstrates 6.7$\times$ real-time latency reduction and 4.7$\times$ end-to-end latency reduction over GPT-4 planner. Our code is available on https://github.com/PKU-SEC-Lab/EfficientNav. Sunjian Zheng, Tong Xie, Tianshi Xu, Bo Yu 0014, Fan Wang 0021, Jie Tang 0003, Shaoshan Liu |
NeurIPS | 7 |
| 2025 | Elevation-Aware Map Matching Model Leveraging Transfer Learning in Sparse Data ConditionsabstractMap matching is a pivotal component of intelligent urban transportation, offering foundational data for technologies such as path planning, traffic analysis, and trajectory analysis. Diverging from conventional rule-based and topological map matching algorithms, we approach the map matching task from a data-driven perspective, presenting an Elevation-Aware Map Matching Model under conditions of sparse data. This paper initiates from the vehicular standpoint, constructing an Elevation-Aware Unit utilizing imagery and sensor data to acquire elevation information for diverse urban roads. Subsequently, this unit is integrated into the map matching model, enhancing the model’s resilience to noise. Concurrently, employing a Fine-tuning transfer learning approach, we formulate a cross-domain map matching model to maximize the reduction of model development costs. The model undergoes testing on real-world datasets, employing four metrics for evaluation. The results indicate the superiority of this map matching model over existing counterparts, particularly in intricate urban road scenarios where the model exhibits outstanding performance. Additionally, we validate the effectiveness of the Elevation-Aware Unit, underscoring the significance of height information for map matching models. Jie Tang 0003, Sunjian Zheng, Bo Yu 0014, Xue (Steve) Liu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | 3-D Line Matching Network Based on Matching Existence Guidance and Knowledge DistillationabstractIn applications, such as scene reconstruction and odometry, accurate matching associations for 3-D lines are crucial. Real-world scenes introduce inconsistencies due to variations in perspective, leading to nonoverlapping data acting as noise. Accurately matching partially overlapping sets of 3-D lines becomes challenging, potentially resulting in failed scene reconstruction and erroneous positioning. Prior approaches relied on the traditional iterative closest line (ICL) methods, involving iterative calculations and sensitivity to initial poses, and were prone to matching failures in low-overlap rate data and singular pattern scenes. Existing 3-D line matching networks either did not consider the noise in 3-D line collections or failed to retain more valid matching pairs, while these models often require a larger number of parameters and inference time. To address these issues, this article proposes matching existence guidance module (MEG)-Net, a Plücker line matching network guided by the existence of matches. It leverages the rich geometric characteristics of 3-D lines represented as Plücker lines, enhancing feature robustness. By guiding the model to handle the noisy data through match existence guidance, it improves the model’s performance on partially overlapping 3-D line data. Experiments on the indoor and outdoor data sets and the Out of Distribution (OOD) data sets demonstrate that the MEG-Net outperforms traditional methods and baseline models in 3-D line matching, with better scalability and noise robustness, achieving state-of-the-art results. Additionally, we propose an innovative knowledge distillation method based on the matching matrices, training a more efficient MEG-Net mini student model with approximately 70% fewer parameters and multiply accumulate operations (MACs), while maintaining superior performance and faster inference speeds on the indoor data sets. Jie Tang 0003, Bo Yu 0014, Xue (Steve) Liu |
IEEE Internet Things J. | 1 |
| 2023 | Invited: Autonomous Driving Digital Twin Empowered Design Automation: An Industry PerspectiveabstractDesigning reliable computing systems for autonomous driving is extremely challenging, as the performance and reliability of the systems have to be thoroughly evaluated under an extremely large amount of driving scenarios. Physically constructing scenarios and conducting testing for autonomous driving systems is time-consuming and expensive. To minimize the need for physical testing and improve development efficiency, we developed a digital-twin-based simulation, which can generate an integral, precise, and comprehensive representation of physical scenarios. In this paper, we share our experiences with the digital-twin-based simulation for autonomous driving, particularly the design requirements and components demanded to facilitate virtual environment construction and design verification, which could greatly improve development efficiency. Bo Yu 0014, Jie Tang 0003, Shaoshan Liu |
DAC | 2 |
| 2023 | HAU$\mathbf {M^3}$: A Height Aware Urban Map Matching Mechanism
Jie Tang 0003, Sunjian Zheng, Bo Yu 0014, Shaoshan Liu |
MobiQuitous (1) | 1 |
| 2023 | Data Fusion in Infrastructure-Augmented Autonomous Driving System: Why? Where? and How?abstractThis article is the first to provide a thorough system design overview along with the fusion methods selection criteria of a real-world cooperative autonomous driving system enabled by the Internet of Things (IoT), named infrastructure-augmented autonomous driving (IAAD). We present an in-depth introduction to the IAAD hardware and software on both road side and vehicle side. We extensively characterize the IAAD system and observe that the network condition fluctuation along the road is the main roadblock for cooperative autonomous driving. To address this challenge, we propose new fusion methods, dubbed “interframe fusion” and “planning fusion” to complement the state-of-the-art “intraframe fusion.” We demonstrate that each fusion method has its own benefit and constraint. In order to select the best fusion method under varying network conditions, we propose “fusion criteria” to instruct the IAAD system to intelligently make the selection and implement a system framework named adaptive spatial-temporal (S–T) choice to realize the adaptive fusion guided by the “fusion criteria.” Our real-world field data verifies that S–T choice has significantly improved autonomous driving’s safety and reliability by decreasing the fusion miss ratio from 30% to 7% and remain the planning displacement error within the 1.7 m instead of 4 m when the network condition exacerbates. Bo Yu 0014, Jie Tang 0003, Shuaiwen Song, Cong Liu 0005, Yang Hu 0001 |
IEEE Internet Things J. | 4 |
| 2023 | An Energy Efficient and Runtime Reconfigurable Accelerator for Robotic LocalizationabstractAccurate and efficient localization of robots under limited on-board resources has fueled specialized localization accelerators. Despite many recent efforts, accelerating robotic localization is still fundamentally challenging. To tackle the challenges, the paper proposes a configurable hardware architecture and a design space optimization method to automatically generate an optimal accelerator design under the design constraints. Data locality, sparsity, and fixed-point arithmetic optimization techniques that are specific to the localization algorithm are exploited to customize the accelerator. In addition, a low-cost runtime configuration mechanism is proposed to enable the accelerator to continuously optimize itself at runtime according to the operating environment to save power while sustaining performance and accuracy. The evaluation on FPGA demonstrates that the proposed accelerator achieves orders of magnitude performance improvement and/or energy savings compared to the software implementation on Intel and Arm CPUs; and substantially outperforms existing FPGA accelerators in terms of performance and energy. Qiang Liu 0011, Yuhui Hao, Weizhuang Liu, Bo Yu 0014, Yiming Gan, Jie Tang 0003, Shaoshan Liu, Yuhao Zhu 0001 |
IEEE Trans. Computers | 6 |
| 2022 | TransMigrator: A Transformer-Based Predictive Page Migration Mechanism for Heterogeneous Memory
Songwen Pei, Yihuan Qian, Jie Tang 0003, Jean-Luc Gaudiot |
NPC | 4 |
| 2022 | Brief Industry Paper: The Necessity of Adaptive Data Fusion in Infrastructure-Augmented Autonomous Driving SystemabstractThis paper is the first to provide a thorough system design overview along with the fusion methods selection criteria of a real-world cooperative autonomous driving system, named Infrastructure-Augmented Autonomous Driving or IAAD. We present an in-depth introduction of the IAAD hardware and software on both road-side and vehicle-side computing/communication platforms. We extensively characterize the IAAD system in the context of real-world deployment scenarios and observe that the network condition fluctuates along the road is currently the main technical roadblock for cooperative autonomous driving. To address this challenge, we propose new fusion methods, dubbed “inter-frame fusion” and “planning fusion” to complement the current state-of-the-art “intra-frame fusion”. We demonstrate that each fusion method has its own benefit and constraint. Adaptively choosing the fusion method according to the real-world condition will benefit the SoV without the violation of the SoV's safety requirements. Shaoshan Liu, Bo Yu 0014, Jie Tang 0003, Shuaiwen Song, Cong Liu 0005, Yang Hu 0001 |
RTAS | 7 |
| 2022 | Rise of the Automotive Health-Domain Controllers: Empowering Healthcare Services in Intelligent VehiclesabstractWe are facing a global healthcare crisis today as the healthcare cost is ever climbing, but with the aging population, government fiscal revenue is ever dropping. To address this imminent problem, we can start by enabling affordable anywhere anytime healthcare access through delivering healthcare services on intelligent vehicles. The foundation upon which healthcare services can be provided on intelligent vehicles is an automotive health-domain controller (AHDC), which is missing today. In this article, for the first time, we explain the necessity and define the functionalities of AHDCs. In addition, we delve into the technical challenges, requirements, and feasibility of integrating multiple key features into AHDCs. It is hoped that this will help the community standardize on-vehicle healthcare services provisioning, and lead to the universal adoption of the revolutionary mobile healthcare system. Shaoshan Liu, Yuzhang Huang, Ao Kong, Jie Tang 0003, Xue (Steve) Liu |
IEEE Internet Things J. | 4 |
| 2021 | On Designing Computing Systems for Autonomous Vehicles: a PerceptIn Case StudyabstractPerceptIn develops and commercializes autonomous vehicles for micromobility around the globe. This paper makes a holistic summary of PerceptIn's development and operating experiences. It provides the business tale behind our product, and presents the development of the computing system for our vehicles. We illustrate the design decision made for the computing system, and show the advantage of offloading localization workloads onto an FPGA platform. Bo Yu 0014, Jie Tang 0003, Shaoshan Liu |
ASP-DAC | 2 |
| 2021 | Streaming Data Priority Scheduling Framework for Autonomous Driving by EdgeabstractIn recent years, intelligent vehicles like autonomous vehicles generate a huge amount of sensing data continuously. The computations on those data streams are far beyond the processing capacity of on-board computing. To deal with the streaming data process in real-time, the deployment of streaming data processing system by edge turns to the first choice in terms of performance. However, the existing frameworks cannot satisfy the complicated demands from autonomous driving tasks and lack the ability in supporting the task priority scheduling. In this paper, we propose a streaming data priority scheduling framework for autonomous driving by edge on Spark Streaming and make an implementation on Spark 2.3.0. The proposed framework can identify the priorities among different data processing tasks and implement the task scheduling based on non-preemptive priority queuing theory. To meet differentiated service level requirements, the proposed non-preemptive priority queuing scheduling mechanism considers the priority category of tasks, the distance between vehicles and edge nodes, and the priority weight of vehicles. Experiments show that this mechanism can effectively identify the priority information of different tasks from different vehicles and reduce the end-to-end latency of high-priority tasks by up to 46% than low-priority tasks. Lingbing Yao, Hang Zhao 0016, Jie Tang 0003, Shaoshan Liu, Jean-Luc Gaudiot |
COMPSAC | 3 |
| 2021 | Invited: Towards Fully Intelligent Transportation through Infrastructure-Vehicle Cooperative Autonomous Driving: Challenges and OpportunitiesabstractThe infrastructure-vehicle cooperative autonomous driving approach relies on the cooperation between intelligent roads and intelligent vehicles. This approach is not only safer but also more economical compared to the traditional on-vehicle-only autonomous driving. In this paper, we introduce the real-world deployment experiences of infrastructure-vehicle cooperative autonomous driving by PerceptIn, where a three-stage development roadmap is taken: infrastructure-augmented autonomous driving (IAAD), infrastructure-guided autonomous driving (IGAD), and infrastructure-planned autonomous driving (IPAD). We then discuss the future research challenges and opportunities for such approach. Shaoshan Liu, Bo Yu 0014, Jie Tang 0003, Qi Zhu 0002 |
DAC | 3 |
| 2021 | Eudoxus: Characterizing and Accelerating Localization in Autonomous Machines Industry Track PaperabstractWe develop and commercialize autonomous machines, such as logistic robots and self-driving cars, around the globe. A critical challenge to our—and any—autonomous machine is accurate and efficient localization under resource constraints, which has fueled specialized localization accelerators recently. Prior acceleration efforts are point solutions in that they each specialize for a specific localization algorithm. In real-world commercial deployments, however, autonomous machines routinely operate under different environments and no single localization algorithm fits all the environments. Simply stacking together point solutions not only leads to cost and power budget overrun, but also results in an overly complicated software stack. This paper demonstrates our new software-hardware co-designed framework for autonomous machine localization, which adapts to different operating scenarios by fusing fundamental algorithmic primitives. Through characterizing the software framework, we identify ideal acceleration candidates that contribute significantly to the end-to-end latency and/or latency variation. We show how to co-design a hardware accelerator to systematically exploit the parallelisms, locality, and common building blocks inherent in the localization framework. We build, deploy, and evaluate an FPGA prototype on our next-generation self-driving cars. To demonstrate the flexibility of our framework, we also instantiate another FPGA prototype targeting drones, which represent mobile autonomous machines. We achieve about $2 \times$ speedup and $4 \times$ energy reduction compared to widely-deployed, optimized implementations on general-purpose platforms. Yiming Gan, Bo Yu 0014, Boyuan Tian, Leimeng Xu, Shaoshan Liu, Qiang Liu 0011, Jie Tang 0003, Yuhao Zhu 0001 |
HPCA | 9 |
| 2021 | Archytas: A Framework for Synthesizing and Dynamically Optimizing Accelerators for Robotic LocalizationabstractDespite many recent efforts, accelerating robotic computing is still fundamentally challenging for two reasons. First, robotics software stack is extremely complicated. Manually designing an accelerator while meeting the latency, power, and resource specifications is unscalable. Second, the environment in which an autonomous machine operates constantly changes; a static accelerator design leads to wasteful computation. Weizhuang Liu, Bo Yu 0014, Yiming Gan, Qiang Liu 0011, Jie Tang 0003, Shaoshan Liu, Yuhao Zhu 0001 |
MICRO | 5 |
| 2021 | A KNN Query Method for Autonomous Driving Sensor Data
Jie Tang 0003, Jiehui Zhang, Zhixin Zeng, Shaoshan Liu |
NPC | 1 |
| 2021 | Brief Industry Paper: An Edge-Based High-Definition Map Crowdsourcing Task Distribution Framework for Autonomous DrivingabstractFacing the difficulty and inefficiency of creating and maintaining High-Definition (HD) maps in our commercial deployments, we have developed an edge-based crowdsourcing task distribution framework for HD Map in autonomous driving. Our key observation is that: HD map data crowdsourcing exhibits the diminishing marginal utility thus there exists an inflection point for maximum utility, meanwhile its premature convergence of utility will leave some map updates not notified in time. Based on this observation, we develop a periodic crowdsourcing task distribution framework. It discretizes the demands for collecting source data into different periods and uses an optimal stopping rule to terminate the data collection for the maximum crowdsourcing utility. The experimental results verify that our crowdsourcing framework can achieve high time coverage and high efficiency with lower cost. Donghua Li, Jie Tang 0003, Shaoshan Liu |
RTAS | 2 |
| 2021 | Brief Industry Paper: The Matter of Time - A General and Efficient System for Precise Sensor Synchronization in Robotic ComputingabstractTime synchronization is a critical task in robotic computing such as autonomous driving. In the past few years, as we developed advanced robotic applications, our synchronization system has evolved as well. In this paper, we first introduce the time synchronization problem and explain the challenges of time synchronization, especially in robotic workloads. Summarizing these challenges, we then present a general hardware synchronization system for robotic computing, which delivers high synchronization accuracy while maintaining low energy and resource consumption. The proposed hardware synchronization system is a key building block in our future robotic products. Shaoshan Liu, Bo Yu 0014, Kunai Zhang, Yisong Qiao, Thomas Yuang Li, Jie Tang 0003, Yuhao Zhu 0001 |
RTAS | 7 |
| 2021 | Brief Industry Paper: An Infrastructure-Aided High Definition Map Data Provisioning Service for Autonomous DrivingabstractAs a fundamental component in the autonomous driving technology stack, High Definition Maps (HD map) provide high-precision descriptions of the environment. It enables extremely accurate perception and localization while improving the efficiency of path planning. However, the HD map's extremely large data volume poses great challenges for the real-time and safety requirements of autonomous driving. Based on our real-world deployment experiences, we first demonstrate how the existing data transmission mechanism is weak in supporting HD map services. To address this problem, we propose an HD map data service mechanism on top of Vehicle-to-Infrastructure (V2I) data transmission under a tight time and energy budget. By this mechanism, the selected road side unit (RSU) nodes cooperate on map provisioning tasks and transmit HD map data proportionately. Furthermore, we model the real-time map data service into a partial knapsack problem and develop a greedy data transmission algorithm. Experimental results confirm that the proposed mechanism can ensure the real-time HD map data service meanwhile meeting the energy limits. Jinliang Xie, Jie Tang 0003, Yanzhi Wang 0001, Qi Zhu 0002, Shaoshan Liu |
RTAS | 2 |
| 2021 | An edge streaming data processing framework for autonomous drivingabstractIn recent years, with the rapid development of sensing technology and the Internet of Things (IoT), sensors play increasingly important roles in traffic control, medical monitoring, industrial production and etc. They generated high volume of data in a streaming way that often need to be processed in real time. Therefore, streaming data computing technology plays an indispensable role in the real-time processing of sensor data in high throughput but low latency. However, there are two problems in deploying streaming data process ability in cloud computing data centre. Firstly, massive sensor nodes simultaneously upload data to the remote cloud computing data centre, which requires a large number of bandwidth resources supports. The existing network infrastructure cannot provide enough bandwidth at a reasonable price. Secondly, due to the geographical distribution characteristics of the cloud computing data centre, there will inevitably be large transmission delay during the process of data transmission. Such end-to-end delay is intolerable to mobile applications especially for those latency sensitive tasks. In view of the above problems, this paper proposes an autonomous driving oriented edge streaming data processing framework, which migrates the computing and storage capability from the remote cloud data centre to the edge data centre. It focuses on the change of vehicle flow in a specific geographical area, and uses the computing power sunk to edge node to process the massive streaming data generated by autonomous vehicles nearby. The proposed framework is implemented on top of Spark Streaming, which builds up a gray model based traffic flow monitor, a traffic prediction orientated prediction layer and a fuzzy control based Batch Interval dynamic adjustment layer for Spark Streaming. It could forecast the variation of sensors data arrive rate, make streaming Batch Interval adjustment in advance and implement real-time streaming process by edge. Therefore, it can realise the monitor and prediction of the data flow changes of the autonomous driving vehicle sensor data in geographical coverage of edge computing node area, meanwhile minimise the end-to-end latency but satisfy the application throughput requirements. The experiments show that it can predict short-term traffic with no more than 4% relative error in a whole day. By making batch consuming rate close to data generating rate, it can maintain system stability well even when arrival data rate changes rapidly. The Batch Interval can be converged to a suitable value in two minutes when data arrival rate is doubled. Compared with vanilla version Spark Streaming, where there has serious task accumulation and introduces large delay, it can reduce 35% latency by squeezing Batch Interval when data arrival rate is low; it also can significantly improve system throughput by only at most 25% Batch Interval increase when data arrival rate is high. Hang Zhao 0016, Linbin Yao, Zhixin Zeng, Donghua Li, Jinliang Xie, Weiling Zhu, Jie Tang 0003 |
Connect. Sci. | 7 |
| 2020 | π-Map: A Decision-Based Sensor Fusion with Global Optimization for Indoor MappingabstractIn this paper, we propose π-map, a tightly coupled fusion mechanism that dynamically consumes LiDAR and sonar data to generate reliable and scalable indoor maps for autonomous robot navigation. The key novelty of π-map over previous attempts is the utilization of a fusion mechanism that works in three stages: the first LiDAR scan matching stage efficiently generates initial key localization poses; the second optimization stage is used to eliminate errors accumulated from the previous stage and guarantees that accurate large-scale maps can be generated; then the final revisit scan fusion stage effectively fuses the LiDAR map and the sonar map to generate a highly accurate representation of the indoor environment. We evaluate π-map on both large and small environments and verify its superiority over existing fusion methods. Zhiliu Yang, Bo Yu 0014, Jie Tang 0003, Shaoshan Liu, Chen Liu 0001 |
IROS | 4 |
| 2020 | Building the Computing System for Autonomous Micromobility Vehicles: Design Constraints and Architectural OptimizationsabstractThis paper presents the computing system design in our commercial autonomous vehicles, and provides a detailed performance, energy, and cost analyses. Drawing from our commercial deployment experience, this paper has two objectives. First, we highlight design constraints unique to autonomous vehicles that might change the way we approach existing architecture problems. Second, we identify new architecture and systems problems that are perhaps less studied before but are critical to autonomous vehicles. Bo Yu 0014, Leimeng Xu, Jie Tang 0003, Shaoshan Liu, Yuhao Zhu 0001 |
MICRO | 4 |
| 2020 | π-Hub: Large-scale video learning, storage, and retrieval on heterogeneous hardware platforms
Jie Tang 0003, Shaoshan Liu, Jie Cao 0003, Bolin Ding, Jean-Luc Gaudiot, Weisong Shi |
Future Gener. Comput. Syst. | 1 |
| 2020 | $\pi$π-BA: Bundle Adjustment Hardware Accelerator Based on Distribution of 3D-Point ObservationsabstractBundle adjustment (BA) is a fundamental optimization technique used in many crucial applications, including 3D scene reconstruction, robotic localization, camera calibration, autonomous driving, street view map generation, and even space exploration etc. Essentially, BA is a joint non-linear optimization problem, and one which can consume a significant amount of time and power, especially for large optimization problems. Previous approaches of optimizing BA performance heavily rely on parallel processing or distributed computing, which trade higher power consumption for higher performance. In this article we propose p-BA, the first hardware-software co-designed BA hardware accelerator that exploits custom hardware to simultaneously achieve higher performance and power efficiency. Specifically, based on our key observation that not all 3D points appear on all images in a BA problem, we designed a Co-Observation Optimization technique to accelerate BA operations with optimized usage of memory and computation resources. In addition, we developed a hardware-friendly differentiation method, which combines the analytic and forward automatic differentiation to calculate derivatives of projection function in the BA problem. We have implemented the proposed design on an embedded FPGA SoC, and experimental results confirm that p-BA outperforms the existing software implementations in terms of performance and power consumption. Qiang Liu 0011, Shuzhen Qin, Bo Yu 0014, Jie Tang 0003, Shaoshan Liu |
IEEE Trans. Computers | 4 |
| 2019 | A DAG Refactor Based Automatic Execution Optimization Mechanism for Spark
Hang Zhao 0016, Donghua Li, Jie Tang 0003, Shaoshan Liu |
NPC | 4 |
| 2019 | Edge Computing for Autonomous Driving: Opportunities and ChallengesabstractSafety is the most important requirement for autonomous vehicles; hence, the ultimate challenge of designing an edge computing ecosystem for autonomous vehicles is to deliver enough computing power, redundancy, and security so as to guarantee the safety of autonomous vehicles. Specifically, autonomous driving systems are extremely complex; they tightly integrate many technologies, including sensing, localization, perception, decision making, as well as the smooth interactions with cloud platforms for high-definition (HD) map generation and data storage. These complexities impose numerous challenges for the design of autonomous driving edge computing systems. First, edge computing systems for autonomous driving need to process an enormous amount of data in real time, and often the incoming data from different sensors are highly heterogeneous. Since autonomous driving edge computing systems are mobile, they often have very strict energy consumption restrictions. Thus, it is imperative to deliver sufficient computing power with reasonable energy consumption, to guarantee the safety of autonomous vehicles, even at high speed. Second, in addition to the edge system design, vehicle-to-everything (V2X) provides redundancy for autonomous driving workloads and alleviates stringent performance and energy constraints on the edge side. With V2X, more research is required to define how vehicles cooperate with each other and the infrastructure. Last, safety cannot be guaranteed when security is compromised. Thus, protecting autonomous driving edge computing systems against attacks at different layers of the sensing and computing stack is of paramount concern. In this paper, we review state-of-the-art approaches in these areas as well as explore potential solutions to address these challenges. Shaoshan Liu, Liangkai Liu, Jie Tang 0003, Bo Yu 0014, Yifan Wang 0005, Weisong Shi |
Proc. IEEE | 3 |
| 2018 | Teaching Autonomous Driving Using a Modular and Integrated ApproachabstractIntroduction: Teaching autonomous driving is a challenging task. Indeed, most existing autonomous driving teaching activities focus on a few of the technologies involved. This not only fails to provide a comprehensive coverage, but also sets a high entry barrier for students with different backgrounds. Objective: The primary objective of this study is to present a modular, integrated approach towards teaching autonomous driving. Methods: We organize the technologies used in autonomous driving into modules. This is described in the textbook we have developed as well as a series of multimedia online lectures designed to provide technical overview for each module. Once the students have understood these modules, the experimental platforms for integration we have developed allow the students to fully understand how the modules interact with each other. Results: To verify this teaching approach, we present three case studies: an introductory class on autonomous driving for students with only a basic technology background; a new session in an existing embedded systems class to demonstrate how embedded system technologies can be applied towards autonomous driving; and an industry professional training session to quickly bring up experienced engineers to work in autonomous driving. The results show that students can maintain a high interest level and make great progress by starting with familiar concepts before moving onto other modules. Conclusions: Autonomous driving is not one single technology, but rather a complex system integrating many technologies. Our modular and integrated approach is an effective method in teaching autonomous driving. Jie Tang 0003, Shaoshan Liu, Songwen Pei, Stéphane Zuckerman, Chen Liu 0001, Weisong Shi, Jean-Luc Gaudiot |
COMPSAC (1) | 1 |
| 2018 | π-SoC: Heterogeneous SoC Architecture for Visual Inertial SLAM ApplicationsabstractIn recent years, we have observed a clear trend in the rapid rise of autonomous vehicles and robotics. One of the core technologies enabling these applications, Simultaneous Localization And Mapping (SLAM), imposes two main challenges: first, these workloads are computationally intensive and they often have real-time requirements; second, these workloads run on battery-powered mobile devices with limited energy budget. Hence, performance should be improved while simultaneously reducing energy consumption, two rather contradicting goals by conventional wisdom. Previous attempts to optimize SLAM performance and energy efficiency usually involve optimizing one function and fail to approach the problem systematically. In this paper, we first study the characteristics of visual inertial SLAM workloads on existing heterogeneous SoCs. Then based on the initial findings, we propose π-SoC, a heterogeneous SoC design that systematically optimize the IO interface, the memory hierarchy, as well as the the hardware accelerator. We implemented this system on a Xilinx Zynq UltraScale MPSoC and was able to deliver over 60 FPS performance with average power less than 5 W. Jie Tang 0003, Bo Yu 0014, Shaoshan Liu, Zhe Zhang 0006, Weikang Fang |
IROS | 1 |
| 2017 | CSAS: Cost-Based Storage Auto-Selection, a Fine Grained Storage Selection Mechanism for Spark
Bo Wang 0139, Jie Tang 0003, Zhimin Gu |
NPC | 2 |
| 2015 | How can Garbage Collection be energy efficient by dynamic offloading?abstractGarbage Collection (GC) is still a major issue in JVM for both mobile and cluster computing. GC offloading is proposed to improve the performance of GC by delivering part or all of the operations into another dedicated GC hardware. However, the traditional offloading just offloads directly not considering the phase change of GC behavior, which can be classified into two different groups: minor GC and major GC. The minor GC is fast and frequently invoked, while major GC is expensive in terms of time but seldom takes place. The direct offloading made GC workload frequently hopping between main processor and GC hardware, introduced a noticeable overhead and offset any possible benefits of workload loading. To solve this issue, we propose to offload GC dynamically by a careful selection of profitable and harmful GC operations. We also made a case study on Apache Spark, a lightning-fast cluster computing platform. It shows dynamic offloading can yield nearly 42.6% performance improvement with a concurrent 32.1% in energy cost reduction. Jie Tang 0003, Chen Liu 0001, Jean-Luc Gaudiot |
ASAP | 1 |
| 2013 | OCP: Offload Co-Processor for energy efficiency in embedded mobile systemsabstractIn current embedded mobile systems design, the application processor (AP) is often woken up to service interrupts and user requests. However, this kind of wakeups from sleep is very expensive in terms of battery usage. In the observation that the operating system/driver workloads are very light-weight, in this paper we propose the Offload Co-Processor (OCP) SoC architecture. In the OCP SoC design, when the device is idle, we offload the operating system workloads (mainly interrupt handling workloads) to an ultra-low-power coprocessor. This way, the co-processor would be able to handle most wake-up requests without awakening the heavy-weight AP, thus avoiding the overhead of AP spin-up/down. Using GPS continuous sampling workload as a case study, we show that the proposed OCP SoC design would extend battery life by 3.5 folds. Jie Tang 0003, Chen Liu 0001, Yu-Liang Chou, Shaoshan Liu |
ASAP | 1 |
| 2013 | Pinned OS/Services: A Case Study of XML Parsing on Intel SCC
Jie Tang 0003, Pollawat Thanarungroj, Chen Liu 0001, Shaoshan Liu, Zhimin Gu, Jean-Luc Gaudiot |
J. Comput. Sci. Technol. | 1 |
| 2013 | Acceleration of XML Parsing through PrefetchingabstractExtensible Markup Language (XML) has become a widely adopted standard for data representation and exchange. However, its features also introduce significant overhead threatening the performance of modern applications. In this paper, we present a study of XML parsing and determine that memory-side data loading in the parsing stage incurs a significant performance overhead, as much as the computation does. Hence, we propose memory-side acceleration which incorporates of data prefetching techniques, and can be applied on top of computation-side acceleration to speed up the XML data parsing. To this end, we study here the impact of our proposed scheme on the performance and energy consumption and demonstrated how it is capable of improving performance by up to 20 percent as well as produce up to 12.77 percent of energy saving when implemented in 32-nm technology. In addition, we implement a prefetcher on an platform in an effort to evaluate its implementation feasibility in terms of area and energy overhead. Jie Tang 0003, Shaoshan Liu, Chen Liu 0001, Zhimin Gu, Jean-Luc Gaudiot |
IEEE Trans. Computers | 1 |
| 2012 | Packer: Parallel Garbage Collection Based on Virtual SpacesabstractThe fundamental challenge of garbage collector (GC) design is to maximize the recycled space with minimal time overhead. For efficient memory management, in many GC designs the heap is divided into large object space (LOS) and normal object space (non-LOS). When either space is full, garbage collection is triggered even though the other space may still have plenty of room, thus leading to inefficient space utilization. Also, space partitioning in existing GC designs implies different GC algorithms for different spaces. This not only prolongs the pause time of garbage collection, but also makes collection inefficient on multiple spaces. To address these problems, we propose Packer, a parallel garbage collection algorithm based on the novel concept of virtual spaces. Instead of physically dividing the heap into multiple spaces, Packer manages multiple virtual spaces in one physical space. With multiple virtual spaces, Packer offers efficient memory management. With one physical space, Packer avoids the problem of an inefficient space utilization. To reduce the garbage collection pause time, we also propose a novel parallelization method that is applicable to multiple virtual spaces. Specifically, we reduce the compacting GC parallelization problem into a discreted acyclic graph (DAG) traversal parallelization problem, and apply it to both normal and large object compaction. Shaoshan Liu, Jie Tang 0003, Ligang Wang 0001, Xiao-Feng Li, Jean-Luc Gaudiot |
IEEE Trans. Computers | 2 |
| 2012 | Achieving middleware execution efficiency: hardware-assisted garbage collection operationsabstractAlthough virtualization technologies bring many benefits to cloud computing environments, as the virtual machines provide more features, the middleware layer has become bloated, introducing a high overhead. Our ultimate goal is to provide hardware-assisted solutions to improve the middleware performance in cloud computing environments. As a starting point, in this paper, we design, implement, and evaluate specialized hardware instructions to accelerate GC operations. We select GC because it is a common component in virtual machine designs and it incurs high performance and energy consumption overheads. We performed a profiling study on various GC algorithms to identify the GC performance hotspots, which contribute to more than 50% of the total GC execution time. By moving these hotspot functions into hardware, we achieved an order of magnitude speedup and significant improvement on energy efficiency. In addition, the results of our performance estimation study indicate that the hardware-assisted GC instructions can reduce the GC execution time by half and lead to a 7% improvement on the overall execution time. Jie Tang 0003, Shaoshan Liu, Zhimin Gu, Xiao-Feng Li, Jean-Luc Gaudiot |
J. Supercomput. | 1 |
| 2011 | Memory-Side Acceleration for XML Parsing
Jie Tang 0003, Shaoshan Liu, Zhimin Gu, Chen Liu 0001, Jean-Luc Gaudiot |
NPC | 1 |
| 2010 | Hardware-assisted middleware: Acceleration of garbage collection operationsabstractAlthough the virtualization technology brings many benefits to cloud computing environments, as the virtual machines provide more features, the middleware layer has become bloated, introducing a high overhead. Our ultimate goal is to provide hardware-assisted solutions to improve the middleware performance in cloud computing environments. As a starting point, in this paper, we design, implement, and evaluate specialized hardware instructions to accelerate GC operations. We select GC because it is a common component in virtual machine designs and it incurs high performance and energy consumption overheads. We performed a profiling study on various GC algorithms to identify the GC performance hotspots, which contribute to more than 50% of the total GC execution time. By moving these hotspot functions into hardware, we managed to achieve an order of magnitude speedup. Jie Tang 0003, Shaoshan Liu, Zhimin Gu, Xiao-Feng Li, Jean-Luc Gaudiot |
ASAP | 1 |