EDBT 2026 Demo / reviewers in the wild / expert
Yunfei Xiang
dblp:240/0099
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | IDEA-GP: Instruction-Driven Architecture with Efficient Online Workload Allocation for Geometric PerceptionabstractThe algorithmic complexity of robotic systems presents significant challenges to achieving generalized acceleration in robot applications.On the one hand, the diversity of operators and computational flows within similar task categories prevents the reuse of specialized computational units.On the other hand, task variations and environmental dynamics can cause workload fluctuations, leading to inefficient resource utilization.This paper focuses on the geometric perception capability of robots, taking localization and mapping as the basic applications, and proposes IDEA-GP, an Instruction-Driven Architecture with Efficient online workload Allocation for Geometric Perception.Built around an array of general computational units designed for spatial positioning representations, IDEA-GP supports a wide range of robot pose-related computational tasks.IDEA-GP employs a compiler to perform online workload analysis and resource allocation.It generates instructions tailored to processing elements (PEs) to schedule computations, thereby accelerating optimization problems and enhancing geometric perception performance.Deployed on the ZCU102 evaluation board, IDEA-GP demonstrates an average speedup of 7.5× over the Intel CPU and 19.7× over the ARM CPU in Simultaneous Localization and Mapping (SLAM) tasks, and a 16.4× speedup over the Intel CPU and 41.6× over the ARM CPU in Structure from Motion (SfM) tasks. Suquan Zhang, Yunfei Xiang, Yuanfan Xu, Qingmin Liao, Yu Wang 0002 |
ISCA | 3 |
| 2025 | REACT3D: Real-time Edge Accelerator for Incremental Training in 3D Gaussian Splatting based SLAM Systemsabstract3D Gaussian Splatting (3DGS) has emerged as a promising approach for high-fidelity scene reconstruction and has been widely adopted in Simultaneous Localization and Mapping (SLAM) systems.3DGS SLAM requires incremental training and rendering of Gaussians in real-time from continuous camera viewpoints.To match the streaming nature of SLAM, 3DGS-based mapping must sustain over 30 frames per second (FPS), which is a widely recognized threshold for maintaining accurate tracking and mapping quality.Existing GPU-based solutions and prior accelerators fall short of this target, primarily due to redundant training computation, unnecessary loss computing, and irregular memory access patterns.To address these challenges, we propose REACT3D, a real-time edge accelerator designed for incremental training in 3DGS SLAM systems.At the algorithmic level, we introduce spatial consistency and convergence aware sparsification, which eliminates redundant computation in both forward and backward rendering by predicting under-optimized regions based on spatial coherence and convergence dynamics.At the architectural level, we design a pixel blockwise fine-grained dataflow to eliminate explicit loss computing, establish a tightly coupled pipeline, and improve hardware utilization.Furthermore, we develop a Content Addressable Memory (CAM)-based Dual-index Gaussian Buffer to resolve discontinuous * Equal contribution. Zhenhua Zhu 0002, Tianchen Zhao, Yunfei Xiang, Huazhong Yang, Yuan Xie 0001, Yu Wang 0002 |
MICRO | 4 |
| 2024 | Invited: Automatic Hardware/Software Design for High-Speed Autonomous Unmanned Aerial Vehicles Guided by a Flight ModelabstractAutonomous Unmanned Aerial Vehicles (UAVs) are on the rise in the industrial and academic communities. Since most UAVs are severely size, weight, and power (SWaP) constrained, building computing system for high-speed UAVs is challenging. Current domain-specific hardware-software (HW-SW) designs for UAVs are mainly bottom-up, focusing on optimizing a single module in the whole system, such as visual-inertial odometry (VIO), depth estimation, or planning. But this leads to underdesign for flight speed as agile navigation depends on a tight combination of multiple modules. To find the optimal HW-SW design for systematic flight performance, we propose a top-down automatic design framework. A flight model is introduced to guide the inter-module and the intra-module HW-SW optimization towards the system-level goal. For the perception algorithms, we define the representative design space. And some critical non-AI operators are accelerated and profiled on embedded GPU to achieve better hardware performance. The design framework is evaluated on a micro UAV equipped with a Nvidia Jetson Orin NX. In a specific navigation scenario, the design found by our framework achieve 40% and 65% increase on flight speed than two manual design methods respectively. Yuanfan Xu, Suquan Zhang, Yunfei Xiang, Hongyang Jia, Yu Wang 0002 |
DAC | 4 |
| 2021 | Prediction interval estimation of landslide displacement using adaptive chicken swarm optimization-tuned support vector machines
Yin Xing, Jianping Yue, Dongjian Cai, Yunfei Xiang |
Appl. Intell. | 6 |
| 2020 | Prior-Attention Residual Learning for More Discriminative COVID-19 Screening in CT ImagesabstractWe propose a conceptually simple framework for fast COVID-19 screening in 3D chest CT images. The framework can efficiently predict whether or not a CT scan contains pneumonia while simultaneously identifying pneumonia types between COVID-19 and Interstitial Lung Disease (ILD) caused by other viruses. In the proposed method, two 3D-ResNets are coupled together into a single model for the two above-mentioned tasks via a novel prior-attention strategy. We extend residual learning with the proposed prior-attention mechanism and design a new so-called prior-attention residual learning (PARL) block. The model can be easily built by stacking the PARL blocks and trained end-to-end using multi-task losses. More specifically, one 3D-ResNet branch is trained as a binary classifier using lung images with and without pneumonia so that it can highlight the lesion areas within the lungs. Simultaneously, inside the PARL blocks, prior-attention maps are generated from this branch and used to guide another branch to learn more discriminative representations for the pneumonia-type classification. Experimental results demonstrate that the proposed framework can significantly improve the performance of COVID-19 screening. Compared to other methods, it achieves a state-of-the-art result. Moreover, the proposed method can be easily extended to other similar clinical applications such as computer-aided detection and diagnosis of pulmonary nodules in CT images, glaucoma lesions in Retina fundus images, etc. Jun Wang 0072, Yiming Bao, Yaofeng Wen, Hongbing Lu, Hu Luo, Yunfei Xiang, Chen Liu 0026, Dahong Qian |
IEEE Trans. Medical Imaging | 6 |