Yiming Gan

dblp:241/3024 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 4 first-author · 18 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Nebula: Infinite-Scale 3D Gaussian Splatting in VR via Collaborative Rendering and Accelerated Stereo Rasterization
abstract
3D Gaussian splatting (3DGS) has drawn significant attention in the architectural community recently. However, current architectural designs often overlook the 3DGS scalability, making them fragile for extremely large-scale 3DGS. Meanwhile, the VR bandwidth requirement makes it impossible to deliver high-fidelity and smooth VR content from the cloud.
Zheng Liu 0022, Xingyang Li, Anbang Wu, Jieru Zhao, Fangxin Liu, Yiming Gan, Jingwen Leng, Yu Feng 0007
ASPLOS (2)7
2026 CHIP-MAP: A Collaborative Optimization Framework for Macro Placement Using Large Language Models
abstract
As integrated circuits continue to grow in both scale and complexity, macro placement plays a critical role in physical design, directly affecting chip-level performance, power, and area (PPA). Traditional macro placement methods, such as simulated annealing, analytical optimization, and reinforcement learning, face limitations including slow convergence, heavy dependence on large datasets, and over-reliance on intermediate PPA indicators rather than final PPA. Large language models (LLMs) offer strong generative power and semantic reasoning that can potentially automate macro layout tasks while addressing the aforementioned problems in traditional methods, but their limited understanding of layout rules and lack of iterative, feedback-driven refinement make direct application challenging. To address this, we propose CHIP-MAP, a macro placement framework based on multi-agent collaboration and feedback-driven optimization. Furthermore, we introduce two innovative tools: the Module Link Weight Analyzer (MWA) and the Standard Cell Usability Score (SCUS), which are designed to guide fine-grained layout refinement. We evaluate CHIP-MAP on five benchmarks ranging from low-power cores to large multi-core processors implemented at 130nm and 45nm technology nodes. Results show that it achieves up to 1.5% area reduction and an average repair of 61.6% of total negative slack (TNS), while also reducing wirelength and improving timing.
Yiming Du, Renye Yan, Yunfan Yang, Frank Qu, Jiajun Tan, ZhiYu Zheng, Yiming Gan, Ling Liang 0003, Zongwei Wang 0001, Yimao Cai
DATE7
2026 SONIC: Smart Optimization for Neural-Integrated CMP with Timing-Aware Fills
abstract
Dummy fill insertion is essential for CMP uniformity but remains challenging due to the nonlinear CMP process, the large optimization space, and timing degradation caused by parasitic coupling. We propose SONIC, a differentiable CMP-driven dummy fill optimization framework that employs a neural CMP simulator to directly optimize planarization objectives using gradient-based methods. SONIC further integrates a timing-aware fill insertion strategy to mitigate coupling capacitance near critical nets. Experimental results demonstrate that SONIC achieves competitive planarization quality with up to 1830× runtime speedup over a full-chip CMP simulator. Compared with the state-of-the-art model-based method, SONIC reduces height variation, line deviation, and outliers by up to 86.16%, 90.10%, and 51.61%, respectively, while achieving a 77.67% runtime reduction and lowering coupling capacitance by 13.05%.
Jiajun Tan, Yiming Du, Yiming Gan, Ling Lang 0002, Yibo Lin, Zongwei Wang 0001, Yimao Cai
DATE4
2026 SPLATONIC: Architectural Support for 3D Gaussian Splatting SLAM via Sparse Processing
abstract
3D Gaussian splatting (3DGS) has emerged as a promising direction for SLAM due to its high-fidelity reconstruction and rapid convergence. However, 3DGS-SLAM algorithms remain impractical for mobile platforms due to their high computational cost, especially for their tracking process. This work introduces Splatonic, a sparse and efficient realtime 3DGS-SLAM algorithm-hardware co-design for resourceconstrained devices. Inspired by classical SLAMs, we propose an adaptive sparse pixel sampling algorithm that reduces the number of rendered pixels by up to$256 \times$while retaining accuracy. To unlock this performance potential on mobile GPUs, we design a novel pixel-based rendering pipeline that improves hardware utilization via Gaussian-parallel rendering and preemptive$\alpha$-checking. Together, these optimizations yield up to$121.7 \times$speedup on the bottleneck stages and$14.6 \times$end-toend speedup on off-the-shelf GPUs. To further address new bottlenecks introduced by our rendering pipeline, we propose a pipelined architecture that simplifies the overall design while addressing newly emerged bottlenecks in projection and aggregation. Evaluated across four 3DGS-SLAM algorithms, Splatonic achieves up to$274.9 \times$speedup and$4738.5 \times$energy savings over mobile GPUs and up to$25.2 \times$speedup and$241.1 \times$energy savings over state-of-the-art accelerators, all with comparable accuracy.
Xiaotong Huang, Tianrui Ma, Yuxiang Xiong, Fangxin Liu, Zhezhi He, Yiming Gan, Zihan Liu 0002, Jingwen Leng, Yu Feng 0007, Minyi Guo
HPCA7
2026 AIRSTONE: Open sourced hardware accelerators and tools for efficient and safe embodied AI computing
Bo Yu 0014, Yuhui Hao, Yiming Gan, Shaoshan Liu
Future Gener. Comput. Syst.3
2026 Corrigendum: Unified and Efficient Factor Graph Accelerator Design for Robotic Optimization
abstract
This is a corrigendum for the article "Unified and Efficient Factor Graph Accelerator Design for Robotic Optimization" published in ACM Trans. Arch. Code Optim. 22, 4, Article 153 (December 2025), 23 pages.
Qiang Liu 0011, Yihao Hua, Yuhui Hao, Bo Yu 0014, Shaoshan Liu, Yiming Gan
ACM Trans. Archit. Code Optim.6
2026 VelKoz: Generating Accelerators for Rigid-Flexible Robots through Domain Specific High-level Synthesis
abstract
Rigid-flexible robots, integrating soft materials with rigid structures, have garnered increasing research interest due to their enhanced capabilities, flexibility, and inherent safety. However, existing control algorithms for these robots often exhibit high computational complexity, hindering real-time implementation. This work proposes VelKoz , an accelerator generation framework tailored for rigid-flexible robot control. It enables users to program in MATLAB and generate synthesizable Verilog code for control algorithms. A key challenge addressed is the integration of robotics domain knowledge with the dataflow representations commonly used in hardware accelerator design. Experimental results demonstrate that the generated accelerators achieve orders-of-magnitude lower latency and energy consumption compared to general-purpose CPUs and outperform customized high-level synthesis (HLS) implementations by 5.3 ×.
Guoshuai Geng, Yuhui Hao, Yinhe Han 0001, Yu Feng 0007, Yiming Gan
ACM Trans. Design Autom. Electr. Syst.5
2025 Aphelios: A Selective Lock-step Neural Processing Unit Design
Yiming Gan, Yuhui Hao, Yinhe Han 0001
ACM Great Lakes Symposium on VLSI2
2025 SLTarch: Towards Scalable Point-Based Neural Rendering by Taming Workload Imbalance and Memory Irregularity
abstract
Rendering is critical in fields like 3D modeling, AR/VR, and autonomous driving, where high-quality, real-time output is essential. Point-based neural rendering (PBNR) offers a photorealistic and efficient alternative to conventional methods, yet it is still challenging to achieve real-time rendering on mobile platforms. We pinpoint two major bottlenecks in PBNR pipelines: LoD search and splatting. LoD search suffers from workload imbalance and irregular memory access, making it inefficient on off-the-shelf GPUs. Meanwhile, splatting introduces severe warp divergence across GPU threads due to its inherent sparsity.To tackle these challenges, we propose SLTarch, an algorithm-architecture co-designed framework. At its core, SLTarch introduces SLTree, a dedicated subtree-based data structure, and LTcore, a specialized hardware architecture tailored for efficient LoD search. Additionally, we co-design a divergence-free splatting algorithm with our simple yet principled hardware augmentation, SPcore, to existing PBNR accelerators. Compared to a mobile GPU, SLTarch achieves 3.9× speedup and 98% energy savings with negligible architecture overhead. Compared to existing accelerator designs, SLTarch achieves 1.8× speedup with 54% energy savings.
Xingyang Li, Yu Feng 0007, Yiming Gan, Jieru Zhao, Zihan Liu 0002, Jingwen Leng, Minyi Guo
ICCAD4
2025 KARMA: Augmenting Embodied AI Agents with Long-and-Short Term Memory Systems
abstract
Embodied AI agents responsible for executing interconnected, long-sequence household tasks often face difficulties with in-context memory, leading to inefficiencies and errors in task execution. To address this issue, we introduce KARMA, an innovative memory system that integrates longterm and short-term memory modules, enhancing large language models (LLMs) for planning in embodied agents through memory-augmented prompting. Karma distinguishes between long-term and short-term memory, with long-term memory capturing comprehensive 3D scene graphs as representations of the environment, while short-term memory dynamically records changes in objects' positions and states. This dualmemory structure allows agents to retrieve relevant past scene experiences, thereby improving the accuracy and efficiency of task planning. Short-term memory employs strategies for effective and adaptive memory replacement, ensuring the retention of critical information while discarding less pertinent data. Compared to state-of-the-art embodied agents enhanced with memory, our memory-augmented embodied AI agent improves success rates by$1.3 \times$and$2.3 \times$in Composite Tasks and Complex Tasks within the AI2-THOR simulator, respectively, and enhances task execution efficiency by$3.4 \times$and$62.7 \times$. Furthermore, we demonstrate that KARMA's plug-and-play capability allows for seamless deployment on real-world robotic systems, such as mobile manipulation platforms. Through this plug-and-play memory system, KARMA significantly enhances the ability of embodied agents to generate coherent and contextually appropriate plans, making the execution of complex household tasks more efficient. Our code is available at https://github.com/WZX0Swarm0Robotics/KARMA/tree/master.
Bo Yu 0014, Junzhe Zhao, Sai Hou, Xing Hu 0001, Yinhe Han 0001, Yiming Gan
ICRA9
2025 Dadu-Corki: Algorithm-Architecture Co-Design for Embodied AI-powered Robotic Manipulation
abstract
Embodied AI robots have the potential to fundamentally improve the way human beings live and manufacture.Continued progress in the burgeoning field of using large language models to control robots depends critically on an efficient computing substrate, and this trend is strongly evident in manipulation tasks.In particular, today's computing systems for embodied AI robots for manipulation tasks are designed purely based on the interest of algorithm developers, where robot actions are divided into a discrete frame basis.Such an execution pipeline creates high latency and energy consumption.This paper proposes Corki, an algorithm-architecture co-design framework for real-time embodied AI-powered robotic manipulation applications.We aim to decouple LLM inference, robotic control, and data communication in the embodied AI robots' compute pipeline.Instead of predicting action for one single frame, * equal contribution.
Yiyang Huang 0002, Yuhui Hao, Bo Yu 0014, Yuxin Yang 0002, Feng Min, Yinhe Han 0001, Lin Ma 0002, Shaoshan Liu, Qiang Liu 0011, Yiming Gan
ISCA11
2025 KINDRED: Heterogeneous Split-Lock Architecture for Safe Autonomous Machines
abstract
With the increasing practicality of autonomous vehicles and drones, the importance of reliability requirements has escalated substantially. In many instances, traditional system designs tend to overlook reliability issues, emphasizing primarily on performance constraints. However, certain designers may opt for a lock-step (redundant) system design, duplicating every component, which in turn can result in significant performance, energy, and cost overheads. In software for autonomous machines, such as self-driving vehicles, performance degradation can increase reaction time, posing safety risks and reducing mission success rates. This article introduces a novel multi-domain lock-step system design, Kindred , which places a strong emphasis on maximizing reliability while minimizing performance overhead. The proposed approach capitalizes on the inherent diversity in fault tolerance among various tasks within autonomous machine software, intelligently scheduling only the vulnerable nodes in the lock-domain. The primary challenge addressed in this study involves the intelligent task scheduling across different domains, complemented by efficient error detection and correction in the lock-domain. In a real system demonstration, we illustrate the effectiveness of Kindred , showcasing its ability to attain the same level of reliability as a full lock-step system while incurring only a mere 2.8% overhead, as opposed to a fully split system, indicating the advantages and potential of our multi-domain lock-step system design in achieving high reliability without compromising performance.
Yiming Gan, Jingwen Leng, Bo Yu 0014, Yuhao Zhu 0001
ACM Trans. Archit. Code Optim.1
2024 ORIANNA: An Accelerator Generation Framework for Optimization-based Robotic Applications
abstract
Despite extensive efforts, existing approaches to design accelerators for optimization-based robotic applications have limitations. Some approaches focus on accelerating general matrix operations, but they fail to fully exploit the specific sparse structure commonly found in many robotic algorithms. On the other hand, certain methods require manual design of dedicated accelerators, resulting in inefficiencies and significant non-recurring engineering (NRE) costs.
Yuhui Hao, Yiming Gan, Bo Yu 0014, Qiang Liu 0011, Yinhe Han 0001, Zishen Wan, Shaoshan Liu
ASPLOS (2)2
2024 Benchmarking and Optimizing Federated Learning with Hardware-related Metrics
Kai Pan, Yapeng Tian, Yinhe Han 0001, Yiming Gan
BMVC4
2023 BLITZCRANK: Factor Graph Accelerator for Motion Planning
abstract
Factor graph is a graph representing the factorization of a probability distribution function and serves as a perfect abstraction in many autonomous machine computing stacks, such as planning, localization, tracking and control, which are challenging tasks for autonomous systems with real-time and energy constraints.In this paper, we present BLITZCRANK, an accelerator for motion planning algorithms using the abstraction of a factor graph. By formulating motion planning as a factor graph inference, we successfully reduce the scale of the problem and utilize the inherent matrix sparsity. BLITZCRANK is able to realize the user-defined optimal design by finding the optimal order of the factor graph inference. With a domain specific balancing order, BLITZCRANK achieves up to 7.4× speed up and 29.7× energy reduction compared to the software implementation on Intel CPU.
Yuhui Hao, Yiming Gan, Bo Yu 0014, Qiang Liu 0011, Shaoshan Liu, Yuhao Zhu 0001
DAC2
2023 An Energy Efficient and Runtime Reconfigurable Accelerator for Robotic Localization
abstract
Accurate and efficient localization of robots under limited on-board resources has fueled specialized localization accelerators. Despite many recent efforts, accelerating robotic localization is still fundamentally challenging. To tackle the challenges, the paper proposes a configurable hardware architecture and a design space optimization method to automatically generate an optimal accelerator design under the design constraints. Data locality, sparsity, and fixed-point arithmetic optimization techniques that are specific to the localization algorithm are exploited to customize the accelerator. In addition, a low-cost runtime configuration mechanism is proposed to enable the accelerator to continuously optimize itself at runtime according to the operating environment to save power while sustaining performance and accuracy. The evaluation on FPGA demonstrates that the proposed accelerator achieves orders of magnitude performance improvement and/or energy savings compared to the software implementation on Intel and Arm CPUs; and substantially outperforms existing FPGA accelerators in terms of performance and energy.
Qiang Liu 0011, Yuhui Hao, Weizhuang Liu, Bo Yu 0014, Yiming Gan, Jie Tang 0003, Shaoshan Liu, Yuhao Zhu 0001
IEEE Trans. Computers5
2022 Crescent: taming memory irregularities for accelerating deep point cloud analytics
abstract
3D perception in point clouds is transforming the perception ability of future intelligent machines. Point cloud algorithms, however, are plagued by irregular memory accesses, leading to massive inefficiencies in the memory sub-system, which bottlenecks the overall efficiency.
Yu Feng 0007, Gunnar Hammonds, Yiming Gan, Yuhao Zhu 0001
ISCA3
2022 Braum: Analyzing and Protecting Autonomous Machine Software Stack
abstract
Autonomous machines, such as Autonomous Vehicles (AV), are vulnerable to a variety of different faults such as radiation-induced soft/transient errors, adversarial attacks, and software bugs, which all jeopardize the reliability of autonomous machines. How vulnerable the AV software stack is to different error sources, however, remains an open question. This paper performs comprehensively fault injections to study how the AV software stack behaves under different error sources. We show that algorithms in an AV software stack inherently possess different forms of masking mechanisms. Based on the characteristic of the inherent fault tolerance mechanisms, we formalize the notion of Fault Tolerance Level (FTL), which quantifies how faults in an algorithm can be masked and/or attenuated without affecting the actuator commands, providing opportunities to relax fault protection. Leveraging the FTL formulation, we propose a dynamic protection system, which, at the high level, spends the limited protection budget (e.g., spatial/temporal redundancy) on the most vulnerable parts of the AV software (i.e., with the lowest FTL). Using Autoware as a case study, we show that our system reduces the error rate of AV software stack by more than 90% with negliaible performance overhead.
Yiming Gan, Paul N. Whatmough, Jingwen Leng, Bo Yu 0014, Shaoshan Liu, Yuhao Zhu 0001
ISSRE1
2021 Eudoxus: Characterizing and Accelerating Localization in Autonomous Machines Industry Track Paper
abstract
We develop and commercialize autonomous machines, such as logistic robots and self-driving cars, around the globe. A critical challenge to our—and any—autonomous machine is accurate and efficient localization under resource constraints, which has fueled specialized localization accelerators recently. Prior acceleration efforts are point solutions in that they each specialize for a specific localization algorithm. In real-world commercial deployments, however, autonomous machines routinely operate under different environments and no single localization algorithm fits all the environments. Simply stacking together point solutions not only leads to cost and power budget overrun, but also results in an overly complicated software stack. This paper demonstrates our new software-hardware co-designed framework for autonomous machine localization, which adapts to different operating scenarios by fusing fundamental algorithmic primitives. Through characterizing the software framework, we identify ideal acceleration candidates that contribute significantly to the end-to-end latency and/or latency variation. We show how to co-design a hardware accelerator to systematically exploit the parallelisms, locality, and common building blocks inherent in the localization framework. We build, deploy, and evaluate an FPGA prototype on our next-generation self-driving cars. To demonstrate the flexibility of our framework, we also instantiate another FPGA prototype targeting drones, which represent mobile autonomous machines. We achieve about $2 \times$ speedup and $4 \times$ energy reduction compared to widely-deployed, optimized implementations on general-purpose platforms.
Yiming Gan, Bo Yu 0014, Boyuan Tian, Leimeng Xu, Shaoshan Liu, Qiang Liu 0011, Jie Tang 0003, Yuhao Zhu 0001
HPCA1
2021 Archytas: A Framework for Synthesizing and Dynamically Optimizing Accelerators for Robotic Localization
abstract
Despite many recent efforts, accelerating robotic computing is still fundamentally challenging for two reasons. First, robotics software stack is extremely complicated. Manually designing an accelerator while meeting the latency, power, and resource specifications is unscalable. Second, the environment in which an autonomous machine operates constantly changes; a static accelerator design leads to wasteful computation.
Weizhuang Liu, Bo Yu 0014, Yiming Gan, Qiang Liu 0011, Jie Tang 0003, Shaoshan Liu, Yuhao Zhu 0001
MICRO3
2020 Low-Latency Proactive Continuous Vision
abstract
Continuous vision is the cornerstone of a diverse range of intelligent applications found on emerging computing platforms such as autonomous machines and Augmented Reality glasses. A critical issue in today's continuous vision systems is their long end-to-end frame latency, which significantly impacts the system agility and user experience. We find that the long latency is fundamentally caused by the serialized execution model of today's continuous vision pipeline, whose key stages, including sensing, imaging, and vision computations, execute sequentially, leading to long frame latency.
Yiming Gan, Yuxian Qiu, Jingwen Leng, Yuhao Zhu 0001
PACT1
2020 TinyLSTMs: Efficient Neural Speech Enhancement for Hearing Aids
abstract
Modern speech enhancement algorithms achieve remarkable noise suppression by means of large recurrent neural networks (RNNs). However, large RNNs limit practical deployment in hearing aid hardware (HW) form-factors, which are battery powered and run on resource-constrained microcontroller units (MCUs) with limited memory capacity and compute capability. In this work, we use model compression techniques to bridge this gap. We define the constraints imposed on the RNN by the HW and describe a method to satisfy them. Although model compression techniques are an active area of research, we are the first to demonstrate their efficacy for RNN speech enhancement, using pruning and integer quantization of weights/activations. We also demonstrate state update skipping, which reduces the computational load. Finally, we conduct a perceptual evaluation of the compressed models to verify audio quality on human raters. Results show a reduction in model size and operations of 11.9$\times$ and 2.9$\times$, respectively, over the baseline for compressed models, without a statistical difference in listening preference and only exhibiting a loss of 0.55dB SDR. Our model achieves a computational latency of 2.39ms, well within the 10ms target and 351$\times$ better than previous work.
Igor Fedorov, Marko Stamenovic, Carl Jensen, Li-Chia Yang, Ari Mandell, Yiming Gan, Matthew Mattina, Paul N. Whatmough
INTERSPEECH6
2020 Ptolemy: Architecture Support for Robust Deep Learning
abstract
Deep learning is vulnerable to adversarial attacks, where carefully-crafted input perturbations could mislead a well-trained Deep Neural Network (DNN) to produce incorrect results. Adversarial attacks jeopardize the safety, security, and privacy of DNN-enabled systems. Today's countermeasures to adversarial attacks either do not have the capability to detect adversarial samples at inference-time, or introduce prohibitively high overhead to be practical at inference-time.We propose Ptolemy, an algorithm-architecture co-designed system that detects adversarial attacks at inference time with low overhead and high accuracy. We exploit the synergies between DNN inference and imperative program execution: an input to a DNN uniquely activates a set of neurons that contribute significantly to the inference output, analogous to the sequence of basic blocks exercised by an input in a conventional program. Critically, we observe that adversarial samples tend to activate distinctive paths from those of benign inputs. Leveraging this insight, we propose an adversarial sample detection framework, which uses canary paths generated from offline profiling to detect adversarial samples at runtime. The Ptolemy compiler along with the co-designed hardware enable efficient execution by exploiting the unique algorithmic characteristics. Extensive evaluations show that Ptolemy achieves higher or similar adversarial sample detection accuracy than today's mechanisms with a much lower (as low as 2%) runtime overhead.
Yiming Gan, Yuxian Qiu, Jingwen Leng, Minyi Guo, Yuhao Zhu 0001
MICRO1