EDBT 2026 Demo / reviewers in the wild / expert
Yiqi Jing
dblp:352/9615
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0004-7690-5356ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GSNorm: An Efficient 3D Gaussian Rendering Accelerator with Splat Normalization and LUT-assist Rasterizationabstract3D Gaussian Splatting recently emerged as the new SOTA approach for many computer graphic tasks. While Gaussian Splatting has demonstrated impressive rendering quality and performance on GPUs, real-time GS rendering on edge devices is still challenging. We identified the unbalanced Rendering pipeline and the uneven Gaussian distribution as the main obstacles to efficient rendering. To address these problems, we present GSNorm, a rendering accelerator with an online quantization preprocessor for per-gaussian coordinate transformation to normalize Gaussian footprints for pixel-wise calculation reduction. A LUT-based quantized rendering design is also presented to break the pipeline data dependency. Furthermore, a depth-guided cluster-sorting unit is incorporated to improve Gaussian sorting efficiency. GSNorm accelerator is implemented and evaluated in TSMC 22 nm technology with several real-world scenes, providing significant rendering efficiency and performance improvements for real-time applications. Peiran Yan, Yiqi Jing, Le Ye |
ASP-DAC | 3 |
| 2025 | Local-GS: An Order-Independent Gaussian Splatting Training Accelerator Exploiting Splat Localityabstract3D Gaussian Splatting has emerged as the SOTA approach for 3D representation and view synthesis. While Gaussian Splatting has demonstrated impressive capability and rendering quality on desktop GPUs, achieving on-demand training on resource-constrained edge devices is still challenging. In this work, we identified the training bottleneck from a few perspectives including algorithm splat locality and the limited memory and hardware under-utilization. To address these problems, we present Local-GS, a 3D Gaussian Splatting training accelerator with order-independent rendering to break the depth-wise data dependency between overlapping Gaussians. We further incorporate a parallel pixel intersection test unit to schedule thread workload based on Gaussian splat locality and improve hardware utilization. A set of unified training-rendering cores are designed to achieve efficient splat-level parallel rendering and gradient propagation. Our Local-GS is implemented in 7 nm and is evaluated by several real-world 3D scenes. Compared to edge Jetson NX GPU, Local-GS achieve 26.9-53 $\times$ training speedup and three orders of magnitude efficiency boost. Qinzhe Zhi, Yiqi Jing, Le Ye, Ru Huang 0001 |
DAC | 3 |
| 2024 | AIG-CIM: A Scalable Chiplet Module with Tri-Gear Heterogeneous Compute-in-Memory for Diffusion AccelerationabstractThe emergence of Diffusion models has gained significant attention in the field of Artificial Intelligence Generated Content. While Diffusion demonstrates impressive image generation capability, it faces hardware deployment challenges due to its unique model architecture and computation requirement. In this paper, we present a hardware accelerator design, i.e. AIG-CIM, which incorporates tri-gear heterogeneous digital compute-in-memory to address the flexible data reuse demands in Diffusion models. Our framework offers a collaborative design methodology for large generative models from the computational circuit-level to the multi-chip-module system-level. We implemented and evaluated the AIG-CIM accelerator using TSMC 22nm technology. For several Diffusion inferences, scalable AIG-CIM chiplets achieve 21.3× latency reduction, up to 231.2× throughput improvement and three orders of magnitude energy efficiency improvement compared to RTX 3090 GPU. Yiqi Jing, Meng Wu 0005, Yufei Ma 0002, Ru Huang 0001, Le Ye |
DAC | 1 |
| 2023 | A Model-Specific End-to-End Design Methodology for Resource-Constrained TinyML HardwareabstractTiny machine learning (TinyML) becomes appealing as it enables machine learning on resource-constrained devices with ultra low energy and small form factor. In this paper, a model-specific end-to-end design methodology is presented for TinyML hardware design. First, we introduce an end-to-end system evaluation method using Roofline models, which considering both AI and other general-purpose computing to guide the architecture design choices. Second, to improve the efficiency of AI computation, we develop an enhanced design space exploration framework, TinyScale, to enable optimal low-voltage operation for energy-efficient TinyML. Finally, we present a use case driven design selection method to search the optimal hardware design across a set of application use cases. Our model-specific design methodology is evaluated on both TSMC 22nm and 55nm technology for MLPerf Tiny benchmark and a keyword spotting (KWS) SoC design. With the help of our end-to-end design methodology, an optimal TinyML hardware can be automatically explored with significant energy and EDP improvements for a diverse of TinyML use cases. Yanchi Dong, Kaixuan Du, Yiqi Jing, Qijun Wang, Pixian Zhan, Fengyun Yan, Yufei Ma 0002, Yun Liang 0001, Le Ye, Ru Huang 0001 |
DAC | 4 |
| 2023 | An Information-Aware Adaptive Data Acquisition System using Level-Crossing ADC with Signal-Dependent Full Scale and Adaptive Resolution for IoT ApplicationsabstractThis paper proposes an information-aware (IA) adaptive data acquisition (ADA) system for the Internet of Things (IoT) applications. The system can obtain valid information adaptively thanks to 1) signal-dependent full-scale feature tracks the amplitude-domain activity of the event; 2) level-crossing (LC) ADC with slope detector delivers the time-domain activity; 3) the IA algorithm determines the quantization resolution according to the detected signal activities. The proposed clock-free event-driven ADA system can reject the redundant data, and compress the valid data from the source, thus saving its power and the power of subsequent data-processing systems. The long-term average power consumption of the system is 128 nW, the resolution varies from 3 to 7 bits according to the input signal state. Compared with conventional ADCs, LC-ADC can compress the data by 2.5x [1]. Further, the proposed system has 15x higher compression ratio (CR) than that of LC-ADC. Yiqi Jing, Zhixuan Wang, Linxiao Shen, Yihan Zhang 0002, Jiayoon Ru, Le Ye |
ISCAS | 1 |
| 2023 | DCIM-3DRec: A 3D Reconstruction Accelerator with Digital Computing-in-Memory and Octree-Based SchedulerabstractLearning-based 3D reconstruction has evolved rapidly with promising quality, while it requires high-performance hardware for interactive applications. In this work, a reconstruction accelerator called DCIM-3DRec is presented which leverages digital computing-in-memory (DCIM) design to facilitate learning-based reconstruction deployment on realtime and low-power edge platforms. The DCIM-3DRec is designed with the following features: a reconfigurable DCIM macro array for high data reuse and macro utilization, and an Octree-based subdivision scheduler for efficient management of 3D space prediction. The DCIM-3DRec accelerator is implemented and evaluated in TSMC 55 nm technology, with a DCIM macro efficiency of 19.4 TOPS/W at INT8. Overall, the DCIM-3DRec accelerator achieves 23× performance gain and four orders of magnitude energy efficiency improvement compared to a Nvidia RTX3090 GPU. Yiqi Jing, Meng Wu 0005, Fengyun Yan, Yufei Ma 0002, Le Ye |
ISLPED | 1 |
| 2023 | Research progress on low-power artificial intelligence of things (AIoT) chip design
Le Ye, Zhixuan Wang, Yufei Ma 0002, Linxiao Shen, Yihan Zhang 0002, Meng Wu 0005, Ying Liu 0069, Yiqi Jing, Hao Zhang 0119, Ru Huang 0001 |
Sci. China Inf. Sci. | 11 |