EDBT 2026 Demo / reviewers in the wild / expert
Yu Feng 0007
dblp:30/4550-7
· DBLP profile ↗
38ranked-venue papers
13as first author
34since 2021 · last 2026
0000-0002-2192-5737ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 11 first-author · 28 since 2021Software engineering, systems software and programming languages · 15 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the (In-)Security of the Shuffling Defense in the Transformer Secure InferenceabstractZhengyi Li, Yakai Wang, Jingwen Leng, Kang Yang, Yu Yu, Jiaping Gui, Yu Feng, Ning Liu, Minyi Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhengyi Li 0002, Yakai Wang, Jingwen Leng, Kang Yang 0002, Yu Yu 0001, Jiaping Gui, Yu Feng 0007, Ning Liu 0007, Minyi Guo |
ACL (1) | 7 |
| 2026 | M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit QuantizationabstractExisting low-bit Microscaling (MX) formats, such as MXFP4, often suffer from substantial accuracy degradation due to the use of a shared scaling factor with the Power-of-Two format. In this work, we explore strategies that introduce minimal metadata to recover accuracy lost during quantization while maintaining high bit efficiency across a wide range of large language models. We propose a complete algorithm-hardware co-design based on flexible metadata, featuring an online quantization with simple encoding. To support the proposed method efficiently, we implement a lightweight hardware unit and integrate it into the accelerator. Evaluation results demonstrate that our method substantially narrows the accuracy gap, achieving on average a 70.63% reduction in accuracy loss compared to MXFP4 and a 37.30% reduction relative to the latest NVFP4 on LLM benchmarks. Furthermore, our design delivers up to 1.91× speedup and 1.75× energy savings over state-of-the-art accelerators. Weiming Hu 0005, Chen Zhang 0001, Cong Guo 0003, Yu Feng 0007, Tianchi Hu, Guanglin Li 0005, Guipeng Hu, Jingwen Leng |
ASPLOS (2) | 6 |
| 2026 | EARTH: An Efficient MoE Accelerator with Entropy-Aware Speculative Prefetch and Result ReuseabstractMixture-of-Experts (MoE) models significantly reduce computation in large language models by activating only a subset of experts per input token, but they introduce severe memory bottlenecks due to the large number of expert parameters. Existing offloading and prefetching strategies either incur accuracy loss, prohibitively high memory traffic, or high decoding overhead, limiting deployment on resource-constrained hardware. In this work, we present EARTH, a hardware–software co-design that addresses these challenges through three key innovations. First, we propose a dual-entropy encoding scheme that decomposes each expert into a high-information base and a delta component, enabling compact storage while preserving accuracy via adaptive precision management. Second, we introduce a delta-aware speculative prefetching and reuse mechanism that preloads base components of predicted experts and selectively fetches deltas, reusing previously computed delta patterns to reduce memory traffic and redundant computation. Third, we design a hardware accelerator that is co-designed to efficiently support this encoding and prefetching strategy, optimizing execution order, parallelism, and memory utilization. Across representative MoE workloads, EARTH reduces data movement overhead, improves prefetch efficiency, and achieves up to 2.10× speedup compared to state-of-the-art baselines, while maintaining high model accuracy. Fangxin Liu, Ning Yang 0012, Jingkui Yang, Zongwu Wang, Chenyang Guan, Yu Feng 0007, Li Jiang 0002, Haibing Guan |
ASPLOS (2) | 6 |
| 2026 | Nebula: Infinite-Scale 3D Gaussian Splatting in VR via Collaborative Rendering and Accelerated Stereo Rasterizationabstract3D Gaussian splatting (3DGS) has drawn significant attention in the architectural community recently. However, current architectural designs often overlook the 3DGS scalability, making them fragile for extremely large-scale 3DGS. Meanwhile, the VR bandwidth requirement makes it impossible to deliver high-fidelity and smooth VR content from the cloud. Zheng Liu 0022, Xingyang Li, Anbang Wu, Jieru Zhao, Fangxin Liu, Yiming Gan, Jingwen Leng, Yu Feng 0007 |
ASPLOS (2) | 9 |
| 2026 | FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core ConnectionabstractThe scaling of computation throughput continues to outpace improvements in memory bandwidth, making many deep learning workloads memory-bound. Kernel fusion is a key technique to alleviate this problem, but the fusion strategies of existing compilers and frameworks are limited to using local scratchpad memory. When the intermediate results exceed the limited capacity (such as FFN), the fusion fails. Although modern GPUs (like the NVIDIA H100) now incorporate an inter-core connection mechanism known as Distributed Shared Memory (DSM)—providing a larger, high-bandwidth, and low-latency on-chip memory pool—this hardware potential has yet to be exploited by software frameworks. To bridge this gap, we present FlashFuser, the first compiler framework to utilize inter-core connection for kernel fusion on modern GPUs. FlashFuser extends established fusion techniques to the DSM domain through three core contributions. First, we propose a powerful DSM-based communication abstraction that formalizes complex cluster-based data exchange patterns, such as reduce, shuffle and multiply. Second, we introduce a dataflow analyzer that generalizes loop scheduling, resource mapping, and tile selection to the distributed memory hierarchy; it determines the optimal execution order and tile sizes by quantifying data movement across memory levels. Finally, FlashFuser integrates these components into a unified search engine that employs analytical cost modeling and DSM-aware pruning strategies to efficiently discover the optimal execution plan. Our evaluation on an NVIDIA H100 GPU shows that FlashFuser reduces memory access by 58 % and delivers kernel speedups of$3.3 x$against highly-tuned libraries and 4.1x against state-of-the-art compilers, resulting in a$1.24 \times$end-to-end speedup. Yangjie Zhou 0001, Zihan Liu 0002, Xinhao Luo, Yijia Diao, Minyi Guo, Jidong Zhai, Yu Feng 0007, Chen Zhang 0001, Anbang Wu, Jingwen Leng |
HPCA | 8 |
| 2026 | SPLATONIC: Architectural Support for 3D Gaussian Splatting SLAM via Sparse Processingabstract3D Gaussian splatting (3DGS) has emerged as a promising direction for SLAM due to its high-fidelity reconstruction and rapid convergence. However, 3DGS-SLAM algorithms remain impractical for mobile platforms due to their high computational cost, especially for their tracking process. This work introduces Splatonic, a sparse and efficient realtime 3DGS-SLAM algorithm-hardware co-design for resourceconstrained devices. Inspired by classical SLAMs, we propose an adaptive sparse pixel sampling algorithm that reduces the number of rendered pixels by up to$256 \times$while retaining accuracy. To unlock this performance potential on mobile GPUs, we design a novel pixel-based rendering pipeline that improves hardware utilization via Gaussian-parallel rendering and preemptive$\alpha$-checking. Together, these optimizations yield up to$121.7 \times$speedup on the bottleneck stages and$14.6 \times$end-toend speedup on off-the-shelf GPUs. To further address new bottlenecks introduced by our rendering pipeline, we propose a pipelined architecture that simplifies the overall design while addressing newly emerged bottlenecks in projection and aggregation. Evaluated across four 3DGS-SLAM algorithms, Splatonic achieves up to$274.9 \times$speedup and$4738.5 \times$energy savings over mobile GPUs and up to$25.2 \times$speedup and$241.1 \times$energy savings over state-of-the-art accelerators, all with comparable accuracy. Xiaotong Huang, Tianrui Ma, Yuxiang Xiong, Fangxin Liu, Zhezhi He, Yiming Gan, Zihan Liu 0002, Jingwen Leng, Yu Feng 0007, Minyi Guo |
HPCA | 10 |
| 2026 | ORANGE: Exploring Ockham's Razor for Neural Rendering by Accelerating 3DGS on NPUs with GEMM-Friendly Blending and Balanced Workloadsabstract3D Gaussian Splatting (3DGS) is an emerging neural rendering technique that delivers efficient and high-fidelity rendering, meeting the growing demands of applications such as AR/VR. As 3DGS is increasingly integrated into diverse applications, DNNs are often deployed alongside it to support tasks such as skeletal pose estimation for human avatars or semantic processing for 3D perception. Unfortunately, existing domain-specific accelerators (DSAs) designed for 3DGS excel at rendering but struggle to execute DNN workloads efficiently. Moreover, these DSAs incur significant design and fabrication costs, limiting their practicality. To address these challenges, we propose ORANGE, a novel approach that enables general-purpose DNN-oriented Neural Processing Units (NPUs) to efficiently execute 3DGS without requiring specialized accelerators. The key insight of ORANGE is that we introduce a GEMM-friendly blending process, which reformulates the conventional 3DGS blending operation to fully utilize the matrix multiplication units prevalent in NPUs during rendering. Additionally, to mitigate workload imbalances caused by variable execution latencies across tiles, we develop a sampling-based latency prediction method paired with a tile batching strategy to minimize idle computing resources. Experiments demonstrate that ORANGE achieves up to$1.67 \times$and$15.5 \times$speedup compared to state-of-the-art 3DGS accelerators and the NVIDIA Xavier NX GPU, respectively, in neural rendering tasks. Our approach offers a cost-effective and versatile solution, adhering to the principle of Ockham's Razor by maximizing efficiency without specialized hardware. Haomin Li 0002, Yun Liang 0001, Fangxin Liu, Zongwu Wang, Yu Feng 0007, Liqiang Lu, Li Jiang 0002, Haibing Guan |
HPCA | 6 |
| 2026 | STEP: Adaptive Spatio-Temporal Expert Prefetching for Low-Latency and Memory-Efficient MoE Inference
Fangxin Liu, Ning Yang 0012, Zongwu Wang, Chenyang Guan, Haomin Li 0002, Yu Feng 0007, Liqiang Lu, Siran Yang, Jiamang Wang, Lin Qu, Li Jiang 0002, Haibing Guan |
ISCA | 6 |
| 2026 | ELSA: An Elastic Snn Inference Architecture for Efficient Neuromorphic Computing
Kang You, Chen Nie, Lee Jun Yan, Ziling Wei, Yu Feng 0007, Honglan Jiang, Zhezhi He |
ISCA | 7 |
| 2026 | APU: Accelerate Point Cloud Neural Networks via Unified Processing-in-SRAM ArchitectureabstractRecent advances in deep learning have expanded point cloud applications by point-based neural networks (PNNs). However, the escalating complexity and computational demands of PNNs overwhelm conventional computers. Specialized PNN accelerators have emerged, significantly outperforming modern CPUs and GPUs. Nevertheless, existing designs remain inefficient when handling performance-critical mapping kernels of PNNs, involving diverse arithmetic functions (e.g., add, multiply, sort) across separate hardware modules. This fragmentation restricts hardware sharing and data locality, leading to area overhead, redundant data movements, and under-utilization. Therefore, a unified and efficient micro-architecture for mapping kernels is needed to enhance performance and reduce data transfers. This paper presents APU, an efficient processing-in-memory (PIM) architecture for PNN acceleration. We introduce the first unified SRAM-PIM micro-architecture that supports all mapping kernels in mainstream PNNs. Data movement is reduced through extensive on-chip memory and maximized data locality viain-situcomputing approach. At the algorithmic level, we introduce mask grouping and aggregation to eliminate costly sorting operations, enabled by hardware support for in-memory vector max-search. This refined strategy reduces computational overhead and data transfers while improving inference accuracy.We further enhance performance by exploiting parallelism across PNN operations and applying mixed-precision quantization. Evaluated on real-world PNN workloads, APU outperforms the state-of-the-art accelerator by 2.54× in speedup and 4.54× in energy saving. Chen Nie, Kang You, Yu Feng 0007, Limin Xiao 0002, Weifeng Zhang 0003, Zhezhi He |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | VelKoz: Generating Accelerators for Rigid-Flexible Robots through Domain Specific High-level SynthesisabstractRigid-flexible robots, integrating soft materials with rigid structures, have garnered increasing research interest due to their enhanced capabilities, flexibility, and inherent safety. However, existing control algorithms for these robots often exhibit high computational complexity, hindering real-time implementation. This work proposes VelKoz , an accelerator generation framework tailored for rigid-flexible robot control. It enables users to program in MATLAB and generate synthesizable Verilog code for control algorithms. A key challenge addressed is the integration of robotics domain knowledge with the dataflow representations commonly used in hardware accelerator design. Experimental results demonstrate that the generated accelerators achieve orders-of-magnitude lower latency and energy consumption compared to general-purpose CPUs and outperform customized high-level synthesis (HLS) implementations by 5.3 ×. Guoshuai Geng, Yuhui Hao, Yinhe Han 0001, Yu Feng 0007, Yiming Gan |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2025 | StreamGrid: Streaming Point Cloud Analytics via Compulsory Splitting and Deterministic TerminationabstractPoint clouds are increasingly important in intelligent applications, but frequent off-chip memory traffic in accelerators causes pipeline stalls and leads to high energy consumption. While conventional line buffer techniques can eliminate off-chip traffic, they cannot be directly applied to point clouds due to their inherent computation patterns. To address this, we introduce two techniques: compulsory splitting and deterministic termination, enabling fully-streaming processing. We further propose StreamGrid, a framework that integrates these techniques and automatically optimizes on-chip buffer sizes. Our evaluation shows StreamGrid reduces on-chip memory by 61.3% and energy consumption by 40.5% with marginal accuracy loss compared to the baselines without our techniques. Additionally, we achieve 10.0× speedup and 3.9× energy efficiency over state-of-the-art accelerators. Yu Feng 0007, Zheng Liu 0022, Weikai Lin, Zihan Liu 0002, Jingwen Leng, Minyi Guo, Zhezhi He, Jieru Zhao, Yuhao Zhu 0001 |
ASPLOS (2) | 1 |
| 2025 | MetaSapiens: Real-Time Neural Rendering with Efficiency-Aware Pruning and Accelerated Foveated RenderingabstractPoint-Based Neural Rendering (PBNR) is emerging as a promising class of rendering techniques, which are permeating all aspects of society, driven by a growing demand for real-time, photorealistic rendering in AR/VR and digital twins. Achieving real-time PBNR on mobile devices is challenging. Weikai Lin, Yu Feng 0007, Yuhao Zhu 0001 |
ASPLOS (1) | 2 |
| 2025 | SNAPPIX: Efficient-Coding-Inspired In-Sensor Compression for Edge VisionabstractEnergy-efficient image acquisition on the edge is crucial for enabling remote sensing applications where the sensor node has weak compute capabilities and must transmit data to a remote server/cloud for processing. To reduce the edge energy consumption, this paper proposes a sensor-algorithm co-designed system called SNAPPIX, which compresses raw pixels in the analog domain inside the sensor. We use coded exposure (CE) as the in-sensor compression strategy as it offers the flexibility to sample, i.e., selectively expose pixels, both spatially and temporally. SnapPix has three contributions. First, we propose a task-agnostic strategy to learn the sampling/exposure pattern based on the classic theory of efficient coding. Second, we codesign the downstream vision model with the exposure pattern to address the pixel-level non-uniformity unique to CE-compressed images. Finally, we propose lightweight augmentations to the image sensor hardware to support our in-sensor CE compression. Evaluating on action recognition and video reconstruction, SnapPix outperforms state-of-the-art video-based methods at the same speed while reducing the energy by up to $15.4 \times$. We have open-sourced the code at: https://github.com/horizonresearch/SnapPix. Weikai Lin, Tianrui Ma, Adith Boloor, Yu Feng 0007, Ruofan Xing, Xuan Zhang 0001, Yuhao Zhu 0001 |
DAC | 4 |
| 2025 | STREAMINGGS: Voxel-Based Streaming 3D Gaussian Splatting with Memory Optimization and Architectural Supportabstract3D Gaussian Splatting (3DGS) has gained popularity for its efficiency and sparse Gaussian-based representation. However, 3DGS struggles to meet the real-time requirement of 90 frames per second (FPS) on resource-constrained mobile devices, achieving only 2 to 9 FPS. Existing accelerators focus on compute efficiency but overlook memory efficiency, leading to redundant DRAM traffic. We introduce STREAMINGGS, a fully streaming 3DGS algorithm-architecture co-design that achieves fine-grained pipelining and reduces DRAM traffic by transforming from a tile-centric rendering to a memory-centric rendering. Results show that our design achieves up to 45.7 × speedup and 62.9 × energy savings over mobile Ampere GPUs. Chenqi Zhang 0002, Yu Feng 0007, Jieru Zhao, Guangda Liu, Wenchao Ding 0001, Chentao Wu, Minyi Guo |
DAC | 2 |
| 2025 | VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM InferenceabstractVector quantization (VQ), which treats a vector as a compression unit, gains increasing research interests for its potential to accelerate large language models (LLMs). Compared to conventional element-wise quantization methods, VQ algorithms can compress weight and KV cache tensors in LLMs with a greater ratio while maintaining the high model accuracy. However, translating a VQ algorithm’s memory reduction into the actual latency improvement is challenging. We profile and analyze the current approach of integrating VQ into computation kernels and show that its major inefficiency lies in the poor access efficiency of codebooks in VQ algorithms and uncoordinated computation dataflow. Meanwhile, the diversity of VQ algorithms (e.g., different vector sizes and entry counts) and LLMs, computation kernels (e.g matrix-matrix/vector multiplication and attention computation) makes it impractical to manually craft efficient kernel implementations for each specific case. In this work, we design and implement VQ-LLM, an efficient fused VQ kernel generation framework. We first introduce a software abstraction called codebook cache to optimize codebook access efficiency and support the integration of VQ with various computations. The codebook cache adaptively stores different entries across the GPU’s memory hierarchy, including off-chip global memory, on-chip shared memory, and registers. Centered around the codebook cache, we design an efficient computation engine that optimizes memory traffic during computations involving codebooks. This compute engine adopts the codebook-centric dataflow and fusion optimizations. Additionally, we provide adaptive heuristics to tailor parameter selection in our optimizations to diverse VQ configurations. Our optimizations achieve the latency reduction of $\mathbf{6 4. 3 6 \%}$ to $\mathbf{9 9. 1 \%}$ compared to existing open-source implementations. A final comparison with state-of-the-art element-wise quantization methods like AWQ and QoQ shows that our VQ-LLM is practically viable, achieving latencies close or even better latencies to those at equivalent bit-widths, potentially offering greater accuracy. Zihan Liu 0002, Xinhao Luo, Junxian Guo, Wentao Ni, Yangjie Zhou 0001, Yue Guan 0003, Cong Guo 0003, Weihao Cui, Yu Feng 0007, Minyi Guo, Yuhao Zhu 0001, Minjia Zhang, Jingwen Leng |
HPCA | 9 |
| 2025 | M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical TypeabstractLarge language models (LLMs) are one of the most important killer computer applications. The recent algorithmic advancement proposes a fine-grained group-wise quantization for LLMs, which treats a small set (e.g., 64) of values in a tensor as a compression unit. It effectively preserves the model accuracy without retraining, and has become the standard approach to efficiently deploy LLMs. On the other hand, there are works that propose various adaptive data types to better adapt to different distributions and further reduce the required bit length for LLMs. In this work, our detailed analysis unveils a key finding that while different tensors exhibit similar distributions, small groups can have markedly different distributions. As such, the group-level diversity requires a new level of adaptivity for which existing adaptive data types fail to provide.In this paper, we propose MANT, a mathematically adaptive numeric type, featuring a more flexible encoding paradigm with a wider range of data distribution and more efficient decoding-computation fusion mechanism to address these challenges. Based on MANT, we develop a supporting framework to assign the appropriate data type for each group adaptively. Meanwhile, the dynamically generated Key-Value (KV) caches in LLMs introduce further complexity for real-time quantization. To tackle this, we propose an efficient real-time quantization mechanism. Besides, we implement a specific processing element (PE) to efficiently support MANT and incorporate a real-time quantization unit. By integrating these components into a systolic array, MANT unifies the group-wise weight and KV cache quantization and addresses the associated challenges. Our evaluation shows achieving, on average, 2.99 × (up to 4.46 ×) speedup and 2.81 × (up to 4.10 ×) energy reduction to the state-of-the-art LLM accelerator. Weiming Hu 0005, Cong Guo 0003, Yu Feng 0007, Renyang Guan, Zhendong Hua, Zihan Liu 0002, Yue Guan 0003, Minyi Guo, Jingwen Leng |
HPCA | 4 |
| 2025 | SLTarch: Towards Scalable Point-Based Neural Rendering by Taming Workload Imbalance and Memory IrregularityabstractRendering is critical in fields like 3D modeling, AR/VR, and autonomous driving, where high-quality, real-time output is essential. Point-based neural rendering (PBNR) offers a photorealistic and efficient alternative to conventional methods, yet it is still challenging to achieve real-time rendering on mobile platforms. We pinpoint two major bottlenecks in PBNR pipelines: LoD search and splatting. LoD search suffers from workload imbalance and irregular memory access, making it inefficient on off-the-shelf GPUs. Meanwhile, splatting introduces severe warp divergence across GPU threads due to its inherent sparsity.To tackle these challenges, we propose SLTarch, an algorithm-architecture co-designed framework. At its core, SLTarch introduces SLTree, a dedicated subtree-based data structure, and LTcore, a specialized hardware architecture tailored for efficient LoD search. Additionally, we co-design a divergence-free splatting algorithm with our simple yet principled hardware augmentation, SPcore, to existing PBNR accelerators. Compared to a mobile GPU, SLTarch achieves 3.9× speedup and 98% energy savings with negligible architecture overhead. Compared to existing accelerator designs, SLTarch achieves 1.8× speedup with 54% energy savings. Xingyang Li, Yu Feng 0007, Yiming Gan, Jieru Zhao, Zihan Liu 0002, Jingwen Leng, Minyi Guo |
ICCAD | 3 |
| 2025 | An Efficient Private GPT Never Autoregressively DecodesabstractThe wide deployment of the generative pre-trained transformer (GPT) has raised privacy concerns for both clients and servers. While cryptographic primitives can be employed for secure GPT inference to protect the privacy of both parties, they introduce considerable performance overhead. To accelerate secure inference, this study proposes a public decoding and secure verification approach that utilizes public GPT models, motivated by the observation that securely decoding one and multiple tokens takes a similar latency. The client uses the public model to generate a set of tokens, which are then securely verified by the private model for acceptance. The efficiency of our approach depends on the acceptance ratio of tokens proposed by the public model, which we improve from two aspects: (1) a private sampling protocol optimized for cryptographic primitives and (2) model alignment using knowledge distillation. Our approach improves the efficiency of secure decoding while maintaining the same level of privacy and generation quality as standard secure decoding. Experiments demonstrate a $2.1\times \sim 6.0\times$ speedup compared to standard decoding across three pairs of public-private models and different network conditions. Zhengyi Li 0002, Yue Guan 0003, Kang Yang 0002, Yu Feng 0007, Ning Liu 0007, Yu Yu 0001, Jingwen Leng, Minyi Guo |
ICML | 4 |
| 2025 | Lumina: Real-Time Neural Rendering by Exploiting Computational Redundancyabstract3D Gaussian Splatting (3DGS) has vastly advanced the pace of neural rendering, but it remains computationally demanding on today's mobile SoCs.To address this challenge, we propose Lumina, a hardware-algorithm co-designed system, which integrates two principal optimizations: a novel algorithm, S 2 , and a radiance caching mechanism, RC, to improve the efficiency of neural rendering.S 2 algorithm exploits temporal coherence in rendering to reduce the computational overhead, while RC leverages the color integration process of 3DGS to decrease the frequency of intensive rasterization computations.Coupled with these techniques, we propose an accelerator architecture, LuminCore, to further accelerate cache lookup and address the fundamental inefficiencies in Rasterization.We show that Lumina achieves 4.5× speedup and 5.3× energy reduction against a mobile Volta GPU, with a marginal quality loss (< 0.2 dB peak signal-to-noise ratio reduction) across synthetic and real-world datasets. Yu Feng 0007, Weikai Lin, Yuge Cheng, Zihan Liu 0002, Jingwen Leng, Minyi Guo, Chen Chen 0067, Shixuan Sun, Yuhao Zhu 0001 |
ISCA | 1 |
| 2025 | ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective PrimitiveabstractLarge language model (LLM) decoding suffers from high latency due to fragmented execution across operators and heavy reliance on off-chip memory for data exchange and reduction.
This execution model limits opportunities for fusion and incurs significant memory traffic and kernel launch overhead.
While modern architectures such as NVIDIA Hopper provide distributed shared memory and low-latency intra-cluster interconnects, they expose only low-level data movement instructions, lacking structured abstractions for collective on-chip communication.
To bridge this software-hardware gap, we introduce two cluster-level communication primitives, ClusterReduce and ClusterGather, which abstract common communication patterns and enable structured, high-speed data exchange and reduction between thread blocks within a cluster, allowing intermediate results to be on-chip without involving off-chip memory.
Building on these abstractions, we design ClusterFusion, an execution framework that schedules communication and computation jointly to expand operator fusion scope by composing decoding stages such as QKV Projection, Attention, and Output Projection into a single fused kernels.
Evaluations on H100 GPUs show that ClusterFusion outperforms state-of-the-art inference frameworks by $1.61\times$ on average in end-to-end latency across different models and configurations. Xinhao Luo, Zihan Liu 0002, Yangjie Zhou 0001, Shihan Fang, Yu Feng 0007, Chen Zhang 0001, Shixuan Sun, Zhenzhe Zheng 0001, Jingwen Leng, Minyi Guo |
NeurIPS | 6 |
| 2025 | PrivateEye: In-Sensor Privacy Preservation Through Optical Feature SeparationabstractWe address privacy issues in applications where images captured by an edge device (camera) are sent to the cloud for inference on utility tasks such as classification. Sending raw images to the cloud exposes them to data sniffing attacks and misuse by untrusted third-party service providers beyond the user's intended tasks. We propose an encoding scheme that not only evades direct visual inspection to the images or image reconstruction, but also prevents sensitive information from being ascertained. Unlike commonly used adversarial learning approaches, the proposed method is two-fold: first, it uses a diffractive optical neural network to spatially separate features corresponding to different tasks on the sensor plane in the optical domain. Then only the pixels corresponding to the utility task region are read. This encoding ensures that private features are never digitally stored on the edge device, thereby preventing privacy leakage. The proposed method successfully reduces the privacy retrieval in binary tasks with minimal accuracy loss (~ 2%) of the utility task, while reducing private task accuracy by ~ 35% and defending against reconstruction attacks with SSIM score of 0.43. Adith Boloor, Weikai Lin, Tianrui Ma, Yu Feng 0007, Yuhao Zhu 0001, Xuan Zhang 0001 |
WACV | 4 |
| 2025 | EDAS: Enabling Fast Data Loading for GPU Serverless ComputingabstractIntegrating GPUs into serverless computing platforms is crucial for improving efficiency. Many GPU functions, such as DNN inferences and scientific services, benefit from GPU usage, which requires only tens to hundreds of milliseconds for pure computation. Under these circumstances, fast data loading is imperative for function performance. However, existing GPU serverless systems face significant data stall issues, leading to extremely low GPU efficiency. Faced with the above problems, we observe opportunities to optimize data loading, such as data preloading and deduplicated data loading. However, these optimizations are impossible in existing GPU serverless systems due to the lack of insights into data information, such as data sizes and read-write attributes of function inputs. To address this, we propose a novel GPU serverless system, EDAS. EDAS first enhances user request specifications, allowing users to annotate data retrieved by GPU functions from the database with additional attributes. Based on this, EDAS takes over data loading from GPU functions and proposes two innovative data loading management schemes: a parallelized data loading scheme and a multi-stage resource exit scheme. Our experimental results show that EDAS reduces function duration by 16.2× and improves system throughput by 1.91× compared with the state-of-the-art serverless platform. Han Zhao 0005, Weihao Cui, Quan Chen 0002, Zijun Li 0001, Zhenhua Han, Yu Feng 0007, Jieru Zhao, Chen Chen 0067, Jingwen Leng, Minyi Guo |
ACM Trans. Archit. Code Optim. | 7 |
| 2024 | Amanda: Unified Instrumentation Framework for Deep Neural NetworksabstractThe success of deep neural networks (DNNs) has sparked efforts to analyze (e.g., tracing) and optimize (e.g., pruning) them. These tasks have specific requirements and ad-hoc implementations in current execution backends like TensorFlow/PyTorch, which require developers to manage fragmented interfaces and adapt their codes to diverse models. In this study, we propose a new framework called Amanda to streamline the development of these tasks. We formalize the implementation of these tasks as neural network instrumentation, which involves introducing instrumentation into the operator level of DNNs. This allows us to abstract DNN analysis and optimization tasks as instrumentation tools on various DNN models. We build Amanda with two levels of APIs to achieve a unified, extensible, and efficient instrumentation design. The user-level API provides a unified operator-grained instrumentation API for different backends. Meanwhile, internally, we design a set of callback-centric APIs for managing and optimizing the execution of original and instrumentation codes in different backends. Through these design principles, the Amanda framework can accommodate a broad spectrum of use cases, such as tracing, profiling, pruning, and quantization, across different backends (e.g., TensorFlow/PyTorch) and execution modes (graph/eager mode). Moreover, our efficient execution management ensures that the performance overhead is typically kept within 5%. Yue Guan 0003, Yuxian Qiu, Jingwen Leng, Fan Yang 0024, Shuo Yu 0006, Yunxin Liu 0001, Yu Feng 0007, Yuhao Zhu 0001, Lidong Zhou, Yun Liang 0001, Chen Zhang 0001, Chao Li 0009, Minyi Guo |
ASPLOS (1) | 7 |
| 2024 | JUNO: Optimizing High-Dimensional Approximate Nearest Neighbour Search with Sparsity-Aware Algorithm and Ray-Tracing Core MappingabstractApproximate nearest neighbor (ANN) search is a widely applied technique in modern intelligent applications, such as recommendation systems and vector databases. Therefore, efficient and high-throughput execution of ANN search has become increasingly important. In this paper, we first characterize the state-of-the-art product quantization-based method of ANN search and identify a significant source of inefficiency in the form of unnecessary pairwise distance calculations and accumulations. To improve efficiency, we propose Juno, an end-to-end ANN search system that adopts a carefully designed sparsity- and locality-aware search algorithm. We also present an efficient hardware mapping that utilizes ray tracing cores in modern GPUs with pipelined execution on tensor cores to execute our sparsity-aware ANN search algorithm. Our evaluations on four datasets from 1 to 100 million search points demonstrate 2.2×-8.5× improvements in search throughput. Moreover, our algorithmic enhancements alone achieve a maximal 2.6× improvement on the hardware without the acceleration of the RT core. Zihan Liu 0002, Wentao Ni, Jingwen Leng, Yu Feng 0007, Cong Guo 0003, Quan Chen 0002, Chao Li 0009, Minyi Guo, Yuhao Zhu 0001 |
ASPLOS (2) | 4 |
| 2024 | AutoVCoder: A Systematic Framework for Automated Verilog Code Generation using LLMsabstractRecently, the use of large language models (LLMs) for software code generation, e.g., C/C++ and Python, has proven a great success. However, LLMs still suffer from low syntactic and functional correctness when it comes to the generation of register-transfer level (RTL) code, such as Verilog. To address this issue, in this paper, we develop AutoVCoder, a systematic open-source framework that significantly improves the LLMs' correctness of generating Verilog code and enhances the quality of its output at the same time. Our framework integrates three novel techniques, including a high-quality hardware dataset generation approach, a two-round LLM fine-tuning method and a domain-specific retrieval-augmented generation (RAG) mechanism. Experimental results demonstrate that AutoVCoder outperforms both industrial and academic LLMs in Verilog code generation. Code and models are available at https://github.com/sjtu-zhao-lab/AutoVCoder. Mingzhe Gao, Jieru Zhao, Zhe Lin 0007, Wenchao Ding 0001, Xiaofeng Hou, Yu Feng 0007, Chao Li 0009, Minyi Guo |
ICCD | 6 |
| 2024 | Cicero: Addressing Algorithmic and Architectural Bottlenecks in Neural Rendering by Radiance Warping and Memory OptimizationsabstractNeural Radiance Field (NeRF) is widely seen as an alternative to traditional physically-based rendering. However, NeRF has not yet seen its adoption in resource-limited mobile systems such as Virtual and Augmented Reality (VR/AR), because it is simply extremely slow. On a mobile Volta GPU, even the state-of-the-art NeRF models generally execute only at 0.8 FPS. We show that the main performance bottlenecks are both algorithmic and architectural. We introduce, Cicero, to tame both forms of inefficiencies. We first introduce two algorithms, one fundamentally reduces the amount of work any NeRF model has to execute, and the other eliminates irregular DRAM accesses. We then describe an on-chip data layout strategy that eliminates SRAM bank conflicts. A pure software implementation of Cicero offers an $8.0 \times$ speed-up and $7.9 \times$ energy saving over a mobile Volta GPU. When compared to a baseline with a dedicated DNN accelerator, our speed-up and energy reduction increase to $28.2 \times$ and $37.8 \times$, respectively - all with minimal quality loss (less than 1.0 dB peak signal-to-noise ratio reduction). Yu Feng 0007, Zihan Liu 0002, Jingwen Leng, Minyi Guo, Yuhao Zhu 0001 |
ISCA | 1 |
| 2024 | BlissCam: Boosting Eye Tracking Efficiency with Learned In-Sensor Sparse SamplingabstractEye tracking is becoming an increasingly important task domain in emerging computing platforms such as Augmented/Virtual Reality (AR/VR). Today’s eye tracking system suffers from long end-to-end tracking latency and can easily eat up half of the power budget of a mobile VR device. Most existing optimization efforts exclusively focus on the computation pipeline by optimizing the algorithm and/or designing dedicated accelerators while largely ignoring the front-end of any eye tracking pipeline: the image sensor. This paper makes a case for co-designing the imaging system with the computing system. In particular, we propose the notion of “in-sensor sparse sampling”, whereby the pixels are drastically downsampled (by $20 \times$) within the sensor. Such in-sensor sampling enhances the overall tracking efficiency by significantly reducing 1) the power consumption of the sensor readout chain and sensor-host communication interfaces, two major power contributors, and 2) the work done on the host, which receives and operates on far fewer pixels. With careful reuse of existing pixel circuitry, our proposed BlissCam requires little hardware augmentation to support the in-sensor operations. Our synthesis results show up to $8.2 \times$ energy reduction and $1.4 \times$ latency reduction over existing eye tracking pipelines. Yu Feng 0007, Tianrui Ma, Yuhao Zhu 0001, Xuan Zhang 0001 |
ISCA | 1 |
| 2024 | Potamoi: Accelerating Neural Rendering via a Unified Streaming ArchitectureabstractNeural Radiance Field (NeRF) has emerged as a promising alternative for photorealistic rendering. Despite recent algorithmic advancements, achieving real-time performance on today’s resource-constrained devices remains challenging. In this article, we identify the primary bottlenecks in current NeRF algorithms and introduce a unified algorithm-architecture co-design, Potamoi , designed to accommodate various NeRF algorithms. Specifically, we introduce a runtime system featuring a plug-and-play algorithm, SpaRW , which significantly reduces the per-frame computational workload and alleviates compute inefficiencies. Furthermore, our unified streaming pipeline coupled with customized hardware support effectively tames both SRAM and DRAM inefficiencies by minimizing repetitive DRAM access and completely eliminating SRAM bank conflicts. When evaluated against a baseline utilizing a dedicated DNN accelerator, our framework demonstrates a speedup and energy reduction of 53.1× and 67.7×, respectively, all while maintaining high visual quality with less than a 1.0 dB reduction in peak signal-to-noise ratio. Yu Feng 0007, Weikai Lin, Zihan Liu 0002, Jingwen Leng, Minyi Guo, Han Zhao 0005, Xiaofeng Hou, Jieru Zhao, Yuhao Zhu 0001 |
ACM Trans. Archit. Code Optim. | 1 |
| 2023 | Invited Paper: Learned In-Sensor Visual Computing: From Compression to EventificationabstractVisual computing is vital for numerous applications. In conventional visual computing systems, CMOS image sensors (CIS) act as pure imaging devices for capturing images, however, recent CIS designs increasingly integrate processing capabilities such as Deep Neural Networks (DNN), which give rise to a notion of in-sensor computing. In this paper, we propose a new concept, learned in-sensor visual computing, which exploits end-to-end optimization of in-sensor processing and downstream vision tasks to achieve better overall algorithm accuracy and adopts hardware/algorithm co-design to achieve ultra-low sensor energy consumption. Two examples of the learned in-sensor visual computing, Leca and EDGAzE, are demonstrated. Yu Feng 0007, Tianrui Ma, Adith Boloor, Yuhao Zhu 0001, Xuan Zhang 0001 |
ICCAD | 1 |
| 2023 | CAMJ: Enabling System-Level Energy Modeling and Architectural Exploration for In-Sensor Visual ComputingabstractCMOS Image Sensors (CIS) are fundamental to emerging visual computing applications. While conventional CIS are purely imaging devices for capturing images, increasingly CIS integrate processing capabilities such as Deep Neural Network (DNN). Computational CIS expand the architecture design space, but to date no comprehensive energy model exists. This paper proposes CamJ, a detailed energy modeling framework that provides a component-level energy breakdown for computational CIS and is validated against nine recent CIS chips. We use CamJ to demonstrate three use-cases that explore architectural trade-offs including computing in vs. off CIS, 2D vs. 3D-stacked CIS design, and analog vs. digital processing inside CIS. The code of CamJ is available at: https://github.com/horizon-research/CamJ. Tianrui Ma, Yu Feng 0007, Xuan Zhang 0001, Yuhao Zhu 0001 |
ISCA | 2 |
| 2023 | Fast and Accurate: Video Enhancement Using Sparse DepthabstractThis paper presents a general framework to build fast and accurate algorithms for video enhancement tasks such as super-resolution, deblurring, and denoising. Essential to our framework is the realization that the accuracy, rather than the density, of pixel flows is what is required for high-quality video enhancement. Most of prior works take the opposite approach: they estimate dense (per-pixel)—but generally less robust—flows, mostly using computationally costly algorithms. Instead, we propose a lightweight flow estimation algorithm; it fuses the sparse point cloud data and (even sparser and less reliable) IMU data available in modern autonomous agents to estimate the flow information. Building on top of the flow estimation, we demonstrate a general framework that integrates the flows in a plug-andplay fashion with different task-specific layers. Algorithms built in our framework achieve 1.78× — 187.41× speedup while providing a 0.42 dB – 6.70 dB quality improvement over competing methods. Yu Feng 0007, Patrick Hansen, Paul N. Whatmough, Guoyu Lu 0001, Yuhao Zhu 0001 |
WACV | 1 |
| 2022 | Crescent: taming memory irregularities for accelerating deep point cloud analyticsabstract3D perception in point clouds is transforming the perception ability of future intelligent machines. Point cloud algorithms, however, are plagued by irregular memory accesses, leading to massive inefficiencies in the memory sub-system, which bottlenecks the overall efficiency. Yu Feng 0007, Gunnar Hammonds, Yiming Gan, Yuhao Zhu 0001 |
ISCA | 1 |
| 2022 | Real-Time Gaze Tracking with Event-Driven Eye SegmentationabstractGaze tracking is increasingly becoming an essential component in Augmented and Virtual Reality. Modern gaze tracking algorithms are heavyweight; they operate at most 5 Hz on mobile processors despite that near-eye cameras comfortably operate at a real-time rate (> 30 Hz). This paper presents a real-time eye tracking algorithm that, on average, operates at 30 Hz on a mobile processor, achieves 0.1°–0.5° gaze accuracies, all the while requiring only 30K parameters, one to two orders of magnitude smaller than state-of-the-art eye tracking algorithms. The crux of our algorithm is an Auto ROI mode, which continuously predicts the Regions of Interest (ROIs) of near-eye images and judiciously processes only the ROIs for gaze estimation. To that end, we introduce a novel, lightweight ROI prediction algorithm by emulating an event camera. We discuss how a software emulation of events enables accurate ROI prediction without requiring special hardware. The code of our paper is available at https://github.com/horizon-research/edgaze. Yu Feng 0007, Nathan Goulding, Hans Reyserhove, Yuhao Zhu 0001 |
VR | 1 |
| 2020 | Real-Time Spatio-Temporal LiDAR Point Cloud CompressionabstractCompressing massive LiDAR point clouds in real-time is critical to autonomous machines such as drones and self-driving cars. While most of the recent prior work has focused on compressing individual point cloud frames, this paper proposes a novel system that effectively compresses a sequence of point clouds. The idea to exploit both the spatial and temporal redundancies in a sequence of point cloud frames. We first identify a key frame in a point cloud sequence and spatially encode the key frame by iterative plane fitting. We then exploit the fact that consecutive point clouds have large overlaps in the physical space, and thus spatially encoded data can be (re-)used to encode the temporal stream. Temporal encoding by reusing spatial encoding data not only improves the compression rate, but also avoids redundant computations, which significantly improves the compression speed. Experiments show that our compression system achieves 40× to 90× compression rate, significantly higher than the MPEG's LiDAR point cloud compression standard, while retaining high end-to-end application accuracies. Meanwhile, our compression system has a compression speed that matches the point cloud generation rate by today LiDARs and out-performs existing compression systems, enabling real-time point cloud transmission. Yu Feng 0007, Shaoshan Liu, Yuhao Zhu 0001 |
IROS | 1 |
| 2020 | Mesorasi: Architecture Support for Point Cloud Analytics via Delayed-AggregationabstractPoint cloud analytics is poised to become a key workload on battery-powered embedded and mobile platforms in a wide range of emerging application domains, such as autonomous driving, robotics, and augmented reality, where efficiency is paramount. This paper proposes Mesorasi, an algorithm-architecture co-designed system that simultaneously improves the performance and energy efficiency of point cloud analytics while retaining its accuracy.Our extensive characterizations of state-of-the-art point cloud algorithms show that, while structurally reminiscent of convolutional neural networks (CNNs), point cloud algorithms exhibit inherent compute and memory inefficiencies due to the unique characteristics of point cloud data. We propose delayed-aggregation, a new algorithmic primitive for building efficient point cloud algorithms. Delayed-aggregation hides the performance bottlenecks and reduces the compute and memory redundancies by exploiting the approximately distributive property of key operations in point cloud algorithms. Delayed-aggregation let point cloud algorithms achieve 1.6× speedup and 51.1% energy reduction on a mobile GPU while retaining the accuracy (-0.9% loss to 1.2% gains). To maximize the algorithmic benefits, we propose minor extensions to contemporary CNN accelerators, which can be integrated into a mobile Systems-on-a-Chip (SoC) without modifying other SoC components. With additional hardware support, Mesorasi achieves up to 3.6× speedup. Yu Feng 0007, Boyuan Tian, Tiancheng Xu, Paul N. Whatmough, Yuhao Zhu 0001 |
MICRO | 1 |
| 2019 | PES: proactive event scheduling for responsive and energy-efficient mobile web computingabstractWeb applications are gradually shifting toward resource-constrained mobile devices. As a result, the Web runtime system must simultaneously address two challenges: responsiveness and energy-efficiency. Conventional Web runtime systems fall short due to their reactive nature: they react to a user event only after it is triggered. The reactive strategy leads to local optimizations that schedule event executions one at a time, missing global optimization opportunities. Yu Feng 0007, Yuhao Zhu 0001 |
ISCA | 1 |
| 2019 | ASV: Accelerated Stereo Vision SystemabstractEstimating depth from stereo vision cameras, i.e., "depth from stereo", is critical to emerging intelligent applications deployed in energy- and performance-constrained devices, such as augmented reality headsets and mobile autonomous robots. While existing stereo vision systems make trade-offs between accuracy, performance and energy-efficiency, we describe ASV, an accelerated stereo vision system that simultaneously improves both performance and energy-efficiency while achieving high accuracy. Yu Feng 0007, Paul N. Whatmough, Yuhao Zhu 0001 |
MICRO | 1 |