Weikai Lin

dblp:156/9803 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0003-3537-4857ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 StreamGrid: Streaming Point Cloud Analytics via Compulsory Splitting and Deterministic Termination
abstract
Point clouds are increasingly important in intelligent applications, but frequent off-chip memory traffic in accelerators causes pipeline stalls and leads to high energy consumption. While conventional line buffer techniques can eliminate off-chip traffic, they cannot be directly applied to point clouds due to their inherent computation patterns. To address this, we introduce two techniques: compulsory splitting and deterministic termination, enabling fully-streaming processing. We further propose StreamGrid, a framework that integrates these techniques and automatically optimizes on-chip buffer sizes. Our evaluation shows StreamGrid reduces on-chip memory by 61.3% and energy consumption by 40.5% with marginal accuracy loss compared to the baselines without our techniques. Additionally, we achieve 10.0× speedup and 3.9× energy efficiency over state-of-the-art accelerators.
Yu Feng 0007, Zheng Liu 0022, Weikai Lin, Zihan Liu 0002, Jingwen Leng, Minyi Guo, Zhezhi He, Jieru Zhao, Yuhao Zhu 0001
ASPLOS (2)3
2025 MetaSapiens: Real-Time Neural Rendering with Efficiency-Aware Pruning and Accelerated Foveated Rendering
abstract
Point-Based Neural Rendering (PBNR) is emerging as a promising class of rendering techniques, which are permeating all aspects of society, driven by a growing demand for real-time, photorealistic rendering in AR/VR and digital twins. Achieving real-time PBNR on mobile devices is challenging.
Weikai Lin, Yu Feng 0007, Yuhao Zhu 0001
ASPLOS (1)1
2025 SNAPPIX: Efficient-Coding-Inspired In-Sensor Compression for Edge Vision
abstract
Energy-efficient image acquisition on the edge is crucial for enabling remote sensing applications where the sensor node has weak compute capabilities and must transmit data to a remote server/cloud for processing. To reduce the edge energy consumption, this paper proposes a sensor-algorithm co-designed system called SNAPPIX, which compresses raw pixels in the analog domain inside the sensor. We use coded exposure (CE) as the in-sensor compression strategy as it offers the flexibility to sample, i.e., selectively expose pixels, both spatially and temporally. SnapPix has three contributions. First, we propose a task-agnostic strategy to learn the sampling/exposure pattern based on the classic theory of efficient coding. Second, we codesign the downstream vision model with the exposure pattern to address the pixel-level non-uniformity unique to CE-compressed images. Finally, we propose lightweight augmentations to the image sensor hardware to support our in-sensor CE compression. Evaluating on action recognition and video reconstruction, SnapPix outperforms state-of-the-art video-based methods at the same speed while reducing the energy by up to $15.4 \times$. We have open-sourced the code at: https://github.com/horizonresearch/SnapPix.
Weikai Lin, Tianrui Ma, Adith Boloor, Yu Feng 0007, Ruofan Xing, Xuan Zhang 0001, Yuhao Zhu 0001
DAC1
2025 Lumina: Real-Time Neural Rendering by Exploiting Computational Redundancy
abstract
3D Gaussian Splatting (3DGS) has vastly advanced the pace of neural rendering, but it remains computationally demanding on today's mobile SoCs.To address this challenge, we propose Lumina, a hardware-algorithm co-designed system, which integrates two principal optimizations: a novel algorithm, S 2 , and a radiance caching mechanism, RC, to improve the efficiency of neural rendering.S 2 algorithm exploits temporal coherence in rendering to reduce the computational overhead, while RC leverages the color integration process of 3DGS to decrease the frequency of intensive rasterization computations.Coupled with these techniques, we propose an accelerator architecture, LuminCore, to further accelerate cache lookup and address the fundamental inefficiencies in Rasterization.We show that Lumina achieves 4.5× speedup and 5.3× energy reduction against a mobile Volta GPU, with a marginal quality loss (< 0.2 dB peak signal-to-noise ratio reduction) across synthetic and real-world datasets.
Yu Feng 0007, Weikai Lin, Yuge Cheng, Zihan Liu 0002, Jingwen Leng, Minyi Guo, Chen Chen 0067, Shixuan Sun, Yuhao Zhu 0001
ISCA2
2025 PowerGS: Display-Rendering Power Co-Optimization for Neural Rendering in Power-Constrained XR Systems
abstract
3D Gaussian Splatting (3DGS) combines classic image-based rendering, point-based graphics, and modern differentiable techniques, and offers an interesting alternative to traditional physically-based rendering. 3DGS-family models are far from efficient for power-constrained Extended Reality (XR) devices, which need to operate at a Watt-level. This paper introduces PowerGS, the first framework to jointly minimize the rendering and display power in 3DGS under a quality constraint. We present a general problem formulation and show that solving the problem amounts to 1) identifying the iso-quality curve(s) in the landscape subtended by the display and rendering power and 2) identifying the power-minimal point on a given curve, which has a closed-form solution given a proper parameterization of the curves. PowerGS also readily supports foveated rendering for further power savings. Extensive experiments and user studies show that PowerGS achieves up to 86% total power reduction compared to state-of-the-art 3DGS models, with minimal loss in both subjective and objective quality. Code is available at https://github.com/horizon-research/PowerGS.
Weikai Lin, Sushant Kondguli, Carl S. Marshall, Yuhao Zhu 0001
SIGGRAPH Asia1
2025 PrivateEye: In-Sensor Privacy Preservation Through Optical Feature Separation
abstract
We address privacy issues in applications where images captured by an edge device (camera) are sent to the cloud for inference on utility tasks such as classification. Sending raw images to the cloud exposes them to data sniffing attacks and misuse by untrusted third-party service providers beyond the user's intended tasks. We propose an encoding scheme that not only evades direct visual inspection to the images or image reconstruction, but also prevents sensitive information from being ascertained. Unlike commonly used adversarial learning approaches, the proposed method is two-fold: first, it uses a diffractive optical neural network to spatially separate features corresponding to different tasks on the sensor plane in the optical domain. Then only the pixels corresponding to the utility task region are read. This encoding ensures that private features are never digitally stored on the edge device, thereby preventing privacy leakage. The proposed method successfully reduces the privacy retrieval in binary tasks with minimal accuracy loss (~ 2%) of the utility task, while reducing private task accuracy by ~ 35% and defending against reconstruction attacks with SSIM score of 0.43.
Adith Boloor, Weikai Lin, Tianrui Ma, Yu Feng 0007, Yuhao Zhu 0001, Xuan Zhang 0001
WACV2
2024 OW3Det: Toward Open-World 3D Object Detection for Autonomous Driving
abstract
Despite their success in LIDAR object detection, modern detectors are vulnerable to uncommon instances and corner cases (e.g., a runaway tire) since they are closed-set and static. Networks under the closed-set setup only predict labels of seen classes, while static models suffer from catastrophic forgetting when gradually learning novel concepts. This motivates us to formulate the open-world 3D object detection task for autonomous driving, which aims to 1) tackle the closed-set issue by identifying unseen instances as unknown and 2) incrementally learn novel classes without forgetting previously obtained knowledge. To achieve the open-world objectives, we propose Open-World 3D Detector (OW3Det), the first framework for open-world 3D object detection. The OW3Det comprises a base detector, a self-supervised unknown identifier, and a knowledge-distillation-restricted incremental learner. Although knowledge distillation facilitates preserving memories, imposing penalties on areas containing unknown objects hinders the incremental learning process. We mitigate this hindrance by employing unknown-driven pivotal mask, which eliminates unnecessary restrictions on regions overlapping with novel instances. Abundant experiments and visualizations demonstrate that the proposed OW3Det attains state-of-the-art performance.
Wenfei Hu, Weikai Lin, Hongyu Fang, Dingsheng Luo
IROS2
2024 Potamoi: Accelerating Neural Rendering via a Unified Streaming Architecture
abstract
Neural Radiance Field (NeRF) has emerged as a promising alternative for photorealistic rendering. Despite recent algorithmic advancements, achieving real-time performance on today’s resource-constrained devices remains challenging. In this article, we identify the primary bottlenecks in current NeRF algorithms and introduce a unified algorithm-architecture co-design, Potamoi , designed to accommodate various NeRF algorithms. Specifically, we introduce a runtime system featuring a plug-and-play algorithm, SpaRW , which significantly reduces the per-frame computational workload and alleviates compute inefficiencies. Furthermore, our unified streaming pipeline coupled with customized hardware support effectively tames both SRAM and DRAM inefficiencies by minimizing repetitive DRAM access and completely eliminating SRAM bank conflicts. When evaluated against a baseline utilizing a dedicated DNN accelerator, our framework demonstrates a speedup and energy reduction of 53.1× and 67.7×, respectively, all while maintaining high visual quality with less than a 1.0 dB reduction in peak signal-to-noise ratio.
Yu Feng 0007, Weikai Lin, Zihan Liu 0002, Jingwen Leng, Minyi Guo, Han Zhao 0005, Xiaofeng Hou, Jieru Zhao, Yuhao Zhu 0001
ACM Trans. Archit. Code Optim.2
2023 Learning Clear Class Separation for Open-set 3D Detector in Autonomous Vehicle via Selective Forgetting
abstract
A trustworthy 3D detector is essential in the perception system of autonomous vehicles, ensuring accurate detection of their surroundings. However, autonomous vehicles have to operate in ever-changing real-world driving scenes, where unknown objects that do not belong to the training set are commonly encountered. Confusion about known and unknown objects could result in severe and dangerous consequences for road safety. To address this problem, we improve the reliability of autonomous driving systems by formulating open-set 3D object detection task. An Open-set 3D Detector (Open3Det) is proposed to reject unknown instances while maintaining performance on known categories. Distinct from 2D objects, clear space separation exists between each 3D instance. Motivated by this, we propose selective forgetting, a novel method capable of filtering out misleading predictions. Given a close-set teacher model, knowledge distillation is introduced to build a open-set student model. The student model preserves its predictions for known objects, whereas predictions of backgrounds and unknown instances are discarded to minimize misleading results. Extensive experiments and visualizations reveal the efficacy of the proposed method.
Wenfei Hu, Weikai Lin, Hongyu Fang, Dingsheng Luo
RO-MAN2
2023 Learning Clear Class Separation for Open-set 3D Detector in Autonomous Vehicle via Selective Forgetting
abstract
A trustworthy 3D detector is essential in the perception system of autonomous vehicles, ensuring accurate detection of their surroundings. However, autonomous vehicles have to operate in ever-changing real-world driving scenes, where unknown objects that do not belong to the training set are commonly encountered. Confusion about known and unknown objects could result in severe and dangerous consequences for road safety. To address this problem, we improve the reliability of autonomous driving systems by formulating open-set 3D object detection task. An Open-set 3D Detector (Open3Det) is proposed to reject unknown instances while maintaining performance on known categories. Distinct from 2D objects, clear space separation exists between each 3D instance. Motivated by this, we propose selective forgetting, a novel method capable of filtering out misleading predictions. Given a close-set teacher model, knowledge distillation is introduced to build a open-set student model. The student model preserves its predictions for known objects, whereas predictions of backgrounds and unknown instances are discarded to minimize misleading results. Extensive experiments and visualizations reveal the efficacy of the proposed method.
Wenfei Hu, Weikai Lin, Hongyu Fang, Dingsheng Luo
RO-MAN2
2016 X-ray computed tomography using sparsity based regularization
Li Liu 0024, Weikai Lin, Mingwu Jin
Neurocomputing2