VLDB 2026 Research / reviewers in the wild / expert
Xin Lou 0001
dblp:133/5204-1
· DBLP profile ↗
44ranked-venue papers
3as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 3 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalarium: A Unified Scala-based Co-Simulation Framework for Agile Chip Development
Yuefeng Zhang, Wenkai Zhou, Binzhe Yuan, Junsheng Chen, Xiangyu Zhang 0002, Hao Geng, Xin Lou 0001 |
ASP-DAC | 8 |
| 2026 | A High-Performance Neural Rendering Accelerator Based on Novel Multi-Level Ray Scheduling and Dual-Process BackendabstractNeural rendering enables photorealistic scene re-construction but remains difficult to deploy on edge devices due to intensive computation, redundant sampling, and memory bandwidth constraints. This work presents a high-performance neural rendering accelerator for real-time embedded rendering. The proposed design integrates: (1) a dual-process backend with fused micro-MLPs to significantly improve sample processing efficiency, (2) multi-resolution spatial partitioning with adaptive ray clustering to exploit sparsity and achieve over 95% cache hit rate, and (3) a multi-level scheduling framework with proactive prefetching to reduce MLP stalls. Implemented on FPGA, the prototype achieves 94.7 FPS at 800×800 resolution with 6.4 W power consumption. An ASIC implementation in 28 nm technology sustains 440 FPS at 268 mW. Experimental results demonstrate state-of-the-art performance and energy efficiency while preserving rendering quality above 30 dB PSNR. Wenkai Zhou, Yuefeng Zhang, Binzhe Yuan, Junsheng Chen, Luntian Zhang, Xiangyu Zhang 0002, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001 |
DATE | 10 |
| 2026 | ZeroBlade: A Spatial Similarity-aware HiSparse MLP Engine for Neural Volume Rendering
Antong Li, Haochuan Wan, Xiangyu Zhang 0002, Xin Lou 0001 |
ISCAS | 4 |
| 2026 | An Efficient Low-Light Object Detection Framework based on Task-Driven Distillation
Wei Zhou 0037, Kangjie Long, Cong Pang, Xiangyu Zhang 0002, Xin Lou 0001 |
ISCAS | 6 |
| 2026 | SCOPE-3D: An Energy Efficient Accelerator for Implicit Neural Representation-based Sparse-view Computed Tomography Reconstruction
Haochuan Wan, Xin Li 0245, Qing Wu 0001, Yuhan Gu, Wenyan Su, Yuyao Zhang 0005, Xin Lou 0001 |
ISCAS | 8 |
| 2026 | WaveMamba: A Vision Backbone Synergizing Feature Extraction and Downsampling
Haitian Yang, Xiangyu Zhang 0002, Yuanmei Zhang, Xin Lou 0001, Wei Zhou 0037 |
ISCAS | 4 |
| 2026 | AlignLite: A Lightweight Framework for Weakly Aligned Multimodal Object Detection
Haitian Yang, Xiangyu Zhang 0002, Yuanmei Zhang, Xin Lou 0001, Wei Zhou 0037 |
ISCAS | 4 |
| 2026 | Unity-EDR: A Hybrid Neural-Mesh Rendering System for Efficient Visual Synthesis
Chaolin Rao, Xiangyu Zhang 0002, Xin Lou 0001 |
ISCAS | 7 |
| 2026 | Terafly: A Multinode FPGA-Based Accelerator Design for Efficient Cooperative Inference in LLMsabstractIn this paper, we propose Terafly, a multi-node accelerator design tailored for efficient Large Language Model (LLM) deployment and inference. Conventional accelerator architectures struggle to effectively handle both the prefill and decode stages during inference. To address this limitation, we introduce a hybrid spatial-temporal architecture that combines the high-throughput advantages of spatial architectures with the flexibility of temporal architectures, enabling it to accommodate the diverse inference patterns of LLMs. In addition, we propose a generation framework to streamline the customization of our LLM-friendly design for various deployment scenarios. Within this framework, users can specify their requirements such as model type, target platform, and performance goals. The framework then generates multiple accelerator nodes and maps them to distinct Super Logic Regions (SLRs) within a single FPGA, enabling cooperative inference under a model parallelism scheme. Through experiments, our generated accelerator can be easily deployed on both Alveo U250 and U50lv cards, serving models ranging from OPT-350M to OPT-1.3B under various performance settings. Notably, when running OPT-1.3B using the generated dual-node accelerator on a single Alveo U50lv card, we achieve an average 1.1x speed-up and a 3.4x improvement in energy efficiency compared to the Nvidia A100 GPU. Jianing Zheng, Gang Chen 0023, Libo Huang 0002, Xin Lou 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | Duplex-GS: Proxy-Guided Weighted Blending for Real-Time Order-Independent Gaussian Splattingabstract3D Gaussian Splatting (3DGS) achieves photorealistic rendering but requires global sorting for α-blending, causing noticeable “popping” artifacts and hindering deployment on edge devices. While sort-free Order-Independent Transparency (OIT) methods circumvent sorting, they introduce “transparency” artifacts and suffer from inefficiency due to the absence of physical constraints. To address these limitations, we present Duplex-GS, a dual-hierarchy framework leveraging proxy-guided spatial organization and a novel hybrid renderer that combines α-blending with reformulated Weighted-Sum Rendering (WSR). We introduce explicit ellipsoidal cell proxies to encapsulate local Gaussians, which enables efficient proxy-level rasterization. This strategy drastically reduces the overhead associated with global sorting. Furthermore, we propose a physically grounded WSR scheme with cell-level early termination, which restores the physical constraints absent in prior OIT-based 3DGS methods, effectively eliminating both popping and transparency artifacts. Extensive experiments on diverse real-world benchmarks demonstrate the effectiveness of the OIT-based paradigm for 3DGS, enabled by a practical dual-hierarchy implementation. Quantitatively, our method delivers high-fidelity real-time rendering, outperforming prior OIT-based 3DGS methods by 1.5×– 4× in speed, while reducing radix-sort cost by 29.8%– 86.9% compared with conventional α-blending without compromising visual quality. Project page: https://duplexgs.github.io/. Jingyi Yu 0001, Xin Lou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | A Real-Time Neural Representation via Algorithm-Hardware Synergy for Sparse-View CT ReconstructionabstractSparse-view computed tomography (SVCT) is an advancement in computed tomography (CT) technology that aims to reduce the radiation dose during imaging. Reconstructing high-quality images from sparse-view (SV) projections is an ill-posed inverse problem. Recently, implicit neural representations (INRs) as a self-supervised paradigm for solving underdetermined inverse problems have demonstrated excellent performance in SVCT reconstruction. However, since INR-based approaches rely on subject-specific training, they require a significant investment of time to optimize from scratch. Consequently, previous INR methods have not been able to meet the requisite timeliness of reconstruction. In our work, we propose RTSyner, an algorithm-hardware collaboration framework that facilitates the real-time efficiency of CT reconstruction. On the algorithmic side, we introduce an efficient coordinate-based feature module that exploits the local latent features as a positional external condition, leveraging the limited structural information of corrupted images derived from the sensory domain. By fusing latent features and coordinate information, the model learns a neural representation of the final tomographic image. On the hardware side, we design a dedicated hardware architecture with a customized algorithm flow to improve reconstruction speed and reduce power consumption. Furthermore, we improve the efficiency of model inference through model quantization, which also facilitates the subsequent deployment of hardware. Our extensive experimental results demonstrate that the RTSyner based on neural representation has achieved real-time SVCT reconstruction through the synergistic acceleration of the algorithm and hardware. We further explore its application potential via volume reconstructions under more complex acquisition geometries. Xin Li 0245, Haochuan Wan, Kangjie Long, Qing Wu 0001, Chenhe Du, Xin Lou 0001, Yuyao Zhang 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2026 | An Energy-Efficient Edge Coprocessor for Neural Rendering With Explicit Data Reuse StrategiesabstractNeural radiance fields (NeRFs) have transformed 3-D reconstruction and rendering, facilitating photorealistic image synthesis from sparse viewpoints. This work introduces an explicit data reuse neural rendering (EDR-NR) architecture, which reduces frequent external memory accesses (EMAs) and cache misses by exploiting the spatial locality from three phases, including rays, ray packets (RPs), and samples. The EDR-NR architecture features a four-stage scheduler that clusters rays on the basis of$Z$-order, prioritize lagging rays when ray divergence happens, reorders RPs based on spatial proximity, and issues samples out-of-orderly (OoO) according to the availability of on-chip feature data. In addition, a four-tier hierarchical RP marching (HRM) technique is integrated with an axis-aligned bounding box (AABB) to facilitate spatial skipping (SS), reducing redundant computations and improving throughput. Moreover, a balanced allocation strategy for feature storage is proposed to mitigate SRAM bank conflicts. Fabricated using a 40-nm process with a die area of 10.5 mm2, the EDR-NR chip demonstrates a$2.41\times $enhancement in normalized energy efficiency, a$1.21\times $improvement in normalized area efficiency, a$1.20\times $increase in normalized throughput, and a 53.42% reduction in on-chip SRAM consumption compared with state-of-the-art accelerators. Binzhe Yuan, Xiangyu Zhang 0002, Yuefeng Zhang, Haochuan Wan, Zhechen Yuan, Junsheng Chen, Yunxiang He, Junran Ding, Chaolin Rao, Wenyan Su, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 15 |
| 2025 | CityGo: Lightweight Urban Modeling and Rendering with Proxy Buildings and Residual GaussiansabstractAccurate and efficient modeling of large-scale urban scenes is critical for applications such as AR navigation, UAV-based inspection, and smart city digital twins. While aerial imagery offers broad coverage and complements limitations of ground-based data, reconstructing city-scale environments from such views remains challenging due to occlusions, incomplete geometry, and high memory demands. Recent advances like 3D Gaussian Splatting (3DGS) improve scalability and visual quality but remain limited by dense primitive usage, long training times, and poor suitability for edge devices. We propose CityGo, a hybrid framework that combines textured proxy geometry with residual and surrounding 3D Gaussians for lightweight, photorealistic rendering of urban scenes from aerial perspectives. Our approach first extracts compact building proxy meshes from MVS point clouds, then uses zero-order SH Gaussians to generate occlusion-free textures via image-based rendering and back-projection. To capture high-frequency details, we introduce residual Gaussians placed based on proxy-photo discrepancies and guided by depth priors. Broader urban context is represented by surrounding Gaussians, with importance-aware downsampling applied to non-critical regions to reduce redundancy. A tailored optimization strategy jointly refines proxy textures and Gaussian parameters, enabling real-time rendering of complex urban scenes on mobile GPUs with significantly reduced training and memory requirements. Extensive experiments on real-world aerial datasets demonstrate that our hybrid representation achieves fastest training speed, while delivering comparable visual fidelity to pure 3D Gaussian Splatting approaches. Furthermore, CityGo enables real-time rendering of large-scale urban scenes on mobile consumer GPUs, with substantially reduced memory usage and energy consumption. Yuhui Zhong, Jiadi Cui, Honglong Zhang, Lan Xu 0003, Xin Lou 0001, Yujiao Shi 0002, Jingyi Yu 0001, Yingliang Zhang |
SIGGRAPH Asia | 8 |
| 2025 | CoARF++: Content-Aware Radiance Field Aligning Model Complexity With Scene IntricacyabstractThis paper introduces the concept of Content-Aware Radiance Fields (CoARF), which adaptively aligns the model complexity with the scene intricacy. By examining the intricacies of radiance fields from three perspectives, model complexity is adapted through scalable feature grids, dynamic neural networks, and model quantization. Specifically, we propose a hash collision detection mechanism that removes redundant feature grid by restricting the valid hash collision to reasonable level, making the space complexity scalable. We introduce an uncertainty-aware decoded layer, where simple points are early-exited to prevent them from being processed by deeper network layers, ensuring computational complexity scalable. Furthermore, we propose Learned Bitwidth Quantization (LBQ) and Adversarial Content-Aware Quantization (A-CAQ) paradigms by making the bitwidth of parameters differentiable and trainable, allowing for adjustable quantization schemes. Building on these techniques, the proposed CoARF++ framework enables a scalable pipeline for radiance fields that is tailored to the unique characteristics of scene complexity and quality requirement. Extensive experiments demonstrate a significant and adjustable reduction in model complexity across various NeRF variants, while maintaining the necessary reconstruction and rendering quality, making it advantageous for the practical deployment of radiance field models. Xue Xian Zheng, Tareq Y. Al-Naffouri, Jingyi Yu 0001, Xin Lou 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | A Neural Rendering Coprocessor With Optimized Ray Representation and MarchingabstractNeural rendering, a transformative approach for 3-D scene reconstruction and rendering, has advanced rapidly in recent years. This article introduces an energy-efficient neural rendering coprocessor that implements the popular and widely used instant neural graphics primitive (Instant-NGP) algorithm. In particular, we address the challenges of limited resources for deploying Instant-NGP on edge by proposing a dedicated architecture, which incorporates three main innovations: 1) we optimize occupancy grid queries in the ray marching module by partitioning the grid and decoupling the query process from sampling point generation, which improves both efficiency and memory usage; 2) we introduce a bilinked list-based ray switching strategy, which ensures continuous pipeline utilization to overcome the inefficiencies caused by sequential processing; and 3) we optimize the hash encoding process by incorporating quantization-aware training (QAT), enabling the hash table to fit into on-chip memory, thereby improving performance on resource-constrained devices. To demonstrate the effectiveness of our architecture, we design and fabricate a proof-of-concept chip using 40-nm CMOS technology and develop a testing system to evaluate its performance. Measurement results validate the advantages of the proposed design, showing that our chip achieves superior energy efficiency compared to both server and edge graphics processing units (GPUs), as well as other state-of-the-art neural rendering chip designs. Zhechen Yuan, Binzhe Yuan, Chaolin Rao, Yiren Zhu, Yunxiang He, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2024 | ZeroTetris: A Spacial Feature Similarity-based Sparse MLP Engine for Neural Volume RenderingabstractNeural Volume Rendering (NVR), a novel paradigm for the longstanding problem of photo-realistic rendering of virtual worlds, has developed explosively in the past three years. The unique and substantial computational requirements of NVR pose challenge on deploying NVR to existing dedicated accelerator for neural networks. In this work, we propose ZeroTetris, a spacial feature similarity-based sparse multilayer perceptron (MLP) hardware accelerator for NVR. By leveraging the unique similarity-based sparsity between adjacent sampling points in NVR models, ZeroTetris efficiently bypass the computation of zero activations, thereby enhancing energy efficiency. Evaluation results affirm the effectiveness of the proposed design, showcasing ZeroTetris's superior performance in both area and power efficiency compared to other dedicated sparse matrix multiplication or MLP accelerator designs. Haochuan Wan, Linjie Ma, Antong Li, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001 |
DAC | 6 |
| 2024 | Content-Aware Radiance Fields: Aligning Model Complexity with Scene Intricacy Through Learned Bitwidth Quantization
Xue Xian Zheng, Jingyi Yu 0001, Xin Lou 0001 |
ECCV (43) | 4 |
| 2024 | A Multi-scale Block PatchMatch-based Unified Algorithm for Efficient 6-D Vision Processingabstract6-D vision, which combines stereo and flow estimation to derive 3-D location and motion information, is crucial for comprehensive real-world perception. This paper presents a multi-scale block PatchMatch (MBPM)-based efficient 6-D vision algorithm with a balance tradeoff between computational complexity and accuracy, making it suitable for embedded application scenarios. By adopting a random search strategy, the proposed algorithm effectively reduces the computational burden without compromising accuracy. Additionally, the block-level processing approach further enhances efficiency by processing data in smaller, manageable chunks. Due to the significantly larger search range in flow estimation compared to stereo, we further incorporate direction initialization and multi-scale PatchMatch propagation techniques. These enhancements are crucial for ensuring reliable accuracy while expediting algorithm convergence. Compared with existing 6-D vision algorithms, the proposed method achieves 16.7% higher accuracy in stereo and 11.1% in flow estimation with much lower complexity. Even when faced with larger searching ranges, the proposed method exhibits consistent accuracy, allowing it to effectively handle fast motion scenarios. Hongyu Wang 0010, Xiangyu Zhang 0002, Xin Lou 0001 |
ISCAS | 3 |
| 2024 | Feature Map Guided Adapter Network for Object Detection in Low-light ConditionsabstractConventional ISP pipelines and image enhancement methods are designed and optimized for human vision, creating a gap between the requirements of computer and human visions. To bridge the requirement gap, we present a co-design framework in which backend computer vision plays a pivotal role in shaping the proceeding image processing algorithm. It features a pre- processing adapter network, responsible for the restoration and enhancement of RAW images from computer vision perspective, especially in challenging environmental conditions. Specifically, we extract feature maps from the backend vision network, utilizing them as constraints for optimizing the preprocessing adapter network. To validate the effectiveness of our proposed framework, we employ object detection in low-light conditions as the computer vision task, with YOLO-v5 as the backbone. Given the considerable noise in low-light images, we compare our results with state-of-the-art denoising algorithms, showcasing the superior performance of our framework. Cong Pang, Wei Zhou 0037, Xiangyu Zhang 0002, Xin Lou 0001 |
ISCAS | 5 |
| 2024 | An Efficient Hardware Volume Renderer for Convolutional Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) has attracted growing attention in the fields of 3D reconstruction and rendering. However, straightforward NeRF algorithms encounter challenges in accurately capturing complex surface details with rich high-frequency information. A recent development known as Convolutional Neural Radiance Field (ConvNeRF) has demonstrated state-of-the-art results for these tasks. But it comes with substantial irregular computational requirements, particularly in the convolutional volume rendering phase. In this paper, we introduce a hardware accelerator designed to enhance the efficiency of convolutional volume rendering in ConvNeRF. Our approach includes the creation of specialized computation modules and corresponding on-chip memory system optimized for seamless support of gated convolutions and skip connections in ConvNeRF. To validate our design, we implement it in VerilogHDL and build a prototype using Field Programmable Gate Array (FPGA). We also map our design to 40nm CMOS technology. The evaluation results underscore the superiority of our accelerator in terms of energy efficiency when compared to an implementation on an NVIDIA 2080Ti GPU, offering approximately 84.6× more frames per watt. Xuexin Wang, Yunxiang He, Xiangyu Zhang 0002, Pingqiang Zhou, Xin Lou 0001 |
ISCAS | 5 |
| 2024 | RAW Images-Based Motion-Assisted Object Detection Accelerator Using Deformable Parts Models Features on 1080p VideosabstractThis paper introduces an end-to-end object detection hardware accelerator that directly processes RAW video signals to generate detection results, enabling a holistic approach to optimization. Unlike existing works that primarily concentrate on the back-end object detector, we explore the redundancy present across multiple stages of the processing pipeline such as the image signal processing (ISP), the temporal correlation in consecutive frames and the back-end detector. A prototype of Deformable Parts Models (DPM)-based accelerator has been successfully validated on the Altera TR5 field-programmable gate array (FPGA) platform. This accelerator demonstrates efficient processing of high-resolution ($1920\times 1080$) videos at 60 frames per second (FPS) while incorporating a 12-scale gradient pyramid and consuming only 130.9 KB blocks of memory. To optimize the search process for motion estimation, we adopt the time division multiplexing (TDM) technology, which effectively reduces both multiplexer usage and memory access. Compared to conventional methods that scan a 1080p frame, the proposed head-based motion search hardware consumes 6.82% of the processing cycles and utilizes merely 6.9 KB of block memory. Evaluation and comparison results demonstrate the effectiveness of the proposed system. Ling Zhang 0010, Xiangyu Zhang 0002, Xin Lou 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Reconfigurable and Energy-Efficient Architecture for Deploying Multi-Layer RNNs on FPGAabstractRecurrent Neural Networks (RNNs) are extensively applied in sequence prediction tasks such as sentiment analysis, and machine translation. Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) are popular recurrent layers for mitigating gradient vanishing challenges. This paper introduces a reconfigurable hardware architecture that supports LSTM, GRU, and Fully Connected (FC) layers to accommodate the diverse structures of RNNs with minimal overhead. The proposed three-mode architecture utilizes dynamically reconfigurable components, allowing seamless mode switching among the three types of layers. Besides, Dynamic Compression (DC) for intermediate results is introduced to minimize precision loss, and layer decomposition as well as processing element grouping techniques are used to improve the processing efficiency. To validate the proposed architecture, a proof-of-concept prototype system using Intel Arria10 FPGA is built. Evaluation results demonstrate a 12% improvement over the state-of-the-art accelerator in energy efficiency when configured as GRU with a layer size set to$512{\times }$512, accompanied by a 93% reduction in block RAM and an 81% reduction in DSP resources. Xiangyu Zhang 0002, Yiren Zhu, Xin Lou 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | Ray Reordering for Hardware-Accelerated Neural Volume RenderingabstractNeural Volume Rendering (NVR) has advanced explosively since the advent of Neural Radiance Field (NeRF), a technique for novel view synthesis of complex scenes based on a finite set of input views. Existing ray casting-based NVR approaches process rays concurrently to leverage parallelism but fails to consider its impact on cache locality, which ultimately undermines the efficiency of corresponding dedicated hardware accelerator designs. We further observed that there exhibits spatial correspondence between features and voxels in NVR that can be exploited by processing in the order of voxel, not ray. This paper introduces a novel approach to meticulously reorder the execution of rays, ensuring that rays with similar memory access patterns are processed in parallel, thereby enhancing cache locality. On the basis of that, we also propose an efficient backend architecture and a corresponding memory subsystem, facilitating accurate data prefetching to hide off-chip memory latency. To validate the proposed architecture, we implement our design in VerilogHDL and evaluate the performance by post-synthesis simulation with real scene data. The evaluation results demonstrate that our design markedly enhances the efficiency of NVR processing, achieving a considerable speedup ($1.62\times $) compared to the state-of-the-art NVR accelerator, while necessitating significantly less silicon area ($5.12\times $) and power ($32.79\times $). Junran Ding, Yunxiang He, Binzhe Yuan, Zhechen Yuan, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | An Efficient Frequency Domain Vision Pipeline From RAW Images to Backend TasksabstractThough high resolution benefits computer vision performance, they are not commonly used in convolutional neural network (CNN)-based vision algorithms due to the limitation of memory and computation resource. Learning in the frequency domain makes high resolution images directly acceptable by CNNs, but the computation, time and energy overhead for pre-processing, including image signal processing (ISP) and domain transformation, can be large. This paper explores different image processing and domain transformation operations and proposes an efficient end-to-end frequency domain learning pipeline from RAW images to vision tasks. In particular, we simplify the pre-processing part by skipping the entire ISP pipeline and replacing the Discrete Cosine Transform (DCT) with a multiplication-free approximated one. Experimental results show that the final vision performance of the proposed pipeline is very close to that of the conventional pipeline, while significant amount of redundant operations can be saved. Wei Zhou 0037, Xiangyu Zhang 0002, Xin Lou 0001 |
ISCAS | 4 |
| 2023 | Analysis and Design of Precision-Scalable Computation Array for Efficient Neural Radiance Field RenderingabstractNeural Radiance Field (NeRF), a disruptive method for 3D representation and rendering, is extremely popular in the field of computer graphics and computer vision in the past three years. The most distinctive feature of NeRF models is their scene representation property, making it possible to quantize the models according to the complexity of the representing scenes. This paper proposes a novel approach to improve the efficiency of NeRF rendering by adopting precision-scalable computation. We first analyze and validate the idea of scene-dependent quantization for NeRF models. Based on that, we further propose look-up table (LUT) processing element (PE)-based precision-scalable computation unit designs. To evaluate the performance of different precision-scalable computing units, we implement these designs and compare the corresponding area, power, speed and energy efficiency. We also compare the proposed designs with existing approaches as well as the fixed precision approach for NeRF rendering tasks. The comparison results show that energy efficiency can be significantly improved by using precision-scalable computation for NeRF. Kangjie Long, Chaolin Rao, Yunxiang He, Zhechen Yuan, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | An Energy-Efficient Accelerator for Medical Image Reconstruction From Implicit Neural RepresentationabstractThis work presents an energy-efficient accelerator for medical image reconstruction from implicit neural representation (INR). The accelerator implements an INR-based algorithm to deliver high-quality medical image reconstruction with arbitrary resolution from a compact implicit format. In particular, we propose a dedicated hardware architecture based on an optimized computation flow for the INR-based reconstruction algorithm, which co-designs data reuse and computation load. The proposed architecture takes in the coordinate of the intersection of three scans and outputs all the voxel intensities, minimizing the data movement between on-chip and off-chip. To validate the proposed accelerator, we build a proof-of-concept prototype demonstration system using field programmable gate array (FPGA). We also map our design to 40nm CMOS technology to measure the performance of the proposed accelerator. The implementation results show that, running at 400MHz, the proposed accelerator is capable of processing medical images with$256\times 256$resolution in real-time at 26.3 frames per second (FPS), with a power consumption of only 795 mW. Comparison results show that the performance, as well as the energy efficiency of the proposed accelerator, outperforms the central processing unit (CPU)-based and graphic processing unit (GPU)-based implementations. Chaolin Rao, Qing Wu 0001, Pingqiang Zhou, Jingyi Yu 0001, Yuyao Zhang 0005, Xin Lou 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | An End-to-end Computer Vision System ArchitectureabstractTo overcome the data movement bottleneck, near-sensor and in-sensor computing are becoming more and more popular. However, in the existing near-/in-sensor computing architectures for vision tasks, the effect of the image signal processing (ISP) pipeline, which is of great importance to the final vision performance [1], is always ignored. In this work, we propose a synthesized RAW image-based end-to-end computer vision paradigm, taking the effect of ISP pipeline into account. In the proposed approach, a generative adversarial network (GAN)-based tool is used to convert the fully processed color images to their corresponding RAW Bayer versions, generating the training data for end-to-end vision models. In the inference stage, RAW images from the sensor are directly fed to the end-to-end model, bypassing the entire ISP pipeline. Experimental results show that by training/tuning the CNN models using synthesized RAW images, it is possible to design an end-to-end (from RAW image to vision task) vision system that directly consumes RAW image data from the sensor with negligible vision performance degradation. By skipping the ISP pipeline, an image sensor can be directly integrated with the back-end vision processor without a complex image processor in the middle, making near-/in-sensor computing a practical approach. Ling Zhang 0010, Wei Zhou 0037, Xiangyu Zhang 0002, Xin Lou 0001 |
ISCAS | 4 |
| 2022 | Radar-Based Human Activity Recognition With 1-D Dense Attention NetworkabstractWith the development of the Internet of things, radar-based human activity recognition is becoming more and more important, because they play an indispensable role in fields such as safety and health monitoring. In this work, a novel network named 1-D dense attention neural network (1-D-DAN) is proposed for the radar-based human activity recognition. In the proposed network, a novel attention mechanism network structure specifically designed for radar spectrogram is proposed, equipping 1-D convolutional network with attention mechanism. With the$x$-axis of the spectrogram represents time and the$y$-axis represents frequency, the proposed attention mechanism includes two branches: 1) time attention branch and 2) frequency attention branch. Moreover, a dense attention operation that can make full use of features in the network is also introduced in the proposed attention mechanism. Experimental results show that compared with the state-of-the-art methods, our proposed 1-D-DAN achieves the highest accuracy in human activity recognition with the lowest computational complexity. Guoji Lai, Xin Lou 0001, Wen Bin Ye 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Ring and Radius Sampling Based Phasor Field Diffraction Algorithm for Non-Line-of-Sight ReconstructionabstractNon-Line-of-Sight (NLOS) imaging reconstructs occluded scenes based on indirect diffuse reflections. The computational complexity and memory consumption of existing NLOS reconstruction algorithms make them challenging to be implemented in real-time. This paper presents a fast and memory-efficient phasor field-diffraction-based NLOS reconstruction algorithm. In the proposed algorithm, the radial property of the Rayleigh Sommerfeld diffraction (RSD) kernels along with the linear property of Fourier transform are utilized to reconstruct the Fourier domain representations of RSD kernels using a set of kernel bases. Moreover, memory consumption is further reduced by sampling the kernel bases in a radius direction and constructing them during the run-time. According to the analysis, the memory efficiency can be improved by as much as 220×. Experimental results show that compared with the original RSD algorithm, the reconstruction time of the proposed algorithm is significantly reduced with little impact on the final imaging quality. Deyang Jiang, Xiaochun Liu, Jianwen Luo 0004, Zhengpeng Liao, Andreas Velten, Xin Lou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Multi-Level Time-Frequency Bins Selection for Direction of Arrival Estimation Using a Single Acoustic Vector SensorabstractIn the context of multi-source direction of arrival (DOA) estimation in an enclosed environment, the challenges include reverberation and overlapping of multiple simultaneous active sources. To address these interferences, the identification of time-frequency (TF) bins dominated by the sources signals is essential. In this work, we propose an intensity vector (IV) based TF bins selection technique for DOA estimation using a single acoustic vector sensor (AVS). The proposed technique involves multi-level inliers selection and outliers removal (MLISOR), which is implemented in three steps. In the first step, we derive the distribution of IVs and then select IVs using a norm metric. In the second step, the regions with the highest local IV density in each time frame are identified. In the third step, we cluster the IVs according to their directions and remove the outliers based on the member-to-centroid angle metric. Simulation results show that both the accuracy and the robustness of the proposed technique outperform the existing techniques. The indoor experimental results also verify that the proposed technique is effective and robust in practical situations. Jianhua Geng, Sifan Wang, Qinglai Liu, Xin Lou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Reconfigurable Nonuniform Filter Bank for Hearing Aid SystemsabstractFilter banks with good reconfigurability are desired in hearing aid systems due to the individual requirements of patients with different hearing loss characteristics. This paper proposes a modularized design approach of a completely reconfigurable nonuniform filter bank for hearing aid systems. The proposed filter bank structure consists of a multiband generation module and a subband extraction module, where the frequency warping technique is adopted. Based on theoretical analyses of second order frequency warping, the subband distribution can be flexibly controlled. Moreover, to adapt the subband extraction module design to the multiband generation module, the relationship between control parameters is analyzed and a mapping formula is derived. Through the co-design of the two modules, superior reconfigurability over existing methods is achieved. Application examples show that the proposed filter bank is able to provide multiple subbands distribution schemes that satisfy the requirements of different audiograms. The proposed filter bank structure has lower complexity, smaller delay and smaller matching errors than existing filter bank structures due to its flexible structure. Ying Wei 0004, Xin Lou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | FPGA Accelerator for Real-Time Non-Line-of-Sight ImagingabstractNon-line-of-sight (NLOS) imaging systems reconstruct hidden scenes using computational methods based on indirect light that diffusely reflected from relay walls. Due to the computation and memory requirements of reconstruction algorithms, real-time NLOS imaging for room-size scenes based on non-confocal data has long been challenging. This paper proposes a field programmable gate array (FPGA) accelerator for the recently proposed Rayleigh-Sommerfeld Diffraction (RSD)-based NLOS reconstruction method. In the proposed accelerator design, ring sampling and radius sampling techniques are proposed to reduce the memory requirements by reconstructing the RSD kernels with a set of kernel bases and ring sampling coefficients during the runtime. Based on that, a customized hardware architecture and the corresponding FPGA design for real-time RSD-based NLOS reconstruction is further proposed. Implementation results show that the proposed FPGA accelerator is capable of reconstructing NLOS scenes at 25 frames per second (FPS), running at a relatively slow clock frequency of 50 MHz. To the best knowledge of the authors, this is the first real-time enabled FPGA accelerator for room-size NLOS imaging with a resolution of$128\times 128$. Zhengpeng Liao, Deyang Jiang, Xiaochun Liu, Andreas Velten, Yajun Ha, Xin Lou 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | A Block PatchMatch-Based Energy-Resource Efficient Stereo Matching Processor on FPGAabstractThis paper presents a field programmable gate array (FPGA)-based, high-performance and energy-resource efficient stereo matching processor. The proposed processor executes block-level PatchMatch-based stereo matching algorithm with a random search strategy to avoid estimation of all disparity levels. To take advantages of different block scales, a coarse-to-fine multi-scale propagation (MSP) scheme is proposed for label update. Based on that, a dedicated hardware architecture is further proposed to explore the benefit of the algorithm. Experimental results show that the proposed FPGA-based processor, running at 350MHz, achieves a peak performance of$1920\times 1080.165$.7 frame per second (FPS) at 128 disparity levels with 3.35W power dissipation. The energy and resource efficiency of the proposed design outperforms state-of-the-art FPGA-based stereo matching processors. When disparity level increases to 256, the computing resource increment of the proposed design is much less than existing designs because random search instead of winner-takes-all (WTA) is utilized. Moreover, unlike existing dedicated stereo matching processors which output only disparity information, the proposed design is also capable of deriving plane slant. This information can be beneficial for follow-up tasks like 3D reconstruction. Hongyu Wang 0010, Wei Zhou 0037, Xiangyu Zhang 0002, Xin Lou 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | A Raw Image-Based End-to-End Object Detection Accelerator Using HOG FeaturesabstractThis paper presents an end-to-end object-detection accelerator that processes raw Bayer images to generate detection results. The accelerator utilizes histogram of oriented gradients (HOG) features in combination with a support vector machine (SVM) classifier. The proposed HOG for raw images (HOGR) skips the image signal processors which consume a significant amount of power (about 2.5X of an accelerator). The proposed architecture temporally partitions the algorithms and time multiplexes the logic such that the accelerator works on the same frequency with the image sensor using affordable resources. The prototype is verified using 1080p ($1920\times 1080$) and VGA ($640\times 480$) raw videos on Altera Arria10 and Cyclone IV field programmable gate array (FPGA) platforms. The accelerator can process 1080p raw videos with a 12-scale pyramid at 60 frames per second (FPS) under the pixel frequency of corresponding image sensor (148.5 MHz), consuming 510 Kbit block memory. To the best knowledge of the authors, this is the first end-to-end HOG+SVM accelerator that takes raw Bayer images as input, skips the ISP pipeline for resource optimization from the system perspective, and is synchronized with the image sensor. Xiangyu Zhang 0002, Ling Zhang 0010, Xin Lou 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | ICARUS: A Specialized Architecture for Neural Radiance Fields RenderingabstractThe practical deployment of Neural Radiance Fields (NeRF) in rendering applications faces several challenges, with the most critical one being low rendering speed on even high-end graphic processing units (GPUs). In this paper, we present ICARUS, a specialized accelerator architecture tailored for NeRF rendering. Unlike GPUs using general purpose computing and memory architectures for NeRF, ICARUS executes the complete NeRF pipeline using dedicated plenoptic cores (PLCore) consisting of a positional encoding unit (PEU), a multi-layer perceptron (MLP) engine, and a volume rendering unit (VRU). A PLCore takes in positions & directions and renders the corresponding pixel colors without any intermediate data going off-chip for temporary storage and exchange, which can be time and power consuming. To implement the most expensive component of NeRF, i.e., the MLP, we transform the fully connected operations to approximated reconfigurable multiple constant multiplications (MCMs), where common subexpressions are shared across different multiplications to improve the computation efficiency. We build a prototype ICARUS using Synopsys HAPS-80 S104, a field programmable gate array (FPGA)-based prototyping system for large-scale integrated circuits and systems design. We evaluate the power-performancearea (PPA) of a PLCore using 40nm LP CMOS technology. Working at 400 MHz, a single PLCore occupies 16.5 mm 2 and consumes 282.8 mW, translating to 0.105 uJ/sample. The results are compared with those of GPU and tensor processing unit (TPU) implementations. Chaolin Rao, Huangjie Yu, Haochuan Wan, Jindong Zhou, Yueyang Zheng, Minye Wu, Anpei Chen, Binzhe Yuan, Pingqiang Zhou, Xin Lou 0001, Jingyi Yu 0001 |
ACM Trans. Graph. | 11 |
| 2021 | Multi-Scale Slanted O(1) Stereo Matching AlgorithmabstractIn this paper, a multi-scale slanted O(1) stereo (MSOS) matching algorithm is proposed. In the proposed MSOS, the concept of multi-scale propagation is introduced to handle the problem of different object size in practical applications. Moreover, the census feature, which is more robust than the original sum of absolute difference (SAD), is employed in the proposed MSOS to compensate the effect of practical situations such as radiometric change. Experimental results show that the matching performance of the proposed MSOS is greatly improved over the original SOS on the outdoor KITTI2015 dataset. Hongyu Wang 0010, Shengyu Gao, Xin Lou 0001 |
ISCAS | 5 |
| 2021 | Lightweight Deep Learning Model in Mobile-Edge Computing for Radar-Based Human Activity RecognitionabstractRadar-based human activity recognition (HAR) has great potential in many fields, such as surveillance, smart homes, and human-computer interaction. Complex deep neural networks have brought significant improvement in classification performance but also a surge of computational cost and the number of parameters, which makes it challenging to deploy in mobile devices. However, the existing studies in this area mainly focus on improving the classification accuracy. In this article, we propose an extremely efficient convolutional neural network (CNN) architecture named Mobile-RadarNet, which is specially designed for human activity classification based on micro-Doppler signatures. The new architecture exploits 1-D depthwise convolutions and pointwise convolutions to build lightweight CNN architecture. The experiments on a seven-class human activity data set demonstrate that the proposed Mobile-RadarNet can achieve high classification accuracy meanwhile to keep the computational complexity at an extremely low level, and thus has great potential to be deployed in the mobile devices. Xin Lou 0001, Wen Bin Ye 0001 |
IEEE Internet Things J. | 2 |
| 2021 | Gradient-Based Feature Extraction From Raw Bayer Pattern ImagesabstractIn this paper, the impact of demosaicing on gradient extraction is studied and a gradient-based feature extraction pipeline based on raw Bayer pattern images is proposed. It is shown both theoretically and experimentally that the Bayer pattern images are applicable to the central difference gradient-based feature extraction algorithms with negligible performance degradation, as long as the arrangement of color filter array (CFA) patterns matches the gradient operators. The color difference constancy assumption, which is widely used in various demosaicing algorithms, is applied in the proposed Bayer pattern image-based gradient extraction pipeline. Experimental results show that the gradients extracted from Bayer pattern images are robust enough to be used in histogram of oriented gradients (HOG)-based pedestrian detection algorithms and shift-invariant feature transform (SIFT)-based matching algorithms. By skipping most of the steps in the image signal processing (ISP) pipeline, the computational complexity and power consumption of a computer vision system can be reduced significantly. Wei Zhou 0037, Ling Zhang 0010, Shengyu Gao, Xin Lou 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | K-SVD Based Denoising Algorithm for DoFP Polarization Image SensorsabstractThis paper presents a novel K times singular value decomposition (K-SVD) based denoising algorithm for the division-of-focal-plane (DoFP) polarization image sensors. In the proposed implementation, the input DoFP image can be expressed by the optimum sparse combination of the dictionary elements via K-SVD and orthogonal matching pursuit (OMP) algorithms. As a result, this implementation is capable of eliminating the Gaussian noise significantly and well-preserving the details and edges of the target DoFP image. Our extensive experimental results on various test images show that the proposed algorithm yields better visual quality and maintains a lower PSNR value while compared with a wide range of previous implementations. Shiting Li, Wen Bin Ye 0001, Huawei Liang, Xiaofang Pan, Xin Lou 0001, Xiaojin Zhao |
ISCAS | 5 |
| 2017 | A passively compensated capacitive sensor readout with biased varactor temperature compensation and temperature coherent quantizationabstractThis paper presents a frequency-mode capacitive sensor readout front-end operating in a wide temperature range. First, a biased varactor temperature compensation (BVTC) is proposed to compensate the aggregate temperature gradients from the sensor and the oscillator circuit, achieving a nullified temperature coefficient for the oscillation frequency. Second, a temperature coherent quantization (TCQ) approach is proposed to enhance the sensor' sensitivity and to provide a self-referenced clock for digitization, whereby the influence from the temperature effect of the clock is minimized through a hybrid down-conversion time-to-digital converter (DTDC). A prototype chip was fabricated using the 0.18-μm CMOS process and it was verified using a commercial 5.92-to-6.53-pF capacitive pressure sensor over a temperature range of −20°C to 120°C. Proven by experiments, the prototype presented as good as ±0.4% full-scale pressure error in the 140°C temperature range. Wang Ling Goh, Kevin Tshun Chuan Chai, Xin Lou 0001, Wen Bin Ye 0001 |
ISCAS | 5 |
| 2017 | Lower Bound Analysis and Perturbation of Critical Path for Area-Time Efficient Multiple Constant MultiplicationsabstractIn this paper, a precise systematic delay model is proposed for the analysis and estimation of critical path delay of multiple constant multiplication (MCM) blocks. For the first time in literature, the mathematical derivation of lower bound of critical path delay of MCM blocks is presented and necessary conditions for achieving the lower bound of critical path delay are discussed. It is shown that the lower bound of critical path delay of MCMs is significantly smaller than that achieved by existing MCM algorithms. An improved genetic algorithm-based approach, with a heuristic algorithm to generate the initial population, is proposed to search for low complexity MCM solutions with the lower bound of critical path delay. This is the first time that design algorithms with gate-level delay control is proposed. Moreover, it is shown that using the information of lower bound of critical path delay, perturbation of timing can be applied to tradeoff the lower bound critical path delay against hardware complexity. It is shown that area-time efficient design of MCM blocks can be obtained by using the proposed techniques. Xin Lou 0001, Ya Jun Yu, Pramod Kumar Meher |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2015 | Fine-grained pipelining for multiple constant multiplicationsabstractMultiple constant multiplication (MCM) is widely used in several digital signal processing applications. Recently, pipelining techniques have been applied to accelerate the computation of the MCM blocks. The existing pipelining techniques consider the adder stage pipelining, i.e., inserting registers between two adjacent adder stages, to reduce the adder depth. However, the critical path may be still long, even though the adder depth is minimized. In this paper, the adder stage pipelining method is analyzed at bit-level and a novel pipelining method is proposed for pipelining the adders in the MCM block. Experimental results show that the proposed pipelining method provides nearly 32% reduction of critical path over the traditional adder stage pipelining in average for several benchmark MCM blocks, while the area and power consumption are maintained. Xin Lou 0001, Pramod Kumar Meher, Ya Jun Yu |
ISCAS | 1 |
| 2015 | Design of high-speed multiplierless linear-phase FIR filtersabstractIn the design of multiplierless FIR filters, researchers have made every effort to reduce the number of adders when coefficients multipliers are realized using adder-and-shift network to decrease the overall chip area. However, with the advance of IC technology, area becomes a less important issue than the speed. In this paper, we propose a speed oriented optimization of linear phase FIR filters, where the length of critical path is used as the criteria in the discrete coefficient search. The length of critical path is measured as the number of cascaded full adders rather than the traditional adder depth. Compared to the area oriented algorithm, the proposed algorithm can generate the filters with much shorter critical path delay and meanwhile the area-delay product is also reduced. Gate level simulations of benchmark filters verify the above claim. Wen Bin Ye 0001, Xin Lou 0001, Ya Jun Yu |
ISCAS | 2 |
| 2014 | High-speed multiplier block design based on bit-level critical path optimizationabstractMultiple constant multiplications (MCM) is a popular technique to implement multiplier blocks with low hardware cost and power consumption. Research works on MCM have been on going for more than two decades. Most algorithms so far have focused on reducing the number of adders and/or adder depth to have low power and/or high speed circuit. However, low adder depth does not guarantee the low critical path examined in bit-level. In this work, we propose an algorithm to optimize the critical path of multiplier blocks in bit-level. Simulation results show that the critical path delay can be reduced by using the proposed algorithm. Xin Lou 0001, Ya Jun Yu, Pramod Kumar Meher |
ISCAS | 1 |