EDBT 2026 Demo / reviewers in the wild / expert
Yunxiang He
dblp:278/8518
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCAR: A Neural Rendering Accelerator with Sparse-aware Sampling and Conflict-Free Encoding
Yunxiang He, Yongzhi Zhang, Qihan Ding, Quanyu Chen |
ISCAS | 1 |
| 2026 | An Energy-Efficient Edge Coprocessor for Neural Rendering With Explicit Data Reuse StrategiesabstractNeural radiance fields (NeRFs) have transformed 3-D reconstruction and rendering, facilitating photorealistic image synthesis from sparse viewpoints. This work introduces an explicit data reuse neural rendering (EDR-NR) architecture, which reduces frequent external memory accesses (EMAs) and cache misses by exploiting the spatial locality from three phases, including rays, ray packets (RPs), and samples. The EDR-NR architecture features a four-stage scheduler that clusters rays on the basis of$Z$-order, prioritize lagging rays when ray divergence happens, reorders RPs based on spatial proximity, and issues samples out-of-orderly (OoO) according to the availability of on-chip feature data. In addition, a four-tier hierarchical RP marching (HRM) technique is integrated with an axis-aligned bounding box (AABB) to facilitate spatial skipping (SS), reducing redundant computations and improving throughput. Moreover, a balanced allocation strategy for feature storage is proposed to mitigate SRAM bank conflicts. Fabricated using a 40-nm process with a die area of 10.5 mm2, the EDR-NR chip demonstrates a$2.41\times $enhancement in normalized energy efficiency, a$1.21\times $improvement in normalized area efficiency, a$1.20\times $increase in normalized throughput, and a 53.42% reduction in on-chip SRAM consumption compared with state-of-the-art accelerators. Binzhe Yuan, Xiangyu Zhang 0002, Yuefeng Zhang, Haochuan Wan, Zhechen Yuan, Junsheng Chen, Yunxiang He, Junran Ding, Chaolin Rao, Wenyan Su, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2025 | Radio Map Reconstruction Based on Nas Enhanced Deep Regularization Completion for Uav CommunicationsabstractThis paper proposes a radio map reconstruction method based on the Neural Architecture Search Enhanced Deep Regularization Model. Traditional radio map reconstruction algorithms face many challenges, such as extremely sparse measurement data and complex wireless signal environments. To solve the problem of unstable output, an external regularizer is introduced to provide additional regularization input in the prediction process to assist the neural network in capturing the explicit prior information from the sparse measurements and performing restoration. Moreover, to ensure the efficiency of capturing implicit prior information, neural architecture search is adopted to optimize the structure of the completion network, further enhancing the robustness of the neural network and the accuracy of reconstruction. The experimental results show that our proposed method is more accurate and stable than the traditional completion methods, especially in the circumstance of a low sampling rate, and can better adapt to the situation of interrupting the probability distribution in unknown regions without a dataset. Yunxiang He, Lingyao Wang, Minxian Shen, Gongrui Huang, Hao Huang 0008, Haitao Zhao 0004 |
VTC2025-Spring | 1 |
| 2025 | A Joint Optimization Framework for Sum-Rate Maximization in Air Reconfigurable Intelligent Surface Assisted MIMO-NOMA SystemsabstractIn this article, a novel multiuser multiple-input-multiple-output (MIMO) communication system for Internet of Things (IoT) is proposed, where the aerial reconfigurable intelligent surface (ARIS) and nonorthogonal multiple access (NOMA) are used as the sum rate enhancement pathway. The base station (BS) has multiple antennas that transmit superimposed signals to multiple users. The passive ARIS serves as a flexible transmit relay to reduce path loss and improve channel gains. Users are divided into several groups based on their channel status, each sharing a radio frequency (RF) chain. To maximize the sum rate of all users, the placement of ARIS, the passive/active beamforming design and the power allocation among users are jointly optimized. As the joint optimization for user grouping, passive/active beamforming and power distribution is formulated as a mixed-integer nonlinear program (MINLP) which is nonconvex and coupled and hence, obtaining an optimal solution is challenging. In this article, the problem is decoupled into three subproblems and solved alternately efficiently. The numerical results demonstrate that the suggested MIMO-ARIS-NOMA system can achieve higher sum rate performance than traditional schemes. Haitao Zhao 0004, Zhipeng Kong, Yunxiang He, Biyao Ding, Hao Huang 0008, Yiyang Ni 0001, Guan Gui 0001, Hikmet Sari, Fumiyuki Adachi |
IEEE Internet Things J. | 3 |
| 2025 | A 0.92-pJ/b 112-Gb/s PAM-4 Transmitter With Bandwidth and Linearity Enhanced Quasi-Voltage-Mode Driver and Reconfigurable Three-Tap T/2-T Variable Fractional-Spaced FFE in 28-nm CMOSabstractThis study presents a 112-Gb/s four-level pulse-amplitude modulation transmitter implemented in a 28-nm CMOS process. An innovative quasi-voltage-mode driver is proposed, which demonstrates similar bandwidth compared to the current-mode logic driver with an approximately 30% power reduction and provides flexible linearity fine-tuning. The TX architectures with and without the pre-driver stage are optimized to exploit the bandwidth limit further. A three-tap T/2–T variable-spaced feed-forward equalizer is designed to realize reconfigurable 0.5-T, 0.6-T, 0.75-T, and 1-T tap delays, which enables customized eye-opening optimization under different data rates and channel responses. The measurement results show that the energy efficiency of the proposed TX is 0.92 pJ/b with a 0.9-Vppd DC output swing, and the best level separation mismatch ratios are 99.8% and 99.6% at 100-Gb/s and 112-Gb/s, respectively. Kanan Wang 0001, Renjie Tang, Shuyi Xiang, Yunxiang He, Xiaoyan Gui |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | A Neural Rendering Coprocessor With Optimized Ray Representation and MarchingabstractNeural rendering, a transformative approach for 3-D scene reconstruction and rendering, has advanced rapidly in recent years. This article introduces an energy-efficient neural rendering coprocessor that implements the popular and widely used instant neural graphics primitive (Instant-NGP) algorithm. In particular, we address the challenges of limited resources for deploying Instant-NGP on edge by proposing a dedicated architecture, which incorporates three main innovations: 1) we optimize occupancy grid queries in the ray marching module by partitioning the grid and decoupling the query process from sampling point generation, which improves both efficiency and memory usage; 2) we introduce a bilinked list-based ray switching strategy, which ensures continuous pipeline utilization to overcome the inefficiencies caused by sequential processing; and 3) we optimize the hash encoding process by incorporating quantization-aware training (QAT), enabling the hash table to fit into on-chip memory, thereby improving performance on resource-constrained devices. To demonstrate the effectiveness of our architecture, we design and fabricate a proof-of-concept chip using 40-nm CMOS technology and develop a testing system to evaluate its performance. Measurement results validate the advantages of the proposed design, showing that our chip achieves superior energy efficiency compared to both server and edge graphics processing units (GPUs), as well as other state-of-the-art neural rendering chip designs. Zhechen Yuan, Binzhe Yuan, Chaolin Rao, Yiren Zhu, Yunxiang He, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | Density Estimation-based Effective Sampling Strategy for Neural RenderingabstractExpanding upon the foundational Neural Radiance Fields (NeRF) framework, neural rendering techniques have seen wide-ranging applications in 3D reconstruction and rendering. Despite the numerous attempts to accelerate the original NeRF, the ray marching process in many neural rendering algorithms remains a performance bottleneck, constraining their rendering speed. In this work, we introduce an innovative method for effective sampling based on density estimation to alleviate this bottleneck. The proposed method can effectively reduce the number of required occupancy grid access without compromising rendering quality. Evaluation and analysis results validate the effectiveness of the proposed approach, delivering notable improvements in rendering efficiency. Yunxiang He |
ISCAS | 1 |
| 2024 | An Efficient Hardware Volume Renderer for Convolutional Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) has attracted growing attention in the fields of 3D reconstruction and rendering. However, straightforward NeRF algorithms encounter challenges in accurately capturing complex surface details with rich high-frequency information. A recent development known as Convolutional Neural Radiance Field (ConvNeRF) has demonstrated state-of-the-art results for these tasks. But it comes with substantial irregular computational requirements, particularly in the convolutional volume rendering phase. In this paper, we introduce a hardware accelerator designed to enhance the efficiency of convolutional volume rendering in ConvNeRF. Our approach includes the creation of specialized computation modules and corresponding on-chip memory system optimized for seamless support of gated convolutions and skip connections in ConvNeRF. To validate our design, we implement it in VerilogHDL and build a prototype using Field Programmable Gate Array (FPGA). We also map our design to 40nm CMOS technology. The evaluation results underscore the superiority of our accelerator in terms of energy efficiency when compared to an implementation on an NVIDIA 2080Ti GPU, offering approximately 84.6× more frames per watt. Xuexin Wang, Yunxiang He, Xiangyu Zhang 0002, Pingqiang Zhou, Xin Lou 0001 |
ISCAS | 2 |
| 2024 | Dentists who can auscultate: Microphone-based toothbrushing quality monitoring system for electronic toothbrush
Jiahe Cui, Di Wu 0070, Yunxiang He, Zhenchao Ouyang |
Expert Syst. Appl. | 3 |
| 2024 | Ray Reordering for Hardware-Accelerated Neural Volume RenderingabstractNeural Volume Rendering (NVR) has advanced explosively since the advent of Neural Radiance Field (NeRF), a technique for novel view synthesis of complex scenes based on a finite set of input views. Existing ray casting-based NVR approaches process rays concurrently to leverage parallelism but fails to consider its impact on cache locality, which ultimately undermines the efficiency of corresponding dedicated hardware accelerator designs. We further observed that there exhibits spatial correspondence between features and voxels in NVR that can be exploited by processing in the order of voxel, not ray. This paper introduces a novel approach to meticulously reorder the execution of rays, ensuring that rays with similar memory access patterns are processed in parallel, thereby enhancing cache locality. On the basis of that, we also propose an efficient backend architecture and a corresponding memory subsystem, facilitating accurate data prefetching to hide off-chip memory latency. To validate the proposed architecture, we implement our design in VerilogHDL and evaluate the performance by post-synthesis simulation with real scene data. The evaluation results demonstrate that our design markedly enhances the efficiency of NVR processing, achieving a considerable speedup ($1.62\times $) compared to the state-of-the-art NVR accelerator, while necessitating significantly less silicon area ($5.12\times $) and power ($32.79\times $). Junran Ding, Yunxiang He, Binzhe Yuan, Zhechen Yuan, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | A Novel Topology Metric for Indoor Point Cloud SLAM Based on Plane Detection Optimization
Zhenchao Ouyang, Jiahe Cui, Yunxiang He, Dongyu Li, Qinglei Hu, Changjie Zhang |
CollaborateCom (3) | 3 |
| 2023 | Analysis and Design of Precision-Scalable Computation Array for Efficient Neural Radiance Field RenderingabstractNeural Radiance Field (NeRF), a disruptive method for 3D representation and rendering, is extremely popular in the field of computer graphics and computer vision in the past three years. The most distinctive feature of NeRF models is their scene representation property, making it possible to quantize the models according to the complexity of the representing scenes. This paper proposes a novel approach to improve the efficiency of NeRF rendering by adopting precision-scalable computation. We first analyze and validate the idea of scene-dependent quantization for NeRF models. Based on that, we further propose look-up table (LUT) processing element (PE)-based precision-scalable computation unit designs. To evaluate the performance of different precision-scalable computing units, we implement these designs and compare the corresponding area, power, speed and energy efficiency. We also compare the proposed designs with existing approaches as well as the fixed precision approach for NeRF rendering tasks. The comparison results show that energy efficiency can be significantly improved by using precision-scalable computation for NeRF. Kangjie Long, Chaolin Rao, Yunxiang He, Zhechen Yuan, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |