EDBT 2026 Demo / reviewers in the wild / expert
Xiangyu Zhang 0002
dblp:95/3760-2
· DBLP profile ↗
20ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0003-3716-4722ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 2 first-author · 17 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalarium: A Unified Scala-based Co-Simulation Framework for Agile Chip Development
Yuefeng Zhang, Wenkai Zhou, Binzhe Yuan, Junsheng Chen, Xiangyu Zhang 0002, Hao Geng, Xin Lou 0001 |
ASP-DAC | 6 |
| 2026 | A High-Performance Neural Rendering Accelerator Based on Novel Multi-Level Ray Scheduling and Dual-Process BackendabstractNeural rendering enables photorealistic scene re-construction but remains difficult to deploy on edge devices due to intensive computation, redundant sampling, and memory bandwidth constraints. This work presents a high-performance neural rendering accelerator for real-time embedded rendering. The proposed design integrates: (1) a dual-process backend with fused micro-MLPs to significantly improve sample processing efficiency, (2) multi-resolution spatial partitioning with adaptive ray clustering to exploit sparsity and achieve over 95% cache hit rate, and (3) a multi-level scheduling framework with proactive prefetching to reduce MLP stalls. Implemented on FPGA, the prototype achieves 94.7 FPS at 800×800 resolution with 6.4 W power consumption. An ASIC implementation in 28 nm technology sustains 440 FPS at 268 mW. Experimental results demonstrate state-of-the-art performance and energy efficiency while preserving rendering quality above 30 dB PSNR. Wenkai Zhou, Yuefeng Zhang, Binzhe Yuan, Junsheng Chen, Luntian Zhang, Xiangyu Zhang 0002, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001 |
DATE | 7 |
| 2026 | ZeroBlade: A Spatial Similarity-aware HiSparse MLP Engine for Neural Volume Rendering
Antong Li, Haochuan Wan, Xiangyu Zhang 0002, Xin Lou 0001 |
ISCAS | 3 |
| 2026 | An Efficient Low-Light Object Detection Framework based on Task-Driven Distillation
Wei Zhou 0037, Kangjie Long, Cong Pang, Xiangyu Zhang 0002, Xin Lou 0001 |
ISCAS | 5 |
| 2026 | WaveMamba: A Vision Backbone Synergizing Feature Extraction and Downsampling
Haitian Yang, Xiangyu Zhang 0002, Yuanmei Zhang, Xin Lou 0001, Wei Zhou 0037 |
ISCAS | 2 |
| 2026 | AlignLite: A Lightweight Framework for Weakly Aligned Multimodal Object Detection
Haitian Yang, Xiangyu Zhang 0002, Yuanmei Zhang, Xin Lou 0001, Wei Zhou 0037 |
ISCAS | 2 |
| 2026 | Unity-EDR: A Hybrid Neural-Mesh Rendering System for Efficient Visual Synthesis
Chaolin Rao, Xiangyu Zhang 0002, Xin Lou 0001 |
ISCAS | 3 |
| 2026 | An Energy-Efficient Edge Coprocessor for Neural Rendering With Explicit Data Reuse StrategiesabstractNeural radiance fields (NeRFs) have transformed 3-D reconstruction and rendering, facilitating photorealistic image synthesis from sparse viewpoints. This work introduces an explicit data reuse neural rendering (EDR-NR) architecture, which reduces frequent external memory accesses (EMAs) and cache misses by exploiting the spatial locality from three phases, including rays, ray packets (RPs), and samples. The EDR-NR architecture features a four-stage scheduler that clusters rays on the basis of$Z$-order, prioritize lagging rays when ray divergence happens, reorders RPs based on spatial proximity, and issues samples out-of-orderly (OoO) according to the availability of on-chip feature data. In addition, a four-tier hierarchical RP marching (HRM) technique is integrated with an axis-aligned bounding box (AABB) to facilitate spatial skipping (SS), reducing redundant computations and improving throughput. Moreover, a balanced allocation strategy for feature storage is proposed to mitigate SRAM bank conflicts. Fabricated using a 40-nm process with a die area of 10.5 mm2, the EDR-NR chip demonstrates a$2.41\times $enhancement in normalized energy efficiency, a$1.21\times $improvement in normalized area efficiency, a$1.20\times $increase in normalized throughput, and a 53.42% reduction in on-chip SRAM consumption compared with state-of-the-art accelerators. Binzhe Yuan, Xiangyu Zhang 0002, Yuefeng Zhang, Haochuan Wan, Zhechen Yuan, Junsheng Chen, Yunxiang He, Junran Ding, Chaolin Rao, Wenyan Su, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | A Multi-scale Block PatchMatch-based Unified Algorithm for Efficient 6-D Vision Processingabstract6-D vision, which combines stereo and flow estimation to derive 3-D location and motion information, is crucial for comprehensive real-world perception. This paper presents a multi-scale block PatchMatch (MBPM)-based efficient 6-D vision algorithm with a balance tradeoff between computational complexity and accuracy, making it suitable for embedded application scenarios. By adopting a random search strategy, the proposed algorithm effectively reduces the computational burden without compromising accuracy. Additionally, the block-level processing approach further enhances efficiency by processing data in smaller, manageable chunks. Due to the significantly larger search range in flow estimation compared to stereo, we further incorporate direction initialization and multi-scale PatchMatch propagation techniques. These enhancements are crucial for ensuring reliable accuracy while expediting algorithm convergence. Compared with existing 6-D vision algorithms, the proposed method achieves 16.7% higher accuracy in stereo and 11.1% in flow estimation with much lower complexity. Even when faced with larger searching ranges, the proposed method exhibits consistent accuracy, allowing it to effectively handle fast motion scenarios. Hongyu Wang 0010, Xiangyu Zhang 0002, Xin Lou 0001 |
ISCAS | 2 |
| 2024 | Feature Map Guided Adapter Network for Object Detection in Low-light ConditionsabstractConventional ISP pipelines and image enhancement methods are designed and optimized for human vision, creating a gap between the requirements of computer and human visions. To bridge the requirement gap, we present a co-design framework in which backend computer vision plays a pivotal role in shaping the proceeding image processing algorithm. It features a pre- processing adapter network, responsible for the restoration and enhancement of RAW images from computer vision perspective, especially in challenging environmental conditions. Specifically, we extract feature maps from the backend vision network, utilizing them as constraints for optimizing the preprocessing adapter network. To validate the effectiveness of our proposed framework, we employ object detection in low-light conditions as the computer vision task, with YOLO-v5 as the backbone. Given the considerable noise in low-light images, we compare our results with state-of-the-art denoising algorithms, showcasing the superior performance of our framework. Cong Pang, Wei Zhou 0037, Xiangyu Zhang 0002, Xin Lou 0001 |
ISCAS | 4 |
| 2024 | An Efficient Hardware Volume Renderer for Convolutional Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) has attracted growing attention in the fields of 3D reconstruction and rendering. However, straightforward NeRF algorithms encounter challenges in accurately capturing complex surface details with rich high-frequency information. A recent development known as Convolutional Neural Radiance Field (ConvNeRF) has demonstrated state-of-the-art results for these tasks. But it comes with substantial irregular computational requirements, particularly in the convolutional volume rendering phase. In this paper, we introduce a hardware accelerator designed to enhance the efficiency of convolutional volume rendering in ConvNeRF. Our approach includes the creation of specialized computation modules and corresponding on-chip memory system optimized for seamless support of gated convolutions and skip connections in ConvNeRF. To validate our design, we implement it in VerilogHDL and build a prototype using Field Programmable Gate Array (FPGA). We also map our design to 40nm CMOS technology. The evaluation results underscore the superiority of our accelerator in terms of energy efficiency when compared to an implementation on an NVIDIA 2080Ti GPU, offering approximately 84.6× more frames per watt. Xuexin Wang, Yunxiang He, Xiangyu Zhang 0002, Pingqiang Zhou, Xin Lou 0001 |
ISCAS | 3 |
| 2024 | RAW Images-Based Motion-Assisted Object Detection Accelerator Using Deformable Parts Models Features on 1080p VideosabstractThis paper introduces an end-to-end object detection hardware accelerator that directly processes RAW video signals to generate detection results, enabling a holistic approach to optimization. Unlike existing works that primarily concentrate on the back-end object detector, we explore the redundancy present across multiple stages of the processing pipeline such as the image signal processing (ISP), the temporal correlation in consecutive frames and the back-end detector. A prototype of Deformable Parts Models (DPM)-based accelerator has been successfully validated on the Altera TR5 field-programmable gate array (FPGA) platform. This accelerator demonstrates efficient processing of high-resolution ($1920\times 1080$) videos at 60 frames per second (FPS) while incorporating a 12-scale gradient pyramid and consuming only 130.9 KB blocks of memory. To optimize the search process for motion estimation, we adopt the time division multiplexing (TDM) technology, which effectively reduces both multiplexer usage and memory access. Compared to conventional methods that scan a 1080p frame, the proposed head-based motion search hardware consumes 6.82% of the processing cycles and utilizes merely 6.9 KB of block memory. Evaluation and comparison results demonstrate the effectiveness of the proposed system. Ling Zhang 0010, Xiangyu Zhang 0002, Xin Lou 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | Reconfigurable and Energy-Efficient Architecture for Deploying Multi-Layer RNNs on FPGAabstractRecurrent Neural Networks (RNNs) are extensively applied in sequence prediction tasks such as sentiment analysis, and machine translation. Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) are popular recurrent layers for mitigating gradient vanishing challenges. This paper introduces a reconfigurable hardware architecture that supports LSTM, GRU, and Fully Connected (FC) layers to accommodate the diverse structures of RNNs with minimal overhead. The proposed three-mode architecture utilizes dynamically reconfigurable components, allowing seamless mode switching among the three types of layers. Besides, Dynamic Compression (DC) for intermediate results is introduced to minimize precision loss, and layer decomposition as well as processing element grouping techniques are used to improve the processing efficiency. To validate the proposed architecture, a proof-of-concept prototype system using Intel Arria10 FPGA is built. Evaluation results demonstrate a 12% improvement over the state-of-the-art accelerator in energy efficiency when configured as GRU with a layer size set to$512{\times }$512, accompanied by a 93% reduction in block RAM and an 81% reduction in DSP resources. Xiangyu Zhang 0002, Yiren Zhu, Xin Lou 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | An Efficient Frequency Domain Vision Pipeline From RAW Images to Backend TasksabstractThough high resolution benefits computer vision performance, they are not commonly used in convolutional neural network (CNN)-based vision algorithms due to the limitation of memory and computation resource. Learning in the frequency domain makes high resolution images directly acceptable by CNNs, but the computation, time and energy overhead for pre-processing, including image signal processing (ISP) and domain transformation, can be large. This paper explores different image processing and domain transformation operations and proposes an efficient end-to-end frequency domain learning pipeline from RAW images to vision tasks. In particular, we simplify the pre-processing part by skipping the entire ISP pipeline and replacing the Discrete Cosine Transform (DCT) with a multiplication-free approximated one. Experimental results show that the final vision performance of the proposed pipeline is very close to that of the conventional pipeline, while significant amount of redundant operations can be saved. Wei Zhou 0037, Xiangyu Zhang 0002, Xin Lou 0001 |
ISCAS | 3 |
| 2022 | An End-to-end Computer Vision System ArchitectureabstractTo overcome the data movement bottleneck, near-sensor and in-sensor computing are becoming more and more popular. However, in the existing near-/in-sensor computing architectures for vision tasks, the effect of the image signal processing (ISP) pipeline, which is of great importance to the final vision performance [1], is always ignored. In this work, we propose a synthesized RAW image-based end-to-end computer vision paradigm, taking the effect of ISP pipeline into account. In the proposed approach, a generative adversarial network (GAN)-based tool is used to convert the fully processed color images to their corresponding RAW Bayer versions, generating the training data for end-to-end vision models. In the inference stage, RAW images from the sensor are directly fed to the end-to-end model, bypassing the entire ISP pipeline. Experimental results show that by training/tuning the CNN models using synthesized RAW images, it is possible to design an end-to-end (from RAW image to vision task) vision system that directly consumes RAW image data from the sensor with negligible vision performance degradation. By skipping the ISP pipeline, an image sensor can be directly integrated with the back-end vision processor without a complex image processor in the middle, making near-/in-sensor computing a practical approach. Ling Zhang 0010, Wei Zhou 0037, Xiangyu Zhang 0002, Xin Lou 0001 |
ISCAS | 3 |
| 2022 | A Block PatchMatch-Based Energy-Resource Efficient Stereo Matching Processor on FPGAabstractThis paper presents a field programmable gate array (FPGA)-based, high-performance and energy-resource efficient stereo matching processor. The proposed processor executes block-level PatchMatch-based stereo matching algorithm with a random search strategy to avoid estimation of all disparity levels. To take advantages of different block scales, a coarse-to-fine multi-scale propagation (MSP) scheme is proposed for label update. Based on that, a dedicated hardware architecture is further proposed to explore the benefit of the algorithm. Experimental results show that the proposed FPGA-based processor, running at 350MHz, achieves a peak performance of$1920\times 1080.165$.7 frame per second (FPS) at 128 disparity levels with 3.35W power dissipation. The energy and resource efficiency of the proposed design outperforms state-of-the-art FPGA-based stereo matching processors. When disparity level increases to 256, the computing resource increment of the proposed design is much less than existing designs because random search instead of winner-takes-all (WTA) is utilized. Moreover, unlike existing dedicated stereo matching processors which output only disparity information, the proposed design is also capable of deriving plane slant. This information can be beneficial for follow-up tasks like 3D reconstruction. Hongyu Wang 0010, Wei Zhou 0037, Xiangyu Zhang 0002, Xin Lou 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | A Raw Image-Based End-to-End Object Detection Accelerator Using HOG FeaturesabstractThis paper presents an end-to-end object-detection accelerator that processes raw Bayer images to generate detection results. The accelerator utilizes histogram of oriented gradients (HOG) features in combination with a support vector machine (SVM) classifier. The proposed HOG for raw images (HOGR) skips the image signal processors which consume a significant amount of power (about 2.5X of an accelerator). The proposed architecture temporally partitions the algorithms and time multiplexes the logic such that the accelerator works on the same frequency with the image sensor using affordable resources. The prototype is verified using 1080p ($1920\times 1080$) and VGA ($640\times 480$) raw videos on Altera Arria10 and Cyclone IV field programmable gate array (FPGA) platforms. The accelerator can process 1080p raw videos with a 12-scale pyramid at 60 frames per second (FPS) under the pixel frequency of corresponding image sensor (148.5 MHz), consuming 510 Kbit block memory. To the best knowledge of the authors, this is the first end-to-end HOG+SVM accelerator that takes raw Bayer images as input, skips the ISP pipeline for resource optimization from the system perspective, and is synchronized with the image sensor. Xiangyu Zhang 0002, Ling Zhang 0010, Xin Lou 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2018 | A Hardware Architecture for Cell-Based Feature-Extraction and Classification Using Dual-Feature SpaceabstractMany computer-vision and machine-learning applications in robotics, mobile, wearable devices, and automotive domains are constrained by their real-time performance requirements. This paper reports a dual-feature-based object recognition coprocessor that exploits both histogram of oriented gradient (HOG) and Haar-like descriptors with a cell-based parallel sliding-window recognition mechanism. The feature extraction circuitry for HOG and Haar-like descriptors is implemented by a pixel-based pipelined architecture, which synchronizes to the pixel frequency from the image sensor. After extracting each cell feature vector, a cell-based sliding window scheme enables parallelized recognition for all windows, which contain this cell. The nearest neighbor search classifier is, respectively, applied to the HOG and Haar-like feature space. The complementary aspects of the two feature domains enable a hardware-friendly implementation of the binary classification for pedestrian detection with improved accuracy. A proof-of-concept prototype chip fabricated in a 65-nm SOI CMOS, having thin gate oxide and buried oxide layers (SOTB CMOS), with 3.22-mm2core area achieves an energy efficiency of 1.52 nJ/pixel and a processing speed of 30 fps for 1024 × 1616-pixel image frames at 200-MHz recognition working frequency and 1-V supply voltage. Furthermore, multiple chips can implement image scaling, since the designed chip has image-size flexibility attributable to the pixel-based architecture. Fengwei An, Xiangyu Zhang 0002, Aiwen Luo, Lei Chen 0001, Hans Jürgen Mattausch |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Resource-Efficient Object-Recognition Coprocessor With Parallel Processing of Multiple Scan Windows in 65-nm CMOSabstractObject recognition offers a more general implementation for vision-based applications. This paper reports a resource-efficient recognition coprocessor with embedded cell-based simplified speeded up robust feature descriptor extraction unit and parallel scan-window (SW) recognition engine, applicable for various mobile scenarios and image sensor types. The feature extraction circuitry with pixel-based pipelined architecture describes the target objects among complex backgrounds, only relying on the pixel frequency from the image sensor. A cell-based SW algorithm enables parallelized recognition in multiple SWs and compatibility to different image sizes. The proposed hardware-friendly object-recognition coprocessor was implemented in 65-nm Silicon on thin BOX CMOS technology with 1.26 mm2core area and can operate down to low supply voltage of 0.5 V. For video graphics array image sizes, the energy efficiency is determined as 910 μJ per frame at 200 MHz and 1-V supply voltage. The coprocessor's classification performance is demonstrated for pedestrian and car detection. Aiwen Luo, Fengwei An, Xiangyu Zhang 0002, Lei Chen 0001, Hans Jürgen Mattausch |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Dynamically reconfigurable system for LVQ-based on-chip learning and recognitionabstractArtificial neural networks implement a simplified model of the human brain and thus specialize on pattern recognition. As an alternative to conventional single-instruction-multiple-data (SIMD) solutions with massive parallelism for self-organizing-map (SOM) neural network models, we report resource-efficient hardware architecture for 1-chip implementation of the learning vector quantization (LVQ) neural network algorithm, which is a variant of SOM. Dynamic configurability for two operation modes is realized through the same circuitry for recognition based on nearest-neighbor matching and on-chip learning based on error back-propagation. Switching between learning and recognition modes is carried out by a pipeline with multiplexers and parallel p-word input (P-MPPI). The multiplexers enable data-flow-path reconfiguration, resulting in a significant reduction of area and power consumption. Thus, the P-MPPI architecture achieves time-domain multiplexing between operation modes as well as area/energy-efficiency by reusing both memory arrays and arithmetic or logic units. Additionally, high flexibility for feature-vector dimension and reference-vector number allows the implementation of many different applications, including continuously adaptive neural systems, on the same hardware platform. A test chip in TSMC 65 nm CMOS has parallel 32-word inputs, 585 K-bit on-chip memory, and achieves high processing throughput of 76.8 Gbps and low power consumption of 27.92 mW (at 150 MHz, 1.0 V supply voltage). Fengwei An, Xiangyu Zhang 0002, Lei Chen 0001, Hans Jürgen Mattausch |
ISCAS | 2 |