EDBT 2026 Demo / reviewers in the wild / expert
Fengwei An
dblp:58/11184
· DBLP profile ↗
23ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-7554-7938ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A High-Precision CORDIC Architecture with Reduced-Convergence Preprocessing and Dynamic Logarithmic Bit-Width Scaling
Zhentao Zheng, Jiaqi Ouyang, Fengwei An |
ISCAS | 6 |
| 2026 | The limits of bio-molecular modeling with large language models: a cross-scale evaluationabstractMOTIVATION: The modeling of bio-molecular system across molecular scales remains a central challenge in scientific research. Large language models (LLMs) are increasingly applied to bio-molecular discovery, yet systematic evaluation across multi-scale biological problems and rigorous assessment of their tool-augmented capabilities remain limited. RESULTS: We reveal a systematic gap between LLM performance and mechanistic understanding through the proposed cross-scale bio-molecular benchmark: BioMol-LLM-Bench, a unified framework comprising 26 downstream tasks that covers 4 distinct difficulty levels, and computational tools are integrated for a more comprehensive evaluation. Evaluation on 13 representative models reveals 4 benchmark-specific observations: chain-of-thought-style training does not consistently improve performance on the evaluated biological tasks; the evaluated hybrid mamba-attention model shows strong performance on long bio-molecular sequence tasks; supervised fine-tuned models show task-specific specialization with reduced performance in some general settings; and current LLMs perform better on classification tasks than on challenging regression tasks under this benchmark setting. AVAILABILITY: Source code is available at https://github.com/AI-HPC-Research-Team/BioMol-LLM-Bench. Yaxin Xu, Yue Zhou 0014, Zhengyu Ma, Fengwei An, Zhixiang Ren |
Bioinform. | 5 |
| 2025 | Pose as Clinical Prior: Learning Dual Representations for Scoliosis Screening
Zirui Zhou, Zizhao Peng, Dongyang Jin, Chao Fan 0001, Fengwei An, Shiqi Yu 0001 |
MICCAI (13) | 5 |
| 2025 | A Reconfigurable Floating-Point Division and Square Root Architecture for High-Precision SoftmaxabstractWith the advancement of deep learning models, the Softmax function with self-attention has become pervasive in everyday applications. As components of the Softmax function and its inputs, both division and square root operations impact its accuracy. However, these two non-linear operations bring significant area and power consumption for hardware implementation. To address these challenges, this paper proposes a reconfigurable floating-point division and square root (FDSR) architecture that achieves low resource consumption and high accuracy for general-purpose computation. The FDSR enhances the traditional non-restoring algorithm by using shift-registers and optimizing the leading-one detection and shift operations, reducing hardware resource usage while maintaining high accuracy (0.5 ULP). In the mantissa calculation, the division can be converted to a square root operation by simply switching the input to the subtractor through multiplexers. Additionally, a triple-mode reconfigurable iteration unit is introduced, featuring a multi-layer variable pipeline architecture to improve adaptability for different applications. By redesigning the pipeline depth and reusing logical units, the FDSR effectively addresses the issue of lengthy iteration cycles in the non-restoring method. Implementation results using 40nm CMOS technology demonstrate that the proposed design achieves a 76.49% power reduction and a 14.69% area reduction for floating-point division compared to Synopsys Design Ware and an 88.05% power reduction and a 90.57% area reduction for floating-point square root. With 28 nm CMOS technology, the FDSR reduces power consumption by 91.55% and reduces area by 64.39% for floating-point division compared to Synopsys Design Ware. On the FPGA platform, the FDSR significantly reduces hardware resource consumption, achieving an 85.23% reduction for floating-point division and 87.81% for floating-point square root, outperforming state-of-the-art designs. Xiwei Fang, Fengwei An |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | Dual-Thread Deflate/Inflate Accelerator With Multicheckpoint Control With High Throughput and Compression Ratio for Bandwidth-Efficient SystemsabstractWith the exponential growth of data volumes in AI training and prediction systems, the cost and resource demands of data transmission have emerged as critical challenges. Lossless data compression effectively reduces data size, transmission bandwidth, and latency while preserving data integrity. This article presents a fully pipelined lossless CODEC integrating Deflate compression and Inflate decompression accelerators. The proposed Deflate implementation employs match filtering and pair merging strategies to enhance compression ratios. We introduce three key innovations for the Inflate decompressor: 1) a dual-thread architecture with multicheckpoint control; 2) optimized end-of-block (EOB) handling in Huffman coding; and 3) a rewinding mechanism in LZ77 decoding. FPGA implementation results demonstrate that our Deflate compressor achieves 16 bytes/cycle throughput with an average compression ratio of 2.26, surpassing state-of-the-art implementations. The 28 nm CMOS implementation shows Inflate decompression throughputs of 1431.85 MB/s (dynamic Huffman) and 1324.26 MB/s (static Huffman) on the Calgary Corpus dataset. Notably, our 28 nm CMOS-based decompressor achieves$1.16\times $higher throughput than recent 14 nm implementations in spite of operating at half their maximum frequency. Yiwei Luo, Jiaqi Ouyang, Lei Chen 0001, Fengwei An |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | ReHIT: Reconfigurable High-Radix Iterative-Taylor Architecture for Ultraprecise Logarithm/Exponential Functions in FPGA-Based Softmax AcceleratorsabstractThe softmax function, as a pivotal component in neural network accelerators, imposes stringent demands on the precision-efficiency tradeoff for logarithmic and exponential computations. This article presents reconfigurable high-radix iterative-Taylor (ReHIT) architecture, a novel hardware framework that synergistically integrates high-radix iterative normalization with optimized Taylor expansion to achieve subunit-in-the-last-place (ULP) precision in floating-point transcendental functions. Our key innovation lies in the hierarchical pretreatment mechanism where high-radix iterations (radix-256/512) systematically decompose input operands into normalized subdomains, enabling subsequent quadratic Taylor approximations with guaranteed convergence. This codesign methodology reduces polynomial orders by 33% compared to conventional approaches while eliminating resource-intensive division operations through shift-and-add transformations. The implemented ReHIT-logarithm (ReHIT-L) and ReHIT-exponential (ReHIT-E) modules demonstrate configurable precision scaling from half to double precision (FP16/32/64), validated through exhaustive error analysis over 1 000 000 random test vectors with worst case errors bounded at 0.78 ULP. Field-programmable gate array (FPGA) implementations on Arria 10/Virtex-7 platforms can achieve up to 13.8% logic resources reduction and 32.3% latency improvement over state-of-the-art designs, with post-synthesis results in Taiwan Semiconductor Manufacturing Company (TSMC) 28-nm showing up to$1.34\times $giga operations per second (GOPS)/W energy efficiency and$3.97\times $GOPS/mm2area efficiency for softmax acceleration. The reconfigurable pipeline of the architecture permits dynamic precision/throughput adaptation, particularly beneficial for quantized neural networks requiring FP16–FP32 hybrid precision. Yangyi Zhang 0002, Lei Chen 0001, Fengwei An |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | Graph convolutional networks for 3D skeleton-based scoliosis screening using gait sequencesabstractAbstract Adolescent idiopathic scoliosis is a significant health concern, ranked as the third most prevalent issue among adolescents after obesity and myopia. Traditional screening methods rely on the use of complex and expensive measuring instruments and expert physicians to interpret X-ray images. These methods can be both time-consuming and inaccessible for widespread screening efforts. To address these challenges, we propose a standardized protocol for the collection of scoliosis gait dataset. This protocol enables the systematic capture of relevant gait characteristics associated with scoliosis, leading to the creation of a comprehensive, annotated dataset tailored for research and diagnostic purposes. Leveraging this dataset, we developed an effective deep learning algorithm based on graph convolutional networks, which outperforms traditional CNN by effectively modeling the complex spatial and temporal dynamics of human gait and posture, leveraging skeletal structure as a graph for more accurate and robust scoliosis screening. We also explored various optimization strategies to enhance the model’s accuracy and efficiency, ensuring robust performance across diverse scenarios. Our innovative approach allows for the rapid and non-invasive recognition of scoliosis. This method is not only scalable but also eliminates the need for specialized equipment or extensive medical expertise, making it ideal for large-scale screening initiatives. By improving the accessibility and efficiency of scoliosis detection, our approach has the potential to facilitate early intervention. Zizhao Peng, Mengying Sun, Yan Wang 0116, Ping Li 0016, Fengwei An |
Vis. Comput. | 7 |
| 2024 | Genetic Quantization-Aware Approximation for Non-Linear Operations in TransformersabstractNon-linear functions are prevalent in Transformers and their lightweight variants, incurring substantial and frequently underestimated hardware costs. Previous state-of-the-art works optimize these operations by piece-wise linear approximation and store the parameters in look-up tables (LUT), but most of them require unfriendly high-precision arithmetics such as FP/INT 32 and lack consideration of integer-only INT quantization. This paper proposed a genetic LUT-Approximation algorithm namely GQA-LUT that can automatically determine the parameters with quantization awareness. The results demonstrate that GQA-LUT achieves negligible degradation on the challenging semantic segmentation task for both vanilla and linear Transformer models. Besides, proposed GQA-LUT enables the employment of INT8-based LUT-Approximation that achieves an area savings of 81.3~81.7% and a power reduction of 79.3~80.2% compared to the high-precision FP/INT 32 alternatives. Code is available at https://github.com/PingchengDong/GQA-LUT. Pingcheng Dong, Yonghao Tan, Tianwei Ni, Yu Liu 0007, Luhong Liang, Shih-Yang Liu, Xijie Huang, Huaiyu Zhu 0004, Fengwei An, Kwang-Ting Cheng |
DAC | 13 |
| 2024 | BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity AllocationabstractLarge language models (LLMs) have demonstrated outstanding performance in various tasks, such as text summarization, text question-answering, and etc. While their performance is impressive, the computational footprint due to their vast number of parameters can be prohibitive. Existing solutions such as SparseGPT and Wanda attempt to alleviate this issue through weight pruning. However, their layer-wise approach results in significant perturbation to the model's output and requires meticulous hyperparameter tuning, such as the pruning rate, which can adversely affect overall model performance. To address this, this paper introduces a novel LLM pruning technique dubbed blockwise parameter-efficient sparsity allocation (BESA) by applying a blockwise reconstruction loss. In contrast to the typical layer-wise pruning techniques, BESA is characterized by two distinctive attributes: i) it targets the overall pruning error with respect to individual transformer blocks, and ii) it allocates layer-specific sparsity in a differentiable manner, both of which ensure reduced performance degradation after pruning. Our experiments show that BESA achieves state-of-the-art performance, efficiently pruning LLMs like LLaMA1, and LLaMA2 with 7B to 70B parameters on a single A100 GPU in just five hours. Code is available at [here](https://github.com/LinkAnonymous/BESA). Peng Xu 0035, Wenqi Shao, Mengzhao Chen, Shitao Tang, Kaipeng Zhang, Peng Gao 0007, Fengwei An, Yu Qiao 0001, Ping Luo 0002 |
ICLR | 7 |
| 2024 | Live Demonstration: A 1920×1080 129fps 4.3pJ/pixel Stereo-Matching Processor for Low-power ApplicationsabstractThis demonstration presents an advanced stereo vision system with high energy efficiency. An ov5640 binocular camera, operating at a maximum of 30 frames/second with FHD (1920×1080) resolution, is employed to capture image pairs. A Spartan-7 FPGA rectifies these images with a calibration map matrix and then channels the pixel stream to the stereo-matching processor in a 28nm CMOS process for depth estimation. The resulting depth map, crucial for tasks like obstacle detection and navigation, is displayed in real-time on the monitor for low-power stereo vision applications. Zhuoyu Chen, Shengming Zhou, Pingcheng Dong, Fengwei An, Lei Chen 0001 |
ISCAS | 6 |
| 2024 | Live Demonstration: A Video Denoising Co-processor with Non-local Means Algorithm for FHD 30fps Image SensorabstractIn this demonstration, a non-local means (NLM) video denoising co-processor with data reuse scheme and dual-clock domain for high resolution image sensor is presented. With an OV5640 camera, the real-time denoising processor can be performed at 30 frames/second for full high definition (FHD 1920×1080) RAW video, a debayer filter then decodes the RAW format to RGB format. Ruoheng Yao, Shengming Zhou, Zhiyue Gao, Yangyi Zhang 0002, Yiwei Luo, Lei Chen 0001, Fengwei An |
ISCAS | 7 |
| 2024 | Gait Patterns as Biomarkers: A Video-Based Approach for Classifying Scoliosis
Zirui Zhou, Zizhao Peng, Chao Fan 0001, Fengwei An, Shiqi Yu 0001 |
MICCAI (5) | 5 |
| 2024 | Stereo Matching Accelerator With Re-Computation Scheme and Data-Reused Pipeline for Autonomous VehiclesabstractBinocular stereo vision is a depth estimation technique by imitating human eyes. It is widely used in various fields, such as self-driving cars, SLAM, and 3D reconstruction. However, designing a hardware architecture that can balance resource utilization, processing speed, and estimation accuracy remains a significant challenge. This paper proposes a compact and efficient hardware-based design that incorporates linear fitting-based cost fusion, disparity optimization with subpixel interpolation, and multi-directional occlusion filling techniques. Firstly, a gradient-enhanced pipelined matching costs architecture with a resource-saving scheme and re-computation paradigm is proposed to improve the accuracy of edge information. Then, we approximate the nonlinear exponential function by linear fitting to save the hardware resource. Moreover, we design the subpixel interpolation with an SRT radix-4 divider to refine the disparity, which significantly enhances the accuracy of the disparity map in real situations. Finally, we proposed a resource-reused architecture for synchronous hole filling and median filter in post-processing. The disparity map quality of the proposed architecture is evaluated on KITTI2015 datasets, which delivers leading accuracy compared to other state-of-the-art works. The architecture has been successfully implemented and demonstrated on the Stratix-V FPGA platform and achieved 54 frames per second operating at 112 MHz under a resolution of$1920\times 1080$. Xiwei Fang, Yunhao Ma, Pingcheng Dong, Zhuoyu Chen, Lei Chen 0001, Fengwei An |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2023 | An Energy-Efficient, Resource-Efficient and High Frame-Rate End-to-End Pedestrian Detector Using HOG-SVM for Intelligent Edge DevicesabstractThis paper proposes a Histogram of Oriented Gradients-Support Vector Machine (HOG-SVM) based pedestrian detector with an end-to-end fully-pipelined architecture to achieve a high frame rate by improving the throughput, and reduce the power consumption by minimizing the data movement. To further improve the energy efficiency under the high frame rate, a bit-width pruning method is used to remove the gray-scale converter's redundant data bit width, and a block-score normalization is employed to significantly reduce the normalizer's required divisions. The reduced computation amount also saves the hardware overhead while maintaining the same calculation accuracy. Besides, a modeling and analysis method of the SVM-classifier-Multiply-ACcumulate (MAC) array is proposed to further improve the energy efficiency and save the logic resources, by optimizing the array size with a hardware utilization of 98.4% while maintaining the same throughput. The FPGA implementation results of$640\times 480$video show a high frame rate of up to 439 fps @143 MHz and a high energy efficiency of 0.76 nJ/pixel with 46.7% fewer LUTs, 22.4% fewer registers, 88.3% fewer DSPs, compared to the state-of-the-art design. The ASIC implementation in 55 nm also confirms a high energy efficiency of 0.35 nJ/pixels at 613 fps and 200 MHz as well as a hardware overhead of 177 k gates and 108 Kbits SRAM. Jianhui Song, Bingqiang Liu, Zixuan Shen, Fengwei An, Chao Wang 0096, Jiang Tang |
IECON | 6 |
| 2023 | A Spatio-Temporal Video Denoising Co-Processor With Adaptive CodecabstractWith the increasing demand for high-resolution video and real-time processing, the limited efficiency of video-denoising algorithms has become a critical factor. This paper proposes a spatio-temporal video denoising co-processor to suppress an image sequence’s spatial and temporal noise. Temporal denoising is achieved by merging the current and previous frames at the pixel level in which the current frame is processed by a spatial filter. After exploiting noise estimation and motion detection, the Wiener filter calculates the merge ratio. Rather than buffering the entire previous frame, the JPEG-like codec can dynamically adjust the compression ratio through a predefined quantization table to satisfy the designed on-chip storage. The experimental results demonstrate that the spatio-temporal denoising co-processor can effectively eliminate the fluctuation of the grayscale value of the noise in videos. Simultaneously, the adaptive codec can reduce the storage space consumption for the frame buffer by at least 80% of the original size. To the best of our knowledge, this is the first fully integrated spatio-temporal denoising co-processor without any external memory. Additionally, the grayscale, RGB, and RAW versions of the co-processor are also implemented on the Stratix V FPGA platform and synthesized in 28nm CMOS technology. Yichen Ouyang, Ruoheng Yao, Zhuoyu Chen, Lei Chen 0001, Fengwei An |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2023 | Anti-Aliasing and Anti-Color-Artifact Demosaicing for High-Resolution CMOS Image SensorabstractDemosaicing is a technique that reconstructs an RGB image from fragmentary color samples sensed by the image sensor. The color filter array (CFA), which is placed over the image sensor, determines the color of each pixel. The most used color filter array is the Bayer CFA. This paper proposes an Anti-Aliasing and Anti-Color-Artifact Demosaicing (AAACA) algorithm for the Bayer pattern and the resource-efficient very-large-scale integration (VLSI) architecture for the proposed algorithm. The AAACA comprises an anti-aliasing approach and a color artifacts filter named color difference-based median filter (CDMF). Compared to the traditional demosaicing methods that equally treat pixels on and not on edges, the AA reconstructs the pixels on edges with different strategies from those not on edges, significantly removing the aliasing around recovered edges. Then the CDMF is utilized to remove the color artifacts of the reconstructed RGB image based on the color difference after median filtering. We respectively simulate the quantitative evaluation and subjective visual quality on McMaster and Kodak datasets. Our experiments reveal that the proposed AAACA algorithm can significantly remove the visual aliasing around recovered edges and greatly reduce color artifacts in demosaiced images compared to the state-of-art demosaicing algorithms. The proposed VLSI architectures can achieve superior visual qualities compared with the previous VLSI implementations under the same process technology conditions with 180nm CMOS technology, an image resolution of$1280\times 720$(HD), and a working frequency of 200MHz. Yangyi Zhang 0002, Zizhao Peng, Lei Chen 0001, Fengwei An |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | A 4.29nJ/pixel Stereo Depth Coprocessor With Pixel Level Pipeline and Region Optimized Semi-Global Matching for IoT ApplicationabstractThe semi-global matching (SGM) algorithm in stereo vision is a well-known depth-estimation method since it can generate dense and robust disparity maps. However, the real-time processing and low power dissipation, the specifications of the Internet-of-Thing (IoT) applications, are challenging for their computational complexity. In this paper, we propose a hardware-oriented SGM algorithm with pixel-level pipeline and region-optimized cost aggregation for high-speed processing and low hardware-resource usage. Firstly, the matching costs in a region are integrated with an optimization strategy to significantly reduce memory usage and improve the processing speed of the cost aggregation. Then, a two-layer parallel two-stage pipeline (TPTP) architecture, which enables pixel-level processing, is designed to calculate two directions (0° and 135°) aggregation to further solve the crucial computational bottleneck of the SGM algorithm. Finally, the architecture is demonstrated on a low-cost XILINX Spartan-7 device and an advanced Stratix-V FPGA device for VGA ($640\times 480$) depth estimation. The experimental results show that the proposed architecture with compact hardware architecture also ensures accuracy. The pixel-level pipeline architecture enables a processing speed of 355 frames per second (fps) at 109MHz on the Spartan-7 FPGA device and 508 fps at 156MHz on the Stratix-V FPGA. Besides, the coprocessor respectively achieves an energy efficiency of 4.74 nJ/pixel with a power dissipation of 517mW and 4.29nJ/pixel with a power dissipation of 669mW on these two FPGAs. Pingcheng Dong, Zhuoyu Chen, Zhuoao Li, Yuzhe Fu, Lei Chen 0001, Fengwei An |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2020 | Energy-Efficient Machine Learning Accelerator for Binary Neural NetworksabstractBinary neural network (BNN) has shown great potential to be implemented with power efficiency and high throughput. Compared with its counterpart, the convolutional neural network (CNN), BNN is trained with binary constrained weights and activations, which are more suitable for edge devices with less computing and storage resource requirements. In this paper, we introduce the BNN characteristics, basic operations and the binarized-network optimization methods. Then we summarize several accelerator designs for BNN hardware implementation by using three mainstream structures, i.e., ReRAM-based crossbar, FPGA and ASIC. Based on the BNN characteristics and hardware custom designs, all these methods achieve massively parallelized computations and highly pipelined data flow to enhance its latency and throughput performance. In addition, the intermediate data with the binary format are stored and processed on chip by constructing the computing-in-memory (CIM) architecture to reduce the off-chip communication costs, including power and latency. Wei Mao 0002, Zhihua Xiao, Peng Xu 0035, Dingbang Liu, Shirui Zhao, Fengwei An, Hao Yu 0001 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2018 | A Hardware Architecture for Cell-Based Feature-Extraction and Classification Using Dual-Feature SpaceabstractMany computer-vision and machine-learning applications in robotics, mobile, wearable devices, and automotive domains are constrained by their real-time performance requirements. This paper reports a dual-feature-based object recognition coprocessor that exploits both histogram of oriented gradient (HOG) and Haar-like descriptors with a cell-based parallel sliding-window recognition mechanism. The feature extraction circuitry for HOG and Haar-like descriptors is implemented by a pixel-based pipelined architecture, which synchronizes to the pixel frequency from the image sensor. After extracting each cell feature vector, a cell-based sliding window scheme enables parallelized recognition for all windows, which contain this cell. The nearest neighbor search classifier is, respectively, applied to the HOG and Haar-like feature space. The complementary aspects of the two feature domains enable a hardware-friendly implementation of the binary classification for pedestrian detection with improved accuracy. A proof-of-concept prototype chip fabricated in a 65-nm SOI CMOS, having thin gate oxide and buried oxide layers (SOTB CMOS), with 3.22-mm2core area achieves an energy efficiency of 1.52 nJ/pixel and a processing speed of 30 fps for 1024 × 1616-pixel image frames at 200-MHz recognition working frequency and 1-V supply voltage. Furthermore, multiple chips can implement image scaling, since the designed chip has image-size flexibility attributable to the pixel-based architecture. Fengwei An, Xiangyu Zhang 0002, Aiwen Luo, Lei Chen 0001, Hans Jürgen Mattausch |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Resource-Efficient Object-Recognition Coprocessor With Parallel Processing of Multiple Scan Windows in 65-nm CMOSabstractObject recognition offers a more general implementation for vision-based applications. This paper reports a resource-efficient recognition coprocessor with embedded cell-based simplified speeded up robust feature descriptor extraction unit and parallel scan-window (SW) recognition engine, applicable for various mobile scenarios and image sensor types. The feature extraction circuitry with pixel-based pipelined architecture describes the target objects among complex backgrounds, only relying on the pixel frequency from the image sensor. A cell-based SW algorithm enables parallelized recognition in multiple SWs and compatibility to different image sizes. The proposed hardware-friendly object-recognition coprocessor was implemented in 65-nm Silicon on thin BOX CMOS technology with 1.26 mm2core area and can operate down to low supply voltage of 0.5 V. For video graphics array image sizes, the energy efficiency is determined as 910 μJ per frame at 200 MHz and 1-V supply voltage. The coprocessor's classification performance is demonstrated for pedestrian and car detection. Aiwen Luo, Fengwei An, Xiangyu Zhang 0002, Lei Chen 0001, Hans Jürgen Mattausch |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Dynamically reconfigurable system for LVQ-based on-chip learning and recognitionabstractArtificial neural networks implement a simplified model of the human brain and thus specialize on pattern recognition. As an alternative to conventional single-instruction-multiple-data (SIMD) solutions with massive parallelism for self-organizing-map (SOM) neural network models, we report resource-efficient hardware architecture for 1-chip implementation of the learning vector quantization (LVQ) neural network algorithm, which is a variant of SOM. Dynamic configurability for two operation modes is realized through the same circuitry for recognition based on nearest-neighbor matching and on-chip learning based on error back-propagation. Switching between learning and recognition modes is carried out by a pipeline with multiplexers and parallel p-word input (P-MPPI). The multiplexers enable data-flow-path reconfiguration, resulting in a significant reduction of area and power consumption. Thus, the P-MPPI architecture achieves time-domain multiplexing between operation modes as well as area/energy-efficiency by reusing both memory arrays and arithmetic or logic units. Additionally, high flexibility for feature-vector dimension and reference-vector number allows the implementation of many different applications, including continuously adaptive neural systems, on the same hardware platform. A test chip in TSMC 65 nm CMOS has parallel 32-word inputs, 585 K-bit on-chip memory, and achieves high processing throughput of 76.8 Gbps and low power consumption of 27.92 mW (at 150 MHz, 1.0 V supply voltage). Fengwei An, Xiangyu Zhang 0002, Lei Chen 0001, Hans Jürgen Mattausch |
ISCAS | 1 |
| 2013 | K-means clustering algorithm for multimedia applications with flexible HW/SW co-design
Fengwei An, Hans Jürgen Mattausch |
J. Syst. Archit. | 1 |
| 2012 | Cluster-Based Prototype Learning System for Multiple Applications with Flexible HW/SW CodesignabstractThis paper proposes a novel hybrid hardware-software (HW/SW) system for K-means-based prototype learning and Nearest-Neighbor (1-NN) classification. We implement a prototype learning system instead of simplifying complex learning algorithms (e.g. neural and fuzzy networks, or SVMs) because this facilitates the adaptability to hardware capabilities and constraints. The K-means algorithm, which is implemented by HW/SW co-design, is effective in improving classification performance and reducing storage requirements. Particularly, the hardware realization is applied to obtain orders of magnitude higher speed for nearest-distance searching, which is the most burdensome performance barrier both in K-means learning and 1-NN classification. We benchmark our multi-purpose learning system against the application of handwritten digit recognition and face recognition to demonstrate its excellent performance, namely high flexibility, fast training, short recognition time and good recognition rate. Fengwei An, Hans Jürgen Mattausch |
PDCAT | 1 |