EDBT 2026 Demo / reviewers in the wild / expert
Jiawei Xu 0002
dblp:79/8798-2
· DBLP profile ↗
9ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-6192-558XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SIMBRAIN: A nonidealities-aware simulation framework for spiking neural networks based on memristor crossbars
Jiawei Xu 0002, Ruisi Shen, Dimitrios Stathis 0001, Lirong Zheng 0001, Zhuo Zou, Ahmed Hemani |
Neurocomputing | 1 |
| 2025 | MemMIMO: A Simulation Framework for Memristor-Based Massive MIMO AccelerationabstractMemristor-based crossbar architectures have proven highly effective for matrix vector multiplication (MVM) operations, making them a promising solution for accelerating the MVMs widely used in precoding algorithms for multiple input multiple output (MIMO) wireless communication systems. However, real-world implementation of memristor-based computing systems face challenges due to commonly observed non-idealities in both the devices themselves and the circuits they’re built into. To facilitate a rapid design flow and investigate the impact of non-idealities, an integrated open-source simulation framework MemMIMO is developed. The simulation framework estimates the accuracy and hardware performance of the computing system, offering a variety of flexible design options. MemMIMO integrates a behavioral model of the mix-signal architecture with a digital front-end. There are three major building blocks in MemMIMO: the device fitting block, the mapping block, and the performance estimation block. These blocks work together to map the complex MVMs in precoding algorithms for MIMO systems to crossbar-based architectures that incorporate memristor models characterized by physical device behavior. Using two typical use cases targeting six-generation (6G) massive MIMO communication as case studies, MemMIMO is used to model different memristor devices, explore the impact of non-idealities on system accuracy, and benchmark circuit-level performance metrics including area, speed, and power. Jiawei Xu 0002, Dimitrios Stathis 0001, Ruisi Shen, Lirong Zheng 0001, Zhuo Zou, Ahmed Hemani |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | FPGA-Based HPC for Associative Memory SystemabstractAssociative memory plays a crucial role in the cognitive capabilities of the human brain. The Bayesian Confidence Propagation Neural Network (BCPNN) is a cortex model capable of emulating brain-like cognitive capabilities, particularly associative memory. However, the existing GPU-based approach for BCPNN simulations faces challenges in terms of time overhead and power efficiency. In this paper, we propose a novel FPGA-based high performance computing (HPC) design for the BCPNN-based associative memory system. Our design endeavors to maximize the spatial and timing utilization of FPGA while adhering to the constraints of the available hardware resources. By incorporating optimization techniques including shared parallel computing units, hybrid-precision computing for a hybrid update mechanism, and the globally asynchronous and locally synchronous (GALS) strategy, we achieve a maximum network size of $150 \times 10$ and a peak working frequency of 100 MHz for the BCPNN-based associative memory system on the Xilinx Alveo U200 Card. The tradeoff between performance and hardware overhead of the design is explored and evaluated. Compared with the GPU counterpart, the FPGA-based implementation demonstrates significant improvements in both performance and energy efficiency, achieving a maximum latency reduction of $33.25 \times$, and a power reduction of over $6.9 \times$, all while maintaining the same network configuration. Yu Yang 0020, Dimitrios Stathis 0001, Ahmed Hemani, Anders Lansner, Jiawei Xu 0002, Lirong Zheng 0001, Zhuo Zou |
ASPDAC | 7 |
| 2023 | ASLog: An Area-Efficient CNN Accelerator for Per-Channel Logarithmic Post-Training QuantizationabstractPost-training quantization (PTQ) has been proven an efficient model compression technique for Convolution Neural Networks (CNNs), without re-training or access to labeled datasets. However, it remains challenging for a CNN accelerator to fulfill the efficiency potential of PTQ methods. A large number of PTQ techniques blindly pursue high theoretic compression effect and accuracy, ignoring their impact on the actual hardware implementation, which causes more hardware overhead than benefit. This paper introduces ASLog, a PTQ-friendly CNN accelerator that explores four key designs in an algorithm-hardware co-optimizing manner: the first practical 4-bit logarithmic PTQ pipeline SLogII, the multiplier-free arithmetic element (AE) design, the energy-efficient bias correction element (BCE) design, and the per-channel quantization friendly (PCF) architecture and dataflow. The proposed SLogII PTQ pipeline can push the limit of logarithmic PTQ to 4-bit with40% lower in power and area consumption compared with a common 8-bit multiplier. The BCE and PCF design proposed in this paper are the first to consider the hardware impact of the widely-used per-channel quantization and bias correction technique, enabling an efficient PTQ-friendly implementation with a small hardware overhead. The ASLog is validated in a UMC 40-nm process, with 12.2 TOPS/W energy efficiency and 0.80 mm2 core area. The ASLog can achieve 336.3 GOPS/mm2 area efficiency and >500 OPs/Byte operational intensity, which map to over$1.85\times $and$1.12\times $improvement compared with the previous related works. Jiawei Xu 0002, Jiangshan Fan, Baolin Nan, Chen Ding 0010, Lirong Zheng 0001, Zhuo Zou, Yuxiang Huan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | Edge-Based Collaborative Training System for Artificial Intelligence-of-ThingsabstractThe descending of intelligence from the cloud to the heterogeneous and low-power edge in the Artificial Intelligence-of-Things prevents uploading user-sensitive information to the cloud. It brings an urgent demand for deploying training tasks collaboratively in industrial scenarios to manage data locally. This article proposes an edge-based collaborative training system for the smart factory which harnesses the intelligence of edge devices by balancing the computational and communicational resources and improving system dependability. Two typical scenarios of parts recognition and defect inspection are evaluated as a case study with our system. The feasibility and dependability of the presented system are verified with a platform composed of eight high-performance (Nvidia Jetson Nano) and eight low-performance edge devices (Raspberry Pi 4B). The efficiency under tradeoff between computational resource and network condition constraints in a cluster is tested to simulate real-case performance in smart factory scenarios. Our platform reaches the peak performance of 1167 images/s training efficiency on ResNet32 under a 125 MB/s bandwidth. Experimental results demonstrate that the proposed design can collaboratively perform training tasks with optimized efficiency and provide dependable collaborations for system fault detection and cluster extension. Yi Jin 0007, Yulong Yan, Yuxiang Huan, Jiawei Xu 0002, Shancang Li, Prosanta Gope, Zhuo Zou, Lirong Zheng 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2021 | Self-aware distributed deep learning framework for heterogeneous IoT edge devices
Yi Jin 0007, Jiawei Cai, Jiawei Xu 0002, Yuxiang Huan, Yulong Yan, Yongliang Guo, Lirong Zheng 0001, Zhuo Zou |
Future Gener. Comput. Syst. | 3 |
| 2021 | IECA: An In-Execution Configuration CNN Accelerator With 30.55 GOPS/mm² Area EfficiencyabstractIt remains challenging for a Convolutional Neural Network (CNN) accelerator to maintain high hardware utilization and low processing latency with restricted on-chip memory. This paper presents an In-Execution Configuration Accelerator (IECA) that realizes an efficient control scheme, exploring architectural data reuse, unified in-execution controlling, and pipelined latency hiding to minimize configuration overhead out of the computation scope. The proposed IECA achieves row-wise convolution with tiny distributed buffers and reduces the size of total on-chip memory by removing 40% of redundant memory storage with shared delay chains. By exploiting a reconfigurable Sequence Mapping Table (SMT) and Finite State Machine (FSM) control, the chip realizes cycle-accurate Processing Element (PE) control, automatic loop tiling and latency hiding without extra time slots for pre-configuration. Evaluated on AlexNet and VGG-16, the IECA retains over 97.3% PE utilization and over 95.6% memory access time hiding on average. The chip is designed and fabricated in a UMC 55-nm process running at a frequency of 250 MHz and achieves an area efficiency of 30.55 GOPS/mm2and 0.244 GOPS/KGE (kilo-gate-equivalent), which makes an over$2.0\times $and$2.1\times $improvement, respectively, compared with that of previous related works. Implementation of the IEC control scheme uses only a 0.55% area of the 2.75 mm2core. Boming Huang, Yuxiang Huan, Haoming Chu, Jiawei Xu 0002, Lizheng Liu, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2020 | A Smart Dental Health-IoT Platform Based on Intelligent Hardware, Deep Learning, and Mobile TerminalabstractThe dental disease is a common disease for a human. Screening and visual diagnosis that are currently performed in clinics possibly cost a lot in various manners. Along with the progress of the Internet of Things (IoT) and artificial intelligence, the internet-based intelligent system have shown great potential in applying home-based healthcare. Therefore, a smart dental health-IoT system based on intelligent hardware, deep learning, and mobile terminal is proposed in this paper, aiming at exploring the feasibility of its application on in-home dental healthcare. Moreover, a smart dental device is designed and developed in this study to perform the image acquisition of teeth. Based on the data set of 12 600 clinical images collected by the proposed device from 10 private dental clinics, an automatic diagnosis model trained by MASK R-CNN is developed for the detection and classification of 7 different dental diseases including decayed tooth, dental plaque, uorosis, and periodontal disease, with the diagnosis accuracy of them reaching up to 90%, along with high sensitivity and high specificity. Following the one-month test in ten clinics, compared with that last month when the platform was not used, the mean diagnosis time reduces by 37.5% for each patient, helping explain the increase in the number of treated patients by 18.4%. Furthermore, application software (APPs) on mobile terminal for client side and for dentist side are implemented to provide service of pre-examination, consultation, appointment, and evaluation. Lizheng Liu, Jiawei Xu 0002, Yuxiang Huan, Zhuo Zou, Shih-Ching Yeh, Lirong Zheng 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | A 3D Tiled Low Power Accelerator for Convolutional Neural NetworkabstractIt remains a challenge to run Deep Learning in devices with stringent power budget in the Internet-of-Things. This paper presents a low-power accelerator for processing Convolutional Neural Networks on the embedded devices. The power reduction is realized by exploring data reuse in three different aspects, with regards to convolution, filter and input features. A systolic-like data flow is proposed and applied to rows of Processing Elements (PEs), which facilitate reusing the data during convolution. Reuse of input features and filters is achieved by arranging the PE array in a 3D tiled architecture, whose dimension is 3 × 14 × 4. Local storage within PEs is therefore reduced and only cost 17.75 kB, which is 20% of the state-of-the-art. With dedicated delay chains in each PE, this accelerator is reconfigurable to suit various parameter settings of convolutional layers. Evaluated in UMC 65 nm low leakage process, the accelerator can reach a peak performance of 84 GOPS and consume only 136 mW at 250 Mhz. Yuxiang Huan, Jiawei Xu 0002, Lirong Zheng 0001, Hannu Tenhunen, Zhuo Zou |
ISCAS | 2 |