Xin'an Wang

dblp:45/6404 · DBLP profile ↗
← Back
20ranked-venue papers
0as first author
8since 2021 · last 2024
0000-0002-9712-8531ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2024 An Energy-Efficient Configurable Coprocessor Based on 1-D CNN for ECG Anomaly Detection
abstract
Many healthcare devices have been widely used for electrocardiogram (ECG) monitoring. However, most of them have relatively low energy efficiency and lack flexibility. A novel ECG coprocessor is proposed in this paper, which can perform efficient ECG anomaly detection. In order to achieve high sensitivity and positive precision of R-peak detection, an algorithm based on Hilbert transform and adaptive threshold comparison is proposed. Also, a flexible one-dimensional convolutional neural network (1-D CNN) based classification engine is adopted, which can be configured with instructions to process various network models for different applications. Good energy efficiency is achieved by combining filter level parallelism and output channel parallelism within the processing element (PE) array with data reuse strategy. A 1-D CNN for arrhythmia detection is proposed to validate the hardware performance. The proposed ECG coprocessor is implemented using 55 nm CMOS technology, occupying an area of 1.39 mm2. At a clock frequency of 100MHz, the energy efficiency is 215.6 nJ/classification. The comparison results show that this design has advantages in energy overhead and detection performance.
Chen Zhang 0031, QianXi Cheng, Changchun Zhou 0001, Xin'an Wang
ISCAS5
2022 OMNET: Real-Time Stereo Matching with Unsupervised Occlusion Mask
abstract
Although CNN has powerful learning capability, it is still difficult for CNN to judge corresponding points in the occlusion region. The ghosting effect in the occlusion region during feature warping is the bottleneck of performance improvement for many stereo matching networks. In this paper, we propose an Occlusion-Aware Refinement Module (OARM), which can learn a rough occlusion map from multi-scale aggregated cost volumes without occlusion supervision to mask and filter pernicious occluded regions in a warped image. Cooperating with well-designed simple yet efficient 2D-based Intra/Cross-Level Aggregation Modules to effectively and efficiently aggregate information of different scales before the disparity refinement stage, OARM helps our proposed OM-Net achieve an error rate of 1.82% (D1-all) on KITTI 2015 dataset, which is even better than most 3D-based networks. Meanwhile, OMNet keeps real-time characteristic and could process a 1248×384 resolution image pair at 26 fps.
Shuiqiang Ye, Xin'an Wang, Yong Zhao 0010
ICIP3
2022 MLP-Stereo: Heterogeneous Feature Fusion in MLP for Stereo Matching
abstract
CNNs’ strong inductive biases of locality and weight sharing provide powerful representation ability and data sample utilization efficiency. However, the weight sharing might smooth out the discrepancy between similar pixels, resulting in the wrong matching between left and right camera-image pair in thin structures region and repetitive texture region. In this paper, we propose a novel Heterogeneous Feature Fusion in MLP (HFF-MLP) for Stereo matching. It employs MLP structure and relaxes the weights sharing in the local spatial region. To this end, pixels in thin structures region and repetitive texture region are dealt with independently using the exclusive weights. Based on HFF -MLP module, we design a real-time network, i.e., MLP-Stereo. Experimental results show that our proposed HFF-MLP achieves competitive results on KITTI 2015 test dataset with the running time of 46 milliseconds. Furthermore, it performs much better than other real-time network in thin structures region and repetitive texture region.
Shuiqiang Ye, Pengcheng Zeng, Xin'an Wang, Yong Zhao 0010
ICIP5
2022 Self-adaptive Multi-scale Aggregation Network for Stereo Matching
abstract
Nowadays stereo matching architectures based on convolutional neural network has achieved remarkable performance. However the existing methods still lack the capability to find the correspondence in ill-posed regions. In this paper we present Self-adaptive Multi-scale Aggregation Network (SMA-Net) for stereo matching. First of all, we construct the cost volume through multi-channels group-wise correlation, which divided the features into groups with different number of channels to enhance the ability of measuring the similarities of features from stereo images. Secondly the self-adaptive cost aggregation is used to regularize the two scale cost volumes from different aggregation branches with intermediate supervise. We conduct comprehensive experiments on SceneFlow, KITTI2012, and KITTI2015 datasets. The competitive results prove that the approach in this paper outperforms many other stereo matching algorithms especially in ill-posed regions.
Shuiqiang Ye, Jiaquan Zhang, Xin'an Wang, Qifei Dai, Zhengzhong Yu, Fuchi Li, Yong Zhao 0010
ICPR4
2022 A Computing-in-Memory SRAM Macro Based on Fully-Capacitive-Coupling With Hierarchical Capacity Attenuator for 4-b MAC Operation
abstract
In this work, we present a fully capacitive-coupling-based SRAM computing-in-memory (CIM) macro aimed at improving the energy efficiency and throughput of edge devices running multi-bit multiply-and-accumulate (MAC) operations. The proposed architecture is built around a customized 9T1C bit-cell in charge-domain computation in a 28nm technology. The proposed design supports 8192 $4{\mathrm{b}}\times 4{\mathrm{b}}$ MAC operations simultaneously. A 4-bit input is generated by DAC, while a 4-bit weight is achieved by a hierarchical capacity attenuator array without additional sharing switches, long sharing time, and complicated controlling signal. To minimize the expensive AD conversion, an input sparsity sensing scheme is proposed, allowing to skip redundant comparators. Access time is 4 ns with 0.9 V power supply at room temperature. The proposed design achieves energy efficiency of 666 TOPS/W and throughput of 4096 GOPS.
Kanglin Xiao, Xiaoxin Cui, Nanbing Pan, Xin'an Wang, Yuan Wang 0001
ISCAS5
2021 An SNN-Based and Neuromorphic-Hardware-Implementable Noise Filter with Self-adaptive Time Window for Event-Based Vision Sensor
abstract
Event-based dynamic vision sensors (DVS), inspired by biological vision systems, lead to new sensing and computing paradigms. The novel sensors output the sensed signal alone with many noise events asynchronously. Data-preprocessing for filtering these noises is significant before utilizing the data in applications such as classification, tracking and motion-data extraction. This paper describes a fully spike-based and neuromorphic-hardware-implementable neural network with a signal-oriented self-adaptive filtering time window for filtering the noise events robustly in the data captured by DVS. In particular, the simple leaky integrate-and-fire (LIF) neuron model is adopted as the basic elements of the network out of the purpose of hardware-friendly. Experiments based on both synthesized data and authentically-captured data are designed for quantitative comparison with traditional DVS noise filters to verify the outperformance of the proposed filter. The main contribution of this work is that the proposed spiking neural network (SNN) based filter achieves higher signal-noise-ratio (SNR) compared to traditional noise filters and performances more robust in the tolerance for changing signals.
Kanglin Xiao, Xiaoxin Cui, Kefei Liu 0002, Xiaole Cui, Xin'an Wang
IJCNN5
2021 Design and Implementation of a Temperature Self-Compensation Balanced Hybrid Ring Oscillator BHRO
abstract
The ring oscillator (RO) is applied to many modern circuits given its simplicity, low-area-cost, low-power. However, the temperature-drifting propagation delay of logic gates makes the RO a difficult solution for frequency reference design in deep submicron process. This work presents the design and implementation of a temperature self-compensation architecture based on a balanced hybrid ring oscillator (BHRO) for precise temperature compensation and clock-on-chip applications. The proposed BHRO is formed by both PTAT (proportional to the absolute temperature) and CTAT (complementary to absolute temperature) delay cells to implements RO-internal compensation. The features include temperature-self-compensation, process insensitive and low-area-cost. Four different test chips are fabricated in the 0.13μm CMOS process. The measurement result exhibits two performance-friendly BHRO architectures with a best temperature coefficient of 31 ppm/C over -55 to 80 , which is among the lowest to our best knowledge. The output compensated frequency is verified to be adjustable varying from 6.95 MHz to 26.5MHz, which are the highest as we have known.
Kanglin Xiao, Bo Wang 0016, Changpei Qiu, Xin'an Wang
ISCAS4
2021 An FPGA-Based Convolutional Neural Network Coprocessor
abstract
In this paper, an FPGA‐based convolutional neural network coprocessor is proposed. The coprocessor has a 1D convolutional computation unit PE in row stationary (RS) streaming mode and a 3D convolutional computation unit PE chain in pulsating array structure. The coprocessor can flexibly control the number of PE array openings according to the number of output channels of the convolutional layer. In this paper, we design a storage system with multilevel cache, and the global cache uses multiple broadcasts to distribute data to local caches and propose an image segmentation method that is compatible with the hardware architecture. The proposed coprocessor implements the convolutional and pooling layers of the VGG16 neural network model, in which the activation value, weight value, and bias value are quantized using 16‐bit fixed‐point quantization, with a peak computational performance of 316.0 GOP/s and an average computational performance of 62.54 GOP/s at a clock frequency of 200 MHz and a power consumption of about 9.25 W.
Changpei Qiu, Xin'an Wang, Tianxia Zhao, Bo Wang 0016
Wirel. Commun. Mob. Comput.2
2020 A Novel Conversion Method for Spiking Neural Network using Median Quantization
abstract
Artificial Neural Networks (ANNs) have achieved great success in the field of computer vision and language understanding. However, it is difficult to deploy these deep learning models on mobile devices because of its massive energy consumption and memory occupation. For another way, highly inspired from biological brain, spiking neural networks (SNNs), are often referred to as the 3-th generation of neural network for its potential superiority in cognitive learning and energy efficiency. Nevertheless, training a deep SNN remains a big challenge. In this paper, we propose a quantized training algorithm for ANNs to minimize spike approximation error, and provide two (temporally or spatially) rate-based conversion methods for SNNs, both of which can be easily mapped to specific neuromorphic platforms. Besides, this novel method can be generalized to various network architectures and adapted to dynamic quantization demand. Experimental results on MNIST and CIFAR-10 dataset demonstrate that the proposed deep spiking neural networks yield the state-of-the-art classification accuracy and need much less operations compared with their ANN counterparts. Our source code will be available upon request for the academic purpose.
Chenglong Zou, Xiaoxin Cui, Jiexian Ge, Hanghang Ma, Xin'an Wang
ISCAS5
2019 Research on feature extraction algorithm for plantar pressure image and gait analysis in stroke patients
Xin'an Wang, Zhuochen Fan, Sixu Zhang
J. Vis. Commun. Image Represent.2
2017 COSY: An Energy-Efficient Hardware Architecture for Deep Convolutional Neural Networks Based on Systolic Array
abstract
Deep convolutional neural networks (CNNs) show extraordinary abilities in artificial intelligence applications, but their large scale of computation usually limits their uses on resource-constrained devices. For CNN's acceleration, exploiting the data reuse of CNNs is an effective way to reduce bandwidth and energy consumption. Row-stationary (RS) dataflow of Eyeriss is one of the most energy-efficient state-of-the-art hardware architectures, but has redundant storage usage and data access, so the data reuse has not been fully exploited. It also requires complex control and is intrinsically unable to skip over zero-valued inputs in timing. In this paper, we present COSY (CNN on Systolic Array), an energy-efficient hardware architecture based on the systolic array for CNNs. COSY adopts the method of systolic array to achieve the storage sharing between processing elements (PEs) in RS dataflow at the RF level, which reduces low-level energy consumption and on-chip storage. Multiple COSY arrays sharing the same storage can execute multiple 2-D convolutions in parallel, further increasing the data reuse in the low-level storage and improving throughput. To compare the energy consumption of COSY and Eyeriss running actual CNN models, we build a process-based energy consumption evaluation system according to the hardware storage hierarchy. The result shows that COSY can achieve an over 15% reduction in energy consumption under the same constraints, improving the theoretical Energy-Delay Product (EDP) and Energy-Delay Squared Product (ED2P) by 1.33× on average. In addition, we prove that COSY has the intrinsic ability for zero-skipping, which can further increase the improvements to 2.25× and 3.83× respectively.
Miren Tian, Mohan Ji, Chenglong Zou, Xin'an Wang, Bo Wang 0016
ICPADS6
2017 Testing of 1TnR RRAM array with sneak path technique
Xiaole Cui, Xiaoxin Cui, Xin'an Wang, Jinfeng Kang
Sci. China Inf. Sci.4
2017 A Robust and Efficient Approach to License Plate Detection
abstract
This paper presents a robust and efficient method for license plate detection with the purpose of accurately localizing vehicle license plates from complex scenes in real time. A simple yet effective image downscaling method is first proposed to substantially accelerate license plate localization without sacrificing detection performance compared with that achieved using the original image. Furthermore, a novel line density filter approach is proposed to extract candidate regions, thereby significantly reducing the area to be analyzed for license plate localization. Moreover, a cascaded license plate classifier based on linear support vector machines using color saliency features is introduced to identify the true license plate from among the candidate regions. For performance evaluation, a data set consisting of 3977 images captured from diverse scenes under different conditions is also presented. Extensive experiments on the widely used Caltech license plate data set and our newly introduced data set demonstrate that the proposed approach substantially outperforms state-of-the-art methods in terms of both detection accuracy and run-time efficiency, increasing the detection ratio from 91.09% to 96.62% while decreasing the run time from 672 to 42 ms for processing an image with a resolution of 1082×728 . The executable code and our collected data set are publicly available.
Yule Yuan, Wenbin Zou, Yong Zhao 0010, Xin'an Wang, Xuefeng Hu, Nikos Komodakis
IEEE Trans. Image Process.4
2014 A novel low-noise high-linearity CMOS transmitter for mobile UHF RFID reader
Xin'an Wang, Jinpeng Shen, Bo Wang 0016, Ru Huang 0001
Sci. China Inf. Sci.2
2013 A super-regenerative pulsed UWB receiver combined with injection-locking
abstract
This paper presents a novel pulsed wideband super-regenerative receiver (SRR). It features the combination of the super-regenerative oscillator with the injection locking oscillator to realize a compact and low power SRR receiver. This receiver allows to calibrating the frequency by injection locking thus suppressing the use of the power consuming PLL. It also benefits the low power characteristics by adopting the pulsed UWB signal working at 3.5–4.5 GHz. The proposed receiver is designed in the CMOS 0.18um technology with 1.2-V supply. This circuit consumes less than 6.6mA (excluding the digital part) for a 60-Mbps pulse rate, with a BER of 10−4.
Tongning Hu, Bo Wang 0016, Jinghai Zhang, Jinpeng Shen, Xin'an Wang
ISCAS7
2013 Fully integrated passive UHF RFID transponder IC with a sensitivity of -12 dBm
abstract
This paper presents a fully integrated passive UHF RFID transponder, which includes an RF/analog front-end, a baseband processor and a 512-bit EEPROM memory. To improve power conversion efficiency, Schottky barrier diode based rectifier is adopted. The design of a novel voltage limiter is discussed in detail, which can provide a stable limiting voltage with low sensitivity to temperature variation and process dispersion. The whole transponder chip is implemented in 0.18 μm CMOS process with a die size of 800×800 μm2. Measurement results show that the total power consumption of the tag chip is only 7.4μW with a sensitivity of -12dBm.
Jinpeng Shen, Xin'an Wang, Bo Wang 0016, Shoucheng Li, Zhengkun Ruan, Xiangrong Zhang
ISCAS2
2013 A novel intelligent verification platform based on a structured analysis model
Xin'an Wang, Zhibin Lian, Yonggui Luo, Ziyi Hu
Sci. China Inf. Sci.2
2012 Theory and verification of operator design methodology
Ziyi Hu, Yong Zhao 0010, Xin'an Wang, Ru Huang 0001, Xing Zhang 0002
Sci. China Inf. Sci.3
2012 A novel compact low-power direct conversion receiver for mobile UHF RFID reader
Xin'an Wang, Jinpeng Shen, Bo Wang 0016, Ru Huang 0001
Sci. China Inf. Sci.2
2008 A Novel Tone Reservation Scheme with Fast Convergence for PAPR Reduction in OFDM Systems
abstract
OFDM is facing great opportunities and challenges in current broadband communication era. These opportunities and challenges derive from the native advantages and disadvantages of OFDM technology respectively. Too high PAPR is one of the main problems that prevent OFDM from being used more generally in broadband systems. Many approaches such as clipping and filtering, coding, SLM, PTS, and tone reservation have been studied to reduce the peak magnitude of OFDM symbols. In these approaches, tone reservation is considered as one of the most promising methods because of no additional distortion, no side information, and low implementation cost. In this paper, a novel tone reservation scheme is presented. Its essential idea is that a subcarrier selected from all reserved subcarriers for PAPR reduction should have a phase close to one of phi, pi/2+phi, pi+phi and -pi/2+phi, at the peak location in time domain, where phi is the phase of the peak sample. This results in no complex multiplication and division in the novel scheme. In addition, the floating positions of such selected subcarriers in frequency domain are helpful to the convergence of the algorithm. The simulation results show that the scheme can provide good performance and fast convergence.
Yuzhong Jiao, Xin'an Wang
CCNC3