EDBT 2026 Demo / reviewers in the wild / expert
Qi Wei 0001
dblp:43/2782-1
· DBLP profile ↗
35ranked-venue papers
0as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 18 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GCC-CIM: A Charge-Domain Compute-in-Memory Macro using Grouped-Row Capacitors and C-2C Ladder with Improved Multi-Bit MAC Linearity
Erxiang Ren, Daniel Zheng Fang, Qi Wei 0001, Fei Qiao |
ISCAS | 5 |
| 2025 | FEI: Fusion Processing of Sensing Energy and Information for Self-Sustainable Infrared Smart Vision SystemabstractIn the natural world, energy and information are deeply entwined, mutually constraining and complementing each other. To exploit this natural merit, this paper proposes a FEI strategy: Fusion processing of sensing Energy and Information for infrared smart vision system. The proposed Information-Power-Coupler (IPCp) takes the ability of simultaneous energy harvesting and low power inpixel computing, which utilizes in-situ coupled energy to process the containing information on the same focal plane. Furthermore, a self-adaptive Intelligent-Power- Controller (IPCtrl) capable of scheduling the harvested energy to complete low power neural network inference is introduced. The implementation of IPC2 system utilizes a software-hardware co-design strategy to exploit the layer-wise characteristic of the computation process and circuit topology, achieving energy-efficient self-sustainable fusion processing of sensing energy and information. Simulation results show that the IPCtrl could supply 594.68nW with the power conversion efficiency of 93.38%, when the harvested energy from the IPCp is 636.84nW. The performance validates the self-sustainability of the system with the self-powered image recognition of a complete network running at 4fps with an accuracy of 99.4%. Haijin Su, Maimaiti Nazhamaiti, Qi Wei 0001, Zheyu Liu, Wenjie Deng, Yongzhe Zhang, Fei Qiao |
ASP-DAC | 6 |
| 2025 | Dcha: Distributed-Centralized Heterogeneous Architecture Enables Efficient Multi-Task Processing for Smart SensingabstractThe rapid development of artificial intelligence (AI) has accelerated the progression of IoT technology into the smart era. Integrating AI processing capabilities into IoT devices to create smart sensing systems holds significant promise. In this work, we propose a distributed-centralized heterogeneous architecture that enables efficient multitask processing for smart sensing. This architecture improves the operational efficiency of sensing systems and enhances the deployment scalability through collaborative computing across end, edge, and center nodes. Specifically, we partition the network in traditional centralized sensing systems into several parts and perform algorithm-hardware co-design for each part on its respective deployment platform. We developed a sample design to validate the proposed architecture. By implementing a lightweight image encoder, we achieved an 88x reduction in encoder parameters and up to 9873x energy gain, facilitating deployment on resource-constrained devices. Experimental results demonstrate that the proposed architecture effectively reduces overall energy consumption by 0.0573x to 0.0889x, while maintaining robust multitask inference capabilities. Moreover, energy consumption reductions of 2.88x to 3.22x on edge nodes and 6311.56x to 10037.23x on end nodes were observed. Erxiang Ren, Cheng Qu, Zheyu Liu, Xinghua Yang, Qi Wei 0001, Fei Qiao |
DATE | 7 |
| 2025 | Denoise on Sensor: A Near-Sensor Compute-in-Memory Macro for Visual Perception Denoising via Concatnation-EliminatingabstractNoise is one of the most common and significant factors leading to image degradation. In recent years, due to the rapid development of neural networks, the performance of denoising algorithms has seen a substantial improvement. However, state-of-the-art denoising models often entail large-scale models and heavy computational requirements, making deployment challenging. Additionally, we have observed that the location of denoisers in the entire image processing pipeline has a significant impact on resource consumption and denoising effectiveness. Deploying the denoiser closer to the image acquisition stage will be more effective in separating noise from the image. In this paper, we propose a near-sensor compute-in-memory macro for visual perception denoising (Denoise on Sensor, DoS) along with its corresponding Edge Denoise U-net (EDU) architecture. DoS employs a mixed-signal circuit implementation for neural network inference, offering a notable advantage in terms of high speed and low power consumption compared to FPGA or GPU based approaches, making it feasible to deploy denoising tasks at the near-sensor edge. The simulation results show that the energy efficiency of DoS can reach 21.98 TOPS/W, and EDU deployed on DoS can achieve around 30dB PSNR and 0.83 SSIM on KODAK, BSD300 and SET14 datasets. Aolin You, Erxiang Ren, Daniel Zheng Fang, Cheng Qu, Qi Wei 0001, Fei Qiao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | Analog Foreground Calibration of High-Resolution SAR ADCabstractTechniques for the foreground calibration of split capacitor successive approximation register (SAR) are discussed. The used methods of digital calibration requires a digital post-processing that can be a limit for the resulting overhead. The presented techniques allow directly obtaining the digital output. Two methods address the mismatch among the unit elements in the most significant byte (MSB) segment and the mismatch at the MSB-LSB interface. A 16-bit SAR ADC with a sampling rate of 1MS/s employs both calibration techniques and verifies the methods. The simulation results show a Schreier merit of 166.3 dB, a signal-to-noise-and-distortion ratio (SNDR) of 87 dB, and an SFDR of -100 dB. A chip fabricated in a 180-nm Bipolar-CMOS-DMOS (BCD) process almost obtains the expected SNDR. The experimental result shows that the spurious-free dynamic range (SFDR) improves from 70.71 to 93.29 dB. Hua Fan 0001, Zhuorui Chen, Franco Maloberti, Qi Wei 0001, Quanyuan Feng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | A 17-mK Resolution ΣΔ-Based Temperature Sensor With an On-Chip Single-Point Calibrated Inaccuracy of ±1.5 °C (3σ)abstractThis work presents the design of a bipolar junction transistor (BJT)-based temperature sensor with a sigma-delta analog-to-digital converter ($\Sigma \Delta $ADC) fabricated in a 0.6$\mu $m BiCMOS process. The sensor utilizes a proportional to absolute temperature (PTAT) current source embedding an auto-zeroing amplifier to bias the BJT core while effectively reducing offset errors. Additionally, this work achieves the combination of the temperature-sensing BJT core and the$\Sigma \Delta $modulator, utilizing dynamic element matching (DEM) technique to accurately generate three different ratios (1:9, 9:1, 5:5) of BJT bias currents to cooperate with the modulator for modulation. The innovative modulation method in this work can reduce the errors caused by the operational amplifier offset voltage and BJT mismatch, completing the digitization of temperature without the need for a reference voltage source. Moreover, the sensor integrates a complete digital signal processor (DSP), allowing on-chip calibration by adjusting the operating cycle of the modulator and the initial value of the counter to obtain accurate and directly readable temperature data. It operates under a 3.3 V supply within a temperature range from −50 °C to 125 °C. Lastly, it achieves a 17 mK resolution and an inaccuracy of ±1.5 °C (3$\sigma $) after an on-chip single-point calibration. Hua Fan 0001, Hongrui Che, Ruoyu Yu, Hongquan Wang, Haizhu Wang, Fumei Liang, Antonio Aprile, Edoardo Bonizzoni, Haishi Wang, Qi Wei 0001, Panfeng Zhao, Quanyuan Feng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 12 |
| 2025 | AM-CIM: Approximate Memory Based Near Sensor Compute-in-Memory Architecture for Keyword SpottingabstractCompute-In-Memory (CIM) has emerged as a promising solution to address the von-Neumann bottleneck, making it a key technology for intelligent computing in edge IoT devices, particularly for real-time applications like keyword spotting (KWS). However, traditional CIM architectures face challenges such as high resource consumption, especially in data conversion, which can significantly impact chip area and energy efficiency. To address these challenges, this work proposes a computational CIM architecture utilizing multilevel analog memory, named AM-CIM, tailored for near-sensor (NS) computation of real-time KWS applications. Additionally, approximate memory technology is integrated into the AM-CIM architecture, employing data resilience scheduling for analog memory which contributes to significant reductions in hardware overhead. This integration facilitates a hardware-software co-design approach. To deploy KWS tasks in AM-CIM, a gated recurrent unit (GRU) network, referred to as MAC-GRU, is implemented. By employing Mel-energy as the input feature at the near-sensor end, the system achieves a 93.13% reduction in feature extraction power consumption. Evaluation results based on TSMC 180-nm technology demonstrate that the AM-CIM architecture achieves an accuracy of 88.51% for 10-keyword classification with a power consumption of$546~\mu W$, while reducing analog memory area by 43.32%. Xiaotao Jia, Guangcai Yuan, Jianyi Yu, Cong Shi 0003, Qi Wei 0001, Youguang Zhang, Weisheng Zhao 0001, Fei Qiao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | NS-Engine: Near-Sensor Neural Network Engine with SRAM-Based Compute-in-Memory MacroabstractSensing devices at edge nodes are usually resource-constrained, such as limited battery capacity and physical size. As a result, there is a substantial demand for enhancing energy efficiency in these devices, which can otherwise hinder the deployment of more complex neural networks. This work proposes an energy-efficient computing engine, which is equipped with appropriate computing power to deploy medium-size neural networks for smart sensing at near-sensor edge nodes. SRAM-based Compute-in-Memory (CIM) macro and end-to-end digital controller comprise the engine. We successfully prototyped this engine using an FPGA platform and conducted a demonstration of image classification on the CIFAR-10 dataset. Furthermore, we implement the above design on a TSMC 40nm process, with post-simulation results indicating an impressive throughput of 368.64GOPS and an energy efficiency of 25.7TOPS/W at 10MHz operating frequency, given the memory capacity is 576kb. Erxiang Ren, Xinghua Yang, Qi Wei 0001, Fei Qiao |
ISCAS | 5 |
| 2023 | A Three-Step Multi-Resolution Time-to-Digital ConverterabstractThis work proposes a three-step multi-resolution time-to-digital converter (TDC) architecture based on the vernier delay line (VDL). The proposed architecture uses a delay-locked loop (DLL) to control TDC with a smooth coarse-to-fine strategy. In addition, the fine TDC uses a combination of multiple resolutions to reduce the number of delay cells and flip-flops. This architecture helps to reduce the area and power consumption and maintains high resolution. We proposed architecture performs better trade-offs between power consumption, linearity, accuracy, and measurement range. The simulation results show that the 7-bit TDC based on VDL designed in 180 nm CMOS achieves 5 ps of time resolution, 0.76/-0.8 LSB DNL and 1.02/-1.39 LSB INL at 100 MHz clock frequency while consuming 3.1 mW, which corresponds to the figure of merit (FoM) of 0.242 pJ/Conv. Jiang Yan, Yu Wang 0002, Fei Qiao, Jiangwei Zhang, Qi Wei 0001, Qingpeng Zhu, Wenxiu Sun, Ge Shi 0001 |
ISCAS | 7 |
| 2023 | Breaking the energy-efficiency barriers for smart sensing applications with "Sensing with Computing" architectures
Xinghua Yang, Zheyu Liu, Kechao Tang, Xunzhao Yin, Cheng Zhuo, Qi Wei 0001, Fei Qiao |
Sci. China Inf. Sci. | 6 |
| 2022 | A 2.17μW@120fps Ultra-Low-Power Dual-Mode CMOS Image Sensor with Senputing ArchitectureabstractThis paper proposes an ultra-low-power CMOS Image Sensor (CIS) chip based on sensing-with-computing (Senputing) architecture to reduce the power bottleneck of vision system. This Senputing chip achieves BNN 1st-layer convolution in analog domain with ultra-low power consumption. It has two working modes, Normal-Sensor (NS) mode and Direct- Photocurrent-Computation (DPC) mode. The prototype measurement results under 65nm CMOS process on MNIST classification task shows that the power of feature map computation is 2.17μW with 120fps frame rates and 98.1% accuracy. The computation efficiency reaches to 11.49TOPs/W, which is 14.8× higher than state-of-art works. Han Xu 0006, Zheyu Liu, Qi Wei 0001, Fei Qiao |
ASP-DAC | 5 |
| 2022 | In-situ self-powered intelligent vision system with inference-adaptive energy scheduling for BNN-based always-on perceptionabstractThis paper proposes an in-situ self-powered BNN-based intelligent visual perception system that harvests light energy utilizing the indispensable image sensor itself. The harvested energy is allocated to the low-power BNN computation modules layer by layer, adopting a light-weighted duty-cycling-based energy scheduler. A software-hardware co-design method, which exploits the layer-wise error tolerance of BNN as well as the computing-error and energy consumption characteristics of the computation circuit, is proposed to determine the parameters of the energy scheduler, achieving high energy efficiency for self-powered BNN inference. Simulation results show that with the proposed inference-adaptive energy scheduling method, self-powered MNIST classification task can be performed at a frame rate of 4 fps if the harvesting power is 1μW, while guaranteeing at least 90% inference accuracy using binary LeNet-5 network. Maimaiti Nazhamaiti, Haijin Su, Han Xu 0006, Zheyu Liu, Fei Qiao, Qi Wei 0001, Zidong Du, Xinghua Yang |
DAC | 6 |
| 2022 | OCTOANTS: A Heterogeneous Lightweight Intelligent Multi-Robot Collaboration System with Resource-constrained IoT DevicesabstractAs the focus on highly intelligent robots continues, a problem that cannot be ignored has emerged: resource con-straints. Considering the game problem of resource limitation and the level of intelligence, we focus on lightweight intelligence. This work is a further refinement of our previous work, a heterogeneous lightweight intelligent multi-robot system. In-spired by the nature creatures “octopus” and “ants”. First, we propose a heterogeneous centralized-distributed architecture, which can make robots collaboration more flexible and non-redundant. Second, to reflect lightweight intelligence, we use the Raspberry Pi, a low computing and power consumption internet of things (IoT) device, as a processing platform and first propose a quantitative definition of the lightweight intelligent system. Then, combining the centralized-distributed architecture and the lightweight computing platform, we propose an adapted algorithm called OCTOANTS and apply it to the simultaneous localization and mapping (SLAM) field. The OCTOANTS architecture consists of one brain and eight tentacles, which can achieve complex things with proper collaboration between them. Finally, we use heterogeneous cameras and heterogeneous algorithms to form a lightweight intelligent collaborative system that can run in the real world. On the low-grade platform Raspberry Pi our heterogeneous tentacles frame rate can reach 41fps and 99.8fps respectively, power consumption is only 2W and 1.2W. At the same time, our heterogeneous system is on average 7.2% more accurate than the state-of-the-art homogeneous system and can be applied to a wider range of application scenarios, demonstrating the superiority and feasibility of our OCTOANTS. Ruiyang Quan, Siqin Qimuge, Peimin Xia, Xin Zan, Fangshi Wang, Changchuan Chen, Qi Wei 0001, Huichan Zhao, Fei Qiao |
IROS | 9 |
| 2022 | Harmonic disturbance observer-based sliding mode control of MEMS gyroscopes
Rui Zhang 0021, Bin Xu 0003, Qi Wei 0001, Pengchao Zhang, Ting Yang 0006 |
Sci. China Inf. Sci. | 3 |
| 2022 | Senputing: An Ultra-Low-Power Always-On Vision Perception Chip Featuring the Deep Fusion of Sensing and ComputingabstractAlways-on intelligent visual perception applications are widely deployed in edges in the AIoT era. In order to eliminate power costs of data conversion and transmission, this paper proposes Senputing, an ultra-low-power processing-in-sensor chip that completely fuses sensing and computing together for a BNN-based hierarchical processing system. This chip could operate in two modes. In computation mode, photocurrents are directly utilized for computing without being converted into voltages, and the computation results of 1-st BNN layer are directly sent out to subsequent BNN processors for an always-on coarse classification, eliminating conversion power and storage cost of raw images. Once an interested objected is detected, this chip switches to sensor mode and sends raw images to potential full-precision processors or cloud servers for fine-grained recognition or segmentation. A$32\times 32$prototype is fabricated with 180nm CMOS process. It accomplishes MNIST dataset classification task with the accuracy of 93.76% and the power consumption of 147nW at 156fps, achieving$13.1\times $energy efficiency compared with state-of-the-art work. Han Xu 0006, Ningchao Lin, Qi Wei 0001, Runsheng Wang, Cheng Zhuo, Xunzhao Yin, Fei Qiao, Huazhong Yang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | RaP-Net: A Region-wise and Point-wise Weighting Network to Extract Robust Features for Indoor LocalizationabstractFeature extraction plays an important role in visual localization. Unreliable features on dynamic objects or repetitive regions will interfere with feature matching and challenge indoor localization greatly. To address the problem, we propose a novel network, RaP-Net, to simultaneously predict region-wise invariability and point-wise reliability, and then extract features by considering both of them. We also introduce a new dataset, named OpenLORIS-Location, to train the proposed network. The dataset contains 1553 images from 93 indoor locations. Various appearance changes between images of the same location are included and can help the model to learn the invariability in typical indoor scenes. Experimental results show that the proposed RaP-Net trained with OpenLORIS-Location dataset achieves excellent performance in the feature matching task and significantly outperforms state-of-the-arts feature algorithms in indoor localization. The RaPNet code and dataset are available at https://github.com/ivipsourcecode/RaP-Net. Dongjiang Li, Jinyu Miao, Xuesong Shi, Qiwei Long, Tianyu Cai, Hongfei Yu, Wei Yang 0029, Haosong Yue, Qi Wei 0001, Fei Qiao |
IROS | 11 |
| 2021 | Project-Based Course in Electronic Engineering EducationabstractIn the teaching of electronic engineering, some practical projects need to be added in order to connect various courses together. To strengthen the connections between courses, the project-based teaching method proposed in this paper advocates linking the knowledge of different courses through the combination of theory and practice. With this as a guide, projects have been set up for students. One of them is about designing and making a managed ethernet switch. In the process of making and completing this project, the students' ability has been significantly improved, which fully proves the benefits of the teaching method. Hua Fan 0001, Bochuan Li, Haizhu Wang, Qi Wei 0001, Quanyuan Feng, Hadi Heidari |
ISCAS | 6 |
| 2021 | A 5.9μW Ultra-Low-Power Dual-Resolution CIS Chip of Sensing-with-Computing for Always-on Intelligent Visual DevicesabstractIn the intelligent IoT edge devices, power consumption is increasing due to the deployment of high-precision algorithms, which greatly limits the working time of the devices. The power of A/D conversion and data transmission has become the bottleneck of traditional visual system. In this paper, a new sensing-with-computing (Senputing) architecture is proposed to reduce this power bottleneck by combining imaging and BNN 1st-layer feature map computation. This Senputing architecture has two working modes, Normal-Sensor mode and Direct-Photocurrent-Computation mode with different resolutions (128×128 and 32×32). An ultra-low-power CMOS image sensor (CIS) chip with Senputing architecture is proposed to verify the feasibility. Our CIS chip is simulated with 180nm CMOS technology, the power of feature map computation is 5.9μW, and the frame rate is 208fps. The computation efficiency reaches to 8.23TOPs/W, which is 10.1 x higher than previous works. Han Xu 0006, Qi Wei 0001, Fei Qiao |
ISCAS | 4 |
| 2021 | A 442.1 nVpp, 13.07 ppm/°C Ultra-Low Noise Bandgap Reference Circuit in 180 nm BCD ProcessabstractThis paper proposed an ultra-low noise bandgap reference (BGR), which could achieve a sub-microvolt level of peak-to peak output noise in the low frequency region and have good load capacity. The noise performance is improved by using a negative feedback loop rather than traditional operational amplifier, which eliminates the noise from operational amplifier. Additionally, detailed analysis of noise is used to optimize the noise performance of the BGR. The driving ability is improved by using current mirrors to avoid the effect of load current on the behavior of BGR. Simulation results show operating at a 5 V supply voltage, over the range from -40 °C to 125 °C, output voltage of the BGR is 1.3345 V of mean value and temperature coefficient is 13.07 ppm/°C. The peak-to-peak output noise of the BGR is 442.1 nV in the band of 0.1 Hz to 10 Hz. The maximum output load current is 444.6 mA. MonteCarlo simulation illustrates process-invariance of this design. Junjun Zou, Qi Wei 0001, Bin Zhou 0006, Xiang Li 0035, Chunge Ju, Rong Zhang 0005, Zhiyong Chen 0004 |
ISCAS | 2 |
| 2021 | Reducing SRAM Reading Power With Column Data Segment and Weights Correlation Enhancement for CNN ProcessingabstractConvolutional neural network (CNN) has been widely deployed in various processors for intelligent visual signal processing. However, the large amount of activations and weights in CNN causes huge power consumption on SRAM access. Data-adaptive SRAM design is a widely studied method to reduce SRAM reading power based on the utilization of data patterns, while current designs only exploit data patterns in a coarse granularity, and have no advantages when faced with randomly distributed weight data. In this article, we propose a hardware–software co-design scheme to reduce SRAM reading power for CNN processing. First, we propose a reconfigurable data-adaptive SRAM architecture with column data segmentation (CDS-RSRAM) to utilize data patterns. Data in one column is partitioned into several segments, and finer-grained data patterns are exploited within each segment for further reading power reduction. Then, a novel training method—minimum segmented neighbor difference (miniSND)—is proposed for enhancing the correlation of weights. MiniSND improves the similarity of weights without classification accuracy degradation, thus weights could benefit from CDS-RSRAM and be read out with less power consumption. Simulation results demonstrate that the co-design scheme saves up to 66%(8b)/89%(2b) power consumption compared with 8T SRAM. Han Xu 0006, Ziru Li, Deliang Fan, Fei Qiao, Qi Wei 0001, Huazhong Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2021 | Serial-Parallel Estimation Model-Based Sliding Mode Control of MEMS GyroscopesabstractThis article proposes a serial-parallel estimation model (SPEM)-based sliding mode control (SMC) of MEMS gyroscope. For the system nonlinearity, the linear-in-parameterized dynamics are formulated and the updating law of the parameter vector is given. For the system uncertainty, the radial basis function (RBF) neural network (NN) is utilized. To improve the approximation accuracy of the compound nonlinearity, the updating laws of the parameter vector and RBF NN weight are constructed by the tracking error and the filtered modeling error derived from SPEM. Furthermore, the fast terminal (FT) SMC is employed to achieve finite-time convergence. The simulation results show that the proposed controller obtains higher tracking accuracy and faster convergence, while the compound nonlinearity approximation is with higher precision. Rui Zhang 0021, Bin Xu 0003, Qi Wei 0001, Ting Yang 0006, Wanliang Zhao, Pengchao Zhang |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | Utilizing Direct Photocurrent Computation and 2D Kernel Scheduling to Improve In-Sensor-Processing EfficiencyabstractDeploying intelligent visual algorithms in terminal devices for always-on sensing is an attractive trend in the IoT era. In-sensor-processing architecture is proposed to reduce power consumption on A/D conversion and data transmission, which performs pre-processing and only converting low-throughput features. However, current designs still require high energy consumption on photoelectric conversion and analog data movement. In this paper, two methods are proposed to improve the energy efficiency of in-sensor-processing architecture, including direct photocurrent computation and 2D kernel scheduling. Photocurrents are directly involved in computation to avoid data conversion; thus the indispensable imaging power is also utilized for computing. Since the location of the pixel data is fixed, data scheduling is conducted on digital weights to eliminate analog data storage and movement. We implement a prototype chip with an array of 32 × 32 units to calculate the first layer of binarized LeNet-5. The post-simulation shows that the proposed architecture reaches the energy efficiency of 11.49TOPs/W, about 14.8x higher than previous works. Han Xu 0006, Maimaiti Nazhamaiti, Yidong Liu, Fei Qiao, Qi Wei 0001, Huazhong Yang |
DAC | 5 |
| 2020 | DXSLAM: A Robust and Efficient Visual SLAM System with Deep FeaturesabstractA robust and efficient Simultaneous Localization and Mapping (SLAM) system is essential for robot autonomy. For visual SLAM algorithms, though the theoretical framework has been well established for most aspects, feature extraction and association is still empirically designed in most cases, and can be vulnerable in complex environments. This paper shows that feature extraction with deep convolutional neural networks (CNNs) can be seamlessly incorporated into a modern SLAM framework. The proposed SLAM system utilizes a state-of-the-art CNN to detect keypoints in each image frame, and to give not only keypoint descriptors, but also a global descriptor of the whole image. These local and global features are then used by different SLAM modules, resulting in much more robustness against environmental changes and viewpoint changes compared with using hand-crafted features. We also train a visual vocabulary of local features with a Bag of Words (BoW) method. Based on the local features, global features, and the vocabulary, a highly reliable loop closure detection method is built. Experimental results show that all the proposed modules significantly outperforms the baseline, and the full system achieves much lower trajectory errors and much higher correct rates on all evaluated data. Furthermore, by optimizing the CNN with Intel OpenVINO toolkit and utilizing the Fast BoW library, the system benefits greatly from the SIMD (single-instruction-multiple-data) techniques in modern CPUs. The full system can run in real-time without any GPU or other accelerators. The code is public at https://github.com/ivipsourcecode/dxslam. Dongjiang Li, Xuesong Shi, Qiwei Long, Shenghui Liu, Wei Yang 0029, Fangshi Wang, Qi Wei 0001, Fei Qiao |
IROS | 7 |
| 2020 | ASP-SIFT: Using Analog Signal Processing Architecture to Accelerate Keypoint Detection of SIFT AlgorithmabstractThe scale-invariant feature transform (SIFT) algorithm is still one of the most reliable image feature extraction methods. Despite its excellent robustness on various image transformations, SIFT's intensive computational burden has been severely preventing it from being used in real-time and energy-efficient embedded machine vision systems. To reduce processing time and energy cost while executing SIFT, an analog signal processing architecture, analog signal processing (ASP)SIFT, is proposed in this article. In ASP-SIFT, the Gaussian pyramid construction, difference-of-Gaussian (DoG) pyramid construction and keypoint locating, which are the primary steps of the keypoint detection part of the SIFT algorithm, are done directly with analog circuit networks. Thus, by completing keypoint detection in the analog domain, the total processing time is approximately equal to the settling time of the circuit network. Besides, by adopting a current-mode circuit network operating in the subthreshold region, the power dissipation would be very low. Simulation results show that the total processing speed for a typical video graphics array (VGA)-format (640 × 480) image is up to 2.3 kframes per second, which is at least 3.26× faster than the state-of-the-art digital hardware accelerators, while the system power is 94.5 mW and the energy consumption is only 40 μJ per frame. Zichen Fan, Zheyu Liu, Zheng Qu 0002, Fei Qiao, Qi Wei 0001, Shuzheng Xu, Huazhong Yang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2019 | Concrete: A Per-layer Configurable Framework for Evaluating DNN with Approximate OperatorsabstractApproximate computing has drawn considerable attention to both academia and industry in the area of DNN hardware. Despite substantial efforts to design approximate circuits and building blocks, the resilience of DNN layers and structures remains an untapped field to explore. This paper presents an efficient framework to evaluate DNN resilience with fine-grained approximate operations, such as multipliers, adders and low-bit operators. The framework can execute large-scale approximate DNNs with relatively less time overhead. Massive experiments are conducted with the proposed framework to reveal the relationship between network structures and error tolerance. Additionally, a case study of fine-tuning the approximate DNN is presented. Zheyu Liu, Guihong Li, Fei Qiao, Qi Wei 0001, Ping Jin, Huazhong Yang |
ICASSP | 4 |
| 2019 | INA: Incremental Network Approximation Algorithm for Limited Precision Deep Neural NetworksabstractApproximate computing is a promising paradigm to deal with large computing workloads in fault-tolerant applications, providing opportunities to improve hardware efficiency of Deep Neural Networks (DNNs). However, it is still difficult to apply highly approximate arithmetics (e.g., multipliers) to DNNs due to the effect of error accumulation and the convergence problem in re-training phase. To tackle this limitation, we propose a hardware-software co-design algorithm, namely Incremental Network Approximation (INA). By addressing the convergence problem, INA promotes fault tolerance of DNNs, and yields more tradeoffs between accuracy and implementation cost. Experiments show that the approximate inference models re-trained by INA could achieve up to 80% hardware reduction in various hardware design level, while the classification accuracy degradation is less than 2%. Moreover, the experiments also exhibit the generality of INA algorithm for applying to various approximate multiplier design. Zheyu Liu, Kaige Jia, Weiqiang Liu 0001, Qi Wei 0001, Fei Qiao, Huazhong Yang |
ICCAD | 4 |
| 2018 | Calibrating process variation at system level with in-situ low-precision transfer learning for analog neural network processorsabstractProcess Variation (PV) may cause accuracy loss of the analog neural network (ANN) processors, and make it hard to be scaled down, as well as feasibility degrading. This paper first analyses the impact of PV on the performance of ANN chips. Then proposes an in-situ transfer learning method at system level to reduce PV's influence with low-precision back-propagation. Simulation results show the proposed method could increase 50% tolerance of operating point drift and 70% ∼ 100% tolerance of mismatch with less than 1% accuracy loss of benchmarks. It also reduces 66.7% memories and has about 50× energy-efficiency improvement of multiplication in the learning stage, compared with the conventional full-precision (32bit float) training system. Kaige Jia, Zheyu Liu, Qi Wei 0001, Fei Qiao, Yi Yang 0039, Hua Fan 0001, Huazhong Yang |
DAC | 3 |
| 2018 | MINTIN: Maxout-Based and Input-Normalized Transformation Invariant Neural NetworkabstractConvolutional Neural Network (CNN) is a powerful model for image classification, but it is insufficient to deal with the spatial variance of the input. This paper presents a Maxout-based and input-normalized transformation invariant neural network (MINTIN), which aims at addressing the nuisance variation of images and accumulating transformation invariance. We introduce an innovative module, the Normalization, and combine it with the Maxout operator. While the former focuses on each image itself, the latter pays attention to augmented versions of input, resulting in fully-utilized information. This combination, which can be inserted into existing CNN architectures, enables the network to learn invariance to rotation and scaling. While the authors of TI-POOLING acclaimed that they reached state-of-the-art results, ours reach a maximum decrease of 0.71%, 0.23% and 0.51% in error rate on MNIST-rot-12k, half-rotated MNIST and scaling MNIST, respectively. The size of the network is also significantly reduced, leading to high computational efficiency. Jingyang Zhang, Kaige Jia, Pengshuai Yang, Fei Qiao, Qi Wei 0001, Huazhong Yang |
ICIP | 5 |
| 2018 | DS-SLAM: A Semantic Visual SLAM towards Dynamic EnvironmentsabstractSimultaneous Localization and Mapping (SLAM) is considered to be a fundamental capability for intelligent mobile robots. Over the past decades, many impressed SLAM systems have been developed and achieved good performance under certain circumstances. However, some problems are still not well solved, for example, how to tackle the moving objects in the dynamic environments, how to make the robots truly understand the surroundings and accomplish advanced tasks. In this paper, a robust semantic visual SLAM towards dynamic environments named DS-SLAM is proposed. Five threads run in parallel in DS-SLAM: tracking, semantic segmentation, local mapping, loop closing and dense semantic map creation. DS-SLAM combines semantic segmentation network with moving consistency check method to reduce the impact of dynamic objects, and thus the localization accuracy is highly improved in dynamic environments. Meanwhile, a dense semantic octo-tree map is produced, which could be employed for high-level tasks. We conduct experiments both on TUM RGB-D dataset and in real-world environment. The results demonstrate the absolute trajectory accuracy in DS-SLAM can be improved one order of magnitude compared with ORB-SLAM2. It is one of the state-of-the-art SLAM systems in high-dynamic environments. Chao Yu 0005, Zuxin Liu, Fugui Xie, Yi Yang 0039, Qi Wei 0001, Fei Qiao |
IROS | 6 |
| 2017 | From "MISSION: IMPOSSIBLE" to mission possible: Fully flexible intelligent contact lens for image classification with analog-to-information processingabstractA prototype of fully flexible intelligent contact lens, which are shown in the impressive action movie series of “MISSION: IMPOSSIBLE”, has become the Possible Mission in this work. Hereon, the system adopts analog-to-information processing method to build a specific Multi-Layer Perceptron network for image classification tasks with flexible devices and circuits, where the information is extracted from raw data of the sensing analog signal directly. Simulated with HSPICE of Level-62 TFT device model, for standard test image data set of MNIST, the classification accuracy of the presented flexible neural network circuit is up to 92.99%; meanwhile, the classification speed is as fast as 10k fps, and the energy consumption is low to only 15.16μJ. Additionally, for the imperfections of flexible devices of larger devices mismatch and process variations, the fault-tolerance of the system has been evaluated as well, which demonstrates the feasibility of the presented methods and lowers the barrier to integrated all kinds of FLEXIBLE Devices into a FULLY FLEXIBLE Systems with sensors, processing parts and even energy harvesting parts, etc., in the future wearable smart terminals. Qin Li 0016, Zheyu Liu, Fei Qiao, Xing Wu 0005, Chaolun Wang, Qi Wei 0001, Huazhong Yang |
ISCAS | 6 |
| 2016 | A precision-improved processing architecture of physical computing for energy-efficient SIFT feature extractionabstractA precision-improved processing architecture of physical computing for energy-efficient SIFT feature extraction algorithm has been proposed in this paper. With the novel physical computing technology of active resistor network (PC: ARN), the SIFT algorithm could be processed in analog signal domain without synchronizing clock signals, which means the complex algorithm could be completed within the setup time of the circuit. Especially for the multi-scale Gaussian convolution of SIFT algorithm, an architecture of two-layer 1-dimension PC: ARN has been adopted to compute the horizontal and vertical 1-dimension gaussian filter, in which way higher accuracy can be obtained when compared with the results processed by a 2D active circuit network. A circuit-level simulation with 65nm CMOS technology has been carried out, which shows the energy consumption of gaussian pyramid multi-scale-filtering hardware architecture is about 25.3pJ, where the size of input frame is assigned as 256×256 pixels. Additionally, the average matching ratio in different image pairs is around 80%. Moreover, integrated into the dominating CMOS image sensor with column-parallel readout technology of analog-to-digital convertor, about 20× speedup can be achieved comparing with previous implementations with FPGA, GPU, etc. Fei Qiao, Xinghua Yang, Qi Wei 0001, Huazhong Yang |
ICASSP | 4 |
| 2015 | Design methodology for approximate accumulator based on statistical error modelabstractApproximate computing technology has aroused growing interest in circuit and system design for its well-performed trade-off between output quality and performance. Numerous basic circuits and system design methodologies for approximate computing have been proposed. Considering that the existing methodologies for the evaluation of tradeoff between output quality and performance is time-consuming, this paper presents a fast design methodology for approximate accumulator based on statistical error model, in which the inexact multistage speculative adder is adopted and modeled for its advantage of compact error pattern. To validate the proposed methodology, Support Vector Machine(SVM) algorithm is analyzed and mapped to a hardware system composed of inexact and accurate computing circuits. Results show that our time for searching the optimal mapping circuits has been saved by 22.08% than functional-based simulation where the final approximate system design achieves 1.57× speedups with 8.56% accuracy degradation. Xinghua Yang, Fei Qiao, Qi Wei 0001, Huazhong Yang |
ASP-DAC | 4 |
| 2015 | Physical computing circuit with no clock to establish Gaussian pyramid of SIFT algorithmabstractPhysical computing scheme of active resistor network is proposed in this paper to set up a multi-scale Gaussian filter, which is also called Gaussian Pyramid in image signal processing. The analog output signal of each image photodiode could be directly processed by the circuit topology of an active resistor network, which has been pre-designed to meet the requirement of the complicated Gaussian Pyramid in SIFT algorithm. Since it is operated with no clock, the physical computing scheme, with only setup time of the whole processing circuits, is much faster than various current digital realizations. The circuit-level simulation with 65nm CMOS technology has been carried out, which shows the energy consumption of the Gaussian Pyramid processing circuit is around 74.79pJ to filter a frame of 256 × 256 pixel image, and the circuit setting time of the processing is about 138.03ps, considering the parasitic capacitances of each nodes of the circuits. Furthermore, the presented circuit is integrated into a smart CMOS image sensor architecture. With the same processing procedure, the new method of active resistor network could achieve 1.58X speedup when compared with its counterpart of FPGA implementation, in which the column-parallel readout technology of analog-to-digital convertor is used. Fei Qiao, Qi Wei 0001, Huazhong Yang |
ISCAS | 3 |
| 2015 | A 14-bit 1.0-GS/s dynamic element matching DAC with >80 dB SFDR up to the NyquistabstractA 14-bit 1.0-GS/s current-steering digital-to-analog converter (DAC) was designed in a 65-nm CMOS process. For such current-steering DACs with a high sampling rate, the code-dependent load variations and switching glitches are a main bottleneck which limits the spurious-free dynamic range (SFDR). Dynamic element matching (DEM) has been an effective solution to randomize these glitches for a higher SFDR and also to reduce the matching requirement of the current cells for an area-efficient design which also improves the SFDR with reduced parasitic capacitance. An effective method named TRI-DEMRZ is proposed in this paper, consisting of time-relaxed interleaving, DEM and return-to-zero encoding. We also apply TRI-DEMRZ in synergy with complementary switched current sources (CSCS) to design the DAC for the purpose of a small die size and enhanced SFDR performance. Post-layout simulations show >80 dB SFDR up to the Nyquist. This DAC has a mixed 1.2 V / 2.5 V power supply and an active area of 0.48 mm2. Xueqing Li 0002, Qi Wei 0001, Huazhong Yang |
ISCAS | 3 |
| 2014 | Design of multi-stage latency adders using detection and sequence-dependence between successive calculationsabstractMulti-stage latency adders based on different prediction schemes have been proved promising to enhance the circuit performance with negligible overhead. This paper presents a novel predictor exploiting both the detection and the sequence-dependence between the successive calculations. The detection of carry-kill pattern of the input data can lower the probability of the operation with multiple clock cycles and the sequence-dependence between the successive calculations is adapted to eliminate redundant cycles. The improved predictors have been inserted into Ripple Carry Adder (RCA) and a multistage latency structure has been setup. Compared with the previous predictors, the proposed one could have the same function with less prediction bits, which results in more energy-efficiency. Simulation results show that 2.41X-3.05X speedups can be achieved than the non-prediction counterpart. Furthermore, a design flow and a method for error control are proposed when applying the adder to approximate computation so that more performance improvement could be obtained after trading off certain precision. Xinghua Yang, Fei Qiao, Qi Wei 0001, Huazhong Yang |
ISCAS | 4 |