VLDB 2026 Research / reviewers in the wild / expert
Fei Qiao
dblp:34/4515
· DBLP profile ↗
96ranked-venue papers
3as first author
59since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 48 · 29 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 16 · 9 since 2021Human-computer interaction and ubiquitous computing · 14 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7Databases, data management, data science and information retrieval · 5 · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Unified Low-Latency ML-KEM Accelerator with Deterministic Packetization and Conflict-Free Memory MappingabstractNational Institute of Standards and Technology (NIST) Federal Information Processing Standards (FIPS) 203 standardizes Module-Lattice-Based Key-Encapsulation Mechanism (ML-KEM), driving the demand for low-latency accelerators that handle heterogeneous pipeline rates and complex on-chip memory provisioning. We present a speed-prioritized, unified ML-KEM-512 accelerator with a 16-lane packet interface across Key Generation (KeyGen), Encapsulation (Encaps), and Decapsulation (Decaps). To bridge irregular data streams, an output shaper aggregates accepted coefficients from rejection sampling into deterministic 16-lane blocks, shortening control critical paths. To sustain parallel Number Theoretic Transform (NTT) and Inverse Number Theoretic Transform (INTT) utilization, memory is partitioned by access semantics: sequential boundary traffic uses hierarchical buffers, while the strided transform domain uses a 16-bank store via distributed Look-Up Table Random Access Memory (LUTRAM). XOR bank mapping and stage-aware scheduling eliminate access conflicts inherent in standard mappings, while a Latest Value Table (LVT)-based scheme provides logical Two-Write Two-Read (2W2R) semantics. Implemented on a Xilinx Artix-7 FPGA, the design operates at 225 MHz, completing KeyGen, Encaps, and Decaps in 1131, 1228, and 1641 cycles. Achieving an 17.78 µs lifecycle latency, it demonstrates superior execution efficiency for the complete ML-KEM-512 flow. Xiangrui Jia, Qingzeng Song, Yongjiang Xue, Weigang Kong, Fei Qiao |
ACM Great Lakes Symposium on VLSI | 5 |
| 2026 | An Agile Design Framework for Resource-Efficient and Parameterizable Edge ISPsabstractFacing the real-time and resource constraints of edge imaging systems, this paper presents a highly parameterizable ISP hardware pipeline in SpinalHDL. To overcome the poor reusability of manually maintained Verilog designs, we adopt a configuration-driven generation paradigm to construct an end-to-end RAW-to-YUV streaming architecture. Under a unified ISPConfig, the pipeline supports flexible module chaining, compile-time structural specialization, and timing-consistent integration. At the microarchitectural level, we introduce a generic Window Generator that combines parameterized line buffers with a deterministic Skid Buffer-based flow-control shell to preserve pixel-level synchronization under backpressure. Experiments on the AMD Xilinx Kria KV260 show that the generated pipeline sustains real-time 1080p60 processing and reduces LUT and DSP utilization by 19.0% and 28.8%, respectively, relative to a functionally matched hand-coded Verilog baseline. The resource gains are mainly attributed to the shared Window Generator, unified inter-stage interfaces, and generator-time elimination of duplicated glue logic. The generated design also agrees well with software references, achieving 39.71 dB PSNR, 0.9636 SSIM, and 1.255 MAE, while representative requirement changes can be completed within a 20-minute RTL-to-simulation regression loop on average across representative modification tasks. Xitong Jiang, Qingzeng Song, Yongjiang Xue, Weigang Kong, Fei Qiao |
ACM Great Lakes Symposium on VLSI | 5 |
| 2026 | FPGA-based Adaptive Texture-Aware Stereo Matching Accelerator for Co-axial RGB-NIR SensorsabstractEdge robots require reliable all-day depth perception, but deploying RGB-NIR heterogeneous sensor fusion is severely constrained by strict size, weight, and power limits. Furthermore, traditional cross-modal registration and dynamic vertical residual compensation rely on external memory buffering, which stalls the continuous hardware dataflow. This paper presents a hardware-algorithm co-designed, DDR-less stereo matching accelerator tailored for co-axial RGB-NIR sensors to enable robust all-day edge vision. By employing a fully pipelined datapath with on-chip line buffers, the design bypasses all external memory dependencies. A cost-domain adaptive texture fusion (CATF) module reuses the Census transform register array to extract texture confidence, adaptively modulating semi-global matching penalty parameters and cross-modal fusion weights with near-zero overhead. To ensure robustness against mechanical vibrations and thermal drifts, an epipolar redundancy mechanism dynamically searches adjacent scanlines, tolerating up to 1-pixel vertical residuals. Implemented on a Zynq-7020 SoC, the system processes 720×540 video streams at 128 FPS. Consuming 35,868 LUTs and 2.43 Mb BRAM, it reduces the mean absolute error by up to 19.9% in all-day scenarios compared to OpenCV SGBM baselines, achieving an effective balance between robust multimodal perception and hardware efficiency. Chenghe Zhang, Qingzeng Song, Yongjiang Xue, Weigang Kong, Fei Qiao |
ACM Great Lakes Symposium on VLSI | 5 |
| 2026 | FPGA-Based High-Parallelism ORB Feature Extraction Accelerator for Visual SLAM
Yongjiang Xue, Fei Qiao, Qingzeng Song |
ISCAS | 4 |
| 2026 | Live Demonstration: An Efficient and Compact Visuo-Tactile Perception System
Cheng Qu, Erxiang Ren, Guangyuan Xu, Fei Qiao |
ISCAS | 7 |
| 2026 | VIR-EH: An Ultra-Low-Power Wide-Spectrum Visible-Infrared Vision Processing Chip with On-Chip Energy Harvesting
Zhilin Tan, Hengrui Guo, Weihao Shan, Jinshui Miao, Fei Qiao |
ISCAS | 7 |
| 2026 | GCC-CIM: A Charge-Domain Compute-in-Memory Macro using Grouped-Row Capacitors and C-2C Ladder with Improved Multi-Bit MAC Linearity
Erxiang Ren, Daniel Zheng Fang, Qi Wei 0001, Fei Qiao |
ISCAS | 6 |
| 2026 | Agent technologies in smart manufacturing: A comprehensive review of evolution, architectures, and edge-cloud deployment
Xifan Yao, Fei Qiao |
Adv. Eng. Informatics | 3 |
| 2026 | Thermal optimization of analog ICs via transistor-array placement
Longtao Jia, Qingduan Meng, Jun Wang 0064, Fei Qiao, Bo Liu 0031 |
Integr. | 6 |
| 2026 | Dynamic Rescheduling for Aircraft Pulsed Assembly Lines: A Distributed Multi-Agent Approach
Juan Liu 0011, Chen Ding 0008, Yumin Ma, Fei Qiao |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | Game Theory-Based Production-Maintenance Collaborative Scheduling Using Imitation-Enhanced Alternating-Training Reinforcement LearningabstractThe collaborative organization of production scheduling and machine maintenance is crucial for achieving effective manufacturing. However, since these two activities are typically managed by different self-interested managers with distinct optimization needs in practice, their collaboration faces the challenge of simultaneous consideration and balanced optimization of both parties’ interests. To this end, this study investigates a novel game theory-based production-maintenance collaborative scheduling problem, wherein the typical objectives individually concerned by two activities are explicitly considered and the decision interactions between activities are modeled as a non-cooperative stochastic game. Nash equilibrium within the game model provides the ideal solution for balancing the interests of both parties. To solve this problem, an imitation-enhanced alternating-training reinforcement learning method is presented. In this method, an imitation learning-based initialization mechanism is designed to accelerate the training of two game agents, and an alternating training mechanism are developed to facilitate two agents efficiently learning the optimal responses to each other’s behaviors and ultimately achieving the decisions that converge to Nash equilibrium. The superiority of the proposed method is verified through comprehensive experiments. Jiaxuan Shi, Fei Qiao, Juan Liu 0011, Yumin Ma |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | TLD-SCA: A Transformer-LSTM Detection Model against Side-Channel Attack in Blockchain Payment ChannelabstractSide-channel attack, which exploits time information leakage during the digital signature generation process, poses a severe threat to the confidentiality and integrity of transactions in blockchain payment channels. Currently, Transformer-based detection methods are effective at capturing long-term dependencies, while Long Short-Term Memory (LSTM) networks excel at modeling short-term dynamic time-series features. However, existing approaches struggle to uniformly model both long-term and short-term time-series features, limiting their performance in anomaly detection for complex transaction sequences. In this article, we propose TLD-SCA, a novel side-channel attack detection model that innovatively integrates Transformer’s capability for global dependency modeling with LSTM’s advantage in capturing local time-series dynamics. This enables long-term and short-term time-series analysis of transaction timing data. Experimental results demonstrate that TLD-SCA significantly outperforms existing methods in terms of accuracy (99.5%), precision (99.2%), and recall (98.3%), thereby providing a higher level of security assurance for blockchain payment channels. Tao Li 0043, Fei Qiao, Kui Lu |
ACM Trans. Web | 2 |
| 2025 | ELTC: An End-to-End Large Language Model-Based Tensor Compilation Optimization Framework
WenBo Ma, Qingzeng Song, Fei Qiao, Yongjiang Xue |
APLAS | 3 |
| 2025 | FEI: Fusion Processing of Sensing Energy and Information for Self-Sustainable Infrared Smart Vision SystemabstractIn the natural world, energy and information are deeply entwined, mutually constraining and complementing each other. To exploit this natural merit, this paper proposes a FEI strategy: Fusion processing of sensing Energy and Information for infrared smart vision system. The proposed Information-Power-Coupler (IPCp) takes the ability of simultaneous energy harvesting and low power inpixel computing, which utilizes in-situ coupled energy to process the containing information on the same focal plane. Furthermore, a self-adaptive Intelligent-Power- Controller (IPCtrl) capable of scheduling the harvested energy to complete low power neural network inference is introduced. The implementation of IPC2 system utilizes a software-hardware co-design strategy to exploit the layer-wise characteristic of the computation process and circuit topology, achieving energy-efficient self-sustainable fusion processing of sensing energy and information. Simulation results show that the IPCtrl could supply 594.68nW with the power conversion efficiency of 93.38%, when the harvested energy from the IPCp is 636.84nW. The performance validates the self-sustainability of the system with the self-powered image recognition of a complete network running at 4fps with an accuracy of 99.4%. Haijin Su, Maimaiti Nazhamaiti, Qi Wei 0001, Zheyu Liu, Wenjie Deng, Yongzhe Zhang, Fei Qiao |
ASP-DAC | 10 |
| 2025 | Dcha: Distributed-Centralized Heterogeneous Architecture Enables Efficient Multi-Task Processing for Smart SensingabstractThe rapid development of artificial intelligence (AI) has accelerated the progression of IoT technology into the smart era. Integrating AI processing capabilities into IoT devices to create smart sensing systems holds significant promise. In this work, we propose a distributed-centralized heterogeneous architecture that enables efficient multitask processing for smart sensing. This architecture improves the operational efficiency of sensing systems and enhances the deployment scalability through collaborative computing across end, edge, and center nodes. Specifically, we partition the network in traditional centralized sensing systems into several parts and perform algorithm-hardware co-design for each part on its respective deployment platform. We developed a sample design to validate the proposed architecture. By implementing a lightweight image encoder, we achieved an 88x reduction in encoder parameters and up to 9873x energy gain, facilitating deployment on resource-constrained devices. Experimental results demonstrate that the proposed architecture effectively reduces overall energy consumption by 0.0573x to 0.0889x, while maintaining robust multitask inference capabilities. Moreover, energy consumption reductions of 2.88x to 3.22x on edge nodes and 6311.56x to 10037.23x on end nodes were observed. Erxiang Ren, Cheng Qu, Zheyu Liu, Xinghua Yang, Qi Wei 0001, Fei Qiao |
DATE | 8 |
| 2025 | Multi-task Low-Level Vision Network for FPGA Deployment: SL-SYENet Joint Optimization Framework
Shenao Li, Yongjiang Xue, Fei Qiao, Qingzeng Song |
ICA3PP (7) | 3 |
| 2025 | FPGA-Based Hybrid GCN-Transformer Accelerator for 3D Pose Estimation
Yongjiang Xue, Fei Qiao, Qingzeng Song |
ICA3PP (6) | 3 |
| 2025 | DB-MFNet: A Dual-Branch Cross-Modal Fusion Network for High-Resolution Remote Sensing Semantic Segmentation
Huiren Hao, Yongjiang Xue, Fei Qiao, Qingzeng Song |
ICIC (5) | 3 |
| 2025 | RCMUNet: An End-to-End Hybrid U-Net Architecture of CNN-Mamba for Low-Light RAW Image Enhancement
Yongjiang Xue, Fei Qiao, Qingzeng Song |
ICIC (14) | 3 |
| 2025 | Mapping at First Sense: A Lightweight Neural Network-Based Indoor Structures Prediction Method for Robot Autonomous ExplorationabstractAutonomous exploration in unknown environments is a critical challenge in robotics, particularly for applications such as indoor navigation, search and rescue, and service robotics. Traditional exploration strategies, such as frontier-based methods, often struggle to efficiently utilize prior knowledge of structural regularities in indoor spaces. To address this limitation, we propose Mapping at First Sense, a lightweight neural network-based approach that predicts unobserved areas in local maps, thereby enhancing exploration efficiency. The core of our method, SenseMapNet, integrates convolutional and transformer-based architectures to infer occluded regions while maintaining computational efficiency for real-time deployment on resource-constrained robots. Additionally, we introduce SenseMapDataset, a curated dataset constructed from KTH and HouseExpo environments, which facilitates training and evaluation of neural models for indoor exploration. Experimental results demonstrate that SenseMapNet achieves an SSIM (structural similarity) of 0.78, LPIPS (perceptual quality) of 0.68, and an FID (feature distribution alignment) of 239.79, outperforming conventional methods in map reconstruction quality. Compared to traditional frontier-based exploration, our method reduces exploration time by 46.5% (from 2335.56s to 1248.68s) while maintaining a high coverage rate (88%) and achieving a reconstruction accuracy of 88%. The proposed method represents a promising step toward efficient, learning-driven robotic exploration in structured environments. Haojia Gao, Haohua Que, Kunrong Li, Weihao Shan, Mingkai Liu, Lei Mu, Xinghua Yang, Fei Qiao |
IJCNN | 10 |
| 2025 | FDSI-RTDETR : A Lightweight Unmanned Aerial Vehicle (UAV) Aerial Image Small Object Detection NetworkabstractUnmanned aerial vehicle (UAV) aerial image detection tasks frequently encounter challenges, including small object detection, significant occlusion, dense target presence, and inconsistent illumination conditions. Traditional object detection algorithms sometimes produce erroneous detections. This study proposes a lightweight enhanced method based on the design concept of RT-DETR. Initially, a lightweight FRESE-Block module is introduced to enhance the feature extraction capabilities of the backbone network. FasterNet is improved through RepConv, while an efficient channel attention mechanism is implemented to represent the interrelations across feature mapping channels, enhancing the capture of small target features. Subsequently, the Dynamic-range Histogram Self-Attention is incorporated into the intra-scale feature interaction module. This approach enables thorough feature extraction at both local and global levels, hence minimizing the false detection rate. Furthermore, the ESlimneck-DASF architecture is suggested to enhance crossscale feature fusion. This framework fully utilizes the benefits of dynamic upsampling and feature fusion, enriching semantic information across various scales. Ultimately, the Inner-MPDIoU loss function with an optimal ratio was chosen to enhance convergence speed and detection accuracy. Experimental findings on the VisDrone2019 and NWPU VHR-10 datasets indicate that the mAP values are 44.1% and 90.7%, representing increases of 1.7% and 2.4% over RT-DETR-r18. Meanwhile, parameters and GFLOPs are diminished by 12.17% and 9.95%. Tianyu Guo 0012, Qingzeng Song, Yongjiang Xue, Fei Qiao |
IJCNN | 4 |
| 2025 | FPGA-Accelerated CNN-Transformer Hybrid Model for Real-Time Semantic Segmentation in Autonomous Driving
Yongjiang Xue, Fei Qiao, Qingzeng Song |
NPC (1) | 3 |
| 2025 | Joint Scheduling-Maintenance Optimization for Non-identical Parallel Batch Processing Machines with Discrete-State Degradation*abstractBatch scheduling optimization is critical for improving equipment utilization in high-value manufacturing. However, high-intensity continuous operations exacerbate machine degradation effects, leading to frequent unplanned down-time and surges in maintenance costs. Existing studies predominantly assume identical machines and idealized maintenance responses, failing to adapt to real-world production scenarios. To this end, this paper investigates a co-optimization problem integrating batch scheduling with maintenance, which fully considers machine differentiation in capacity and degradation rates. We establish a discrete degradation state model for machines and design state-dependent maintenance policies. For non-identical machine capacities, we develop an adaptive capacity batch formation heuristic (ACBFLPT). To address state observation latency in multi-machine synchronous decision-making, we propose a Sequential QMIX Adaptive Batching (SQAB) algorithm that integrates a sequential decision-making mechanism based on the QMIX framework with ACBFLPT. The performance of our method has been validated through extensive comparative and ablation experiments. Xiong Zheng, Guichen Yan, Fei Qiao |
SMC | 3 |
| 2025 | Production-logistics collaborative scheduling in dynamic flexible job shops using nested-hierarchical deep reinforcement learning
Jiaxuan Shi, Fei Qiao, Juan Liu 0011, Yumin Ma, Dongyuan Wang, Chen Ding 0008 |
Adv. Eng. Informatics | 2 |
| 2025 | A new data-driven production scheduling method based on digital twin for smart shop floors
Yumin Ma, Luyao Li, Jiaxuan Shi, Juan Liu 0011, Fei Qiao |
Expert Syst. Appl. | 5 |
| 2025 | Asynchronous Multi-Agent Collaborative Framework for Integrated Production and Procurement Optimization in RefineryabstractThe refinery industry operates as a highly complex system characterized by dynamic interactions among numerous processes, resources, and decision-making strategies. Within this context, production planning and crude oil procurement management are critical to ensuring operational efficiency and profitability. However, traditional optimization approaches often neglect the distinct decision-making time scales of these two functions, limiting their ability to address market dynamics and supply chain complexities effectively. To overcome these challenges, this study introduces an asynchronous multi-agent collaborative optimization framework that enhances the coordination between production planning and crude oil procurement in refinery operations. By enabling production and procurement agents to operate on independent time scales, the framework adapts to fluctuations in crude oil prices and variations in product demand. The study further extends the classical multi-agent reinforcement learning algorithm into an asynchronous paradigm, introducing Async-MAPPO, which allows agents to independently optimize decisions using real-time data. This approach mitigates operational delays and enhances adaptability. Experimental evaluations demonstrate that the proposed method significantly improves production efficiency, reduces operational costs, and strengthens the refinery’s resilience to market fluctuations, underscoring the critical role of asynchronous optimization in refinery operations. Kai Wang 0024, Fei Qiao, Hanli Wang |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Denoise on Sensor: A Near-Sensor Compute-in-Memory Macro for Visual Perception Denoising via Concatnation-EliminatingabstractNoise is one of the most common and significant factors leading to image degradation. In recent years, due to the rapid development of neural networks, the performance of denoising algorithms has seen a substantial improvement. However, state-of-the-art denoising models often entail large-scale models and heavy computational requirements, making deployment challenging. Additionally, we have observed that the location of denoisers in the entire image processing pipeline has a significant impact on resource consumption and denoising effectiveness. Deploying the denoiser closer to the image acquisition stage will be more effective in separating noise from the image. In this paper, we propose a near-sensor compute-in-memory macro for visual perception denoising (Denoise on Sensor, DoS) along with its corresponding Edge Denoise U-net (EDU) architecture. DoS employs a mixed-signal circuit implementation for neural network inference, offering a notable advantage in terms of high speed and low power consumption compared to FPGA or GPU based approaches, making it feasible to deploy denoising tasks at the near-sensor edge. The simulation results show that the energy efficiency of DoS can reach 21.98 TOPS/W, and EDU deployed on DoS can achieve around 30dB PSNR and 0.83 SSIM on KODAK, BSD300 and SET14 datasets. Aolin You, Erxiang Ren, Daniel Zheng Fang, Cheng Qu, Qi Wei 0001, Fei Qiao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2025 | AM-CIM: Approximate Memory Based Near Sensor Compute-in-Memory Architecture for Keyword SpottingabstractCompute-In-Memory (CIM) has emerged as a promising solution to address the von-Neumann bottleneck, making it a key technology for intelligent computing in edge IoT devices, particularly for real-time applications like keyword spotting (KWS). However, traditional CIM architectures face challenges such as high resource consumption, especially in data conversion, which can significantly impact chip area and energy efficiency. To address these challenges, this work proposes a computational CIM architecture utilizing multilevel analog memory, named AM-CIM, tailored for near-sensor (NS) computation of real-time KWS applications. Additionally, approximate memory technology is integrated into the AM-CIM architecture, employing data resilience scheduling for analog memory which contributes to significant reductions in hardware overhead. This integration facilitates a hardware-software co-design approach. To deploy KWS tasks in AM-CIM, a gated recurrent unit (GRU) network, referred to as MAC-GRU, is implemented. By employing Mel-energy as the input feature at the near-sensor end, the system achieves a 93.13% reduction in feature extraction power consumption. Evaluation results based on TSMC 180-nm technology demonstrate that the AM-CIM architecture achieves an accuracy of 88.51% for 10-keyword classification with a power consumption of$546~\mu W$, while reducing analog memory area by 43.32%. Xiaotao Jia, Guangcai Yuan, Jianyi Yu, Cong Shi 0003, Qi Wei 0001, Youguang Zhang, Weisheng Zhao 0001, Fei Qiao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2024 | NS-Engine: Near-Sensor Neural Network Engine with SRAM-Based Compute-in-Memory MacroabstractSensing devices at edge nodes are usually resource-constrained, such as limited battery capacity and physical size. As a result, there is a substantial demand for enhancing energy efficiency in these devices, which can otherwise hinder the deployment of more complex neural networks. This work proposes an energy-efficient computing engine, which is equipped with appropriate computing power to deploy medium-size neural networks for smart sensing at near-sensor edge nodes. SRAM-based Compute-in-Memory (CIM) macro and end-to-end digital controller comprise the engine. We successfully prototyped this engine using an FPGA platform and conducted a demonstration of image classification on the CIFAR-10 dataset. Furthermore, we implement the above design on a TSMC 40nm process, with post-simulation results indicating an impressive throughput of 368.64GOPS and an energy efficiency of 25.7TOPS/W at 10MHz operating frequency, given the memory capacity is 576kb. Erxiang Ren, Xinghua Yang, Qi Wei 0001, Fei Qiao |
ISCAS | 6 |
| 2024 | Cambricon-M: A Fibonacci-Coded Charge-Domain SRAM-Based CIM Accelerator for DNN InferenceabstractCharge-domain SRAM-based Computing-in-memory (CIM) proves to be a promising method for DNN inference, and benefits from avoiding data movement between computing units and memory. However, the high resolution Analog-to-Digital Converters (ADCs) dominates the energy consumption (up to 64%), limiting the energy efficiency of SRAM-CIM architectures. The main reason is the wide range of input analog values, requiring high resolution ADCs to convert the high precision averaged analog voltages into high bitwidth digital data. In this paper, to reduce the ADC overhead, we propose Cambricon-M, a novel Fibonacci-coded SRAM-based charge-domain CIM accelerator for DNN inference. Cambricon-M features the Fibonacci coding, which guarantees low density of ‘1’ in operands (i.e., the adjacent two bits of each ‘1’ are both ‘0’), narrowing the output voltage range and enabling low resolution ADCs. Further, Cambricon-M exploits the high bit-level sparsity to address the extra energy and area overhead caused by the larger bitwidth in Fibonacci coding. Specifically, Cambricon-M proposes zero-skipping methods to reduce ineffectual input/output, and the bit-slice based compression method to reduce memory capacity/bandwidth pressure. Experimental results show that Cambricon-M reduces ADC energy by 68.7%, and improves the energy efficiency 3.48× and 1.62× compared to TPUv4 and an ISAAC-based charge-domain SRAM-CIM accelerator. Hongrui Guo, Mo Zou, Yifan Hao 0001, Zidong Du, Erxiang Ren, Yang Liu 0466, Yongwei Zhao 0001, Tianrui Ma, Rui Zhang 0040, Xing Hu 0001, Fei Qiao, Zhiwei Xu 0002, Qi Guo 0001, Tianshi Chen 0002 |
MICRO | 11 |
| 2024 | An Incremental Remaining Useful Life Prediction Method Based on Wasserstein GAN and Knowledge DistillationabstractThe precise and prompt estimation of remaining useful life (RUL) for equipment under diverse operating conditions can assist in proactive equipment maintenance and prevent failures that may lead to financial loss and casualties. This article proposes a novel task-incremental RUL prediction method based on Wasserstein GAN with gradient penalty and Knowledge Distillation (WGAN-KD) to achieve high-precision and rapid prediction. WGAN-KD develops a dual old task retention model to ensure the retention of old tasks while facilitating the acquisition of new ones. To evaluate the performance of WGAN-KD, several experiments were conducted on rolling bearings under diverse operating conditions. The experimental results demonstrated that WGAN-KD outperforms the compared incremental learning methods in accuracy under different operating conditions. Furthermore, it maintains high-precision prediction while enhancing training efficiency compared to batch learning methods. Xiaorui He, Fei Qiao, Jiaxuan Shi |
SMC | 3 |
| 2024 | Hybrid Variable Neighborhood Search Algorithm for the Multi-objective Distributed Permutation Flowshop Scheduling Problem with Sequence-Dependent Setup TimesabstractThis paper addresses the multi-objective distributed permutation flowshop scheduling problem with sequence-dependent setup times (MODPFSP_SDST), whose optimization objectives are the makespan, the total energy consumption and the noise emission. An effective hybrid variable neighborhood search (HVNS) algorithm is proposed for solving this problem. HVNS utilizes a probabilistic matrix to store structural information of the promising solution, and then employs matrix-based search operators to execute variable neighborhood search. Additionally, the objective-oriented local intensification search is integrated to further enhance solution quality. Simulations and comparisons demonstrate the effectiveness of HVNS in solving the MODPFSP_SDST. Mingzhe She, Fei Qiao, Yumin Ma, Jiakang Ai, Juan Liu 0011 |
SMC | 2 |
| 2024 | Production-Logistics Collaborative Scheduling in Dynamic Flexible Job Shops via Multi-Objective Deep Reinforcement LearningabstractProduction scheduling and logistics scheduling are vital means for organizing manufacturing activities in flexible job shops. Given the intricate coupling relationship between them, corresponding collaborative scheduling becomes urgent need and challenging. Meanwhile, the actual manufacturing process is inevitably affected by disturbances, necessitating the consideration of dynamic environments. To this end, this study investigates a new production-logistics collaborative scheduling problem in dynamic flexible job shops (PLCSP-DFJS). The high-frequency disturbance of new job arrivals is incorporated into the PLCSP-DFJS, and two objectives, namely makespan and total logistics cost, are optimized. A multi-objective deep reinforcement learning (MODRL) method is presented to solve PLCSP-DFJS. In MODRL, a weight-decomposition and neighborhood-inheritance training mechanism is devised to obtain the near-optimal Pareto front, and a dual-level channel-driven framework capable of achieving decentralized decision-making of production and logistics is designed. The performance of MODRL is verified through experiments conducted in an aviation component production shop. Jiaxuan Shi, Fei Qiao, Yumin Ma |
SMC | 2 |
| 2023 | Memory-Efficient and Real-Time SPAD-based dToF Depth Sensor with Spatial and Statistical CorrelationabstractSingle Photon Avalanche Diode (SPAD)-based direct time-of-flight (dToF) depth sensors are widely used in Internet of Things (IoT) devices due to their high accuracy. Existing SPAD-based dToF sensors measure depth by continually accumulating the depth-measured value in a histogram. However, histogram-based methods typically have low convergence speed (~10 frames per second (FPS)) and large memory overhead (MB-level), hindering their use in real-time embedded IoT devices. To overcome these two challenges, we propose SSC, a histogram-free Spatial and Statistical Correlation based depth measurement method. On the one hand, SSC applies the spatial correlation of the adjacent pixels to accelerate the convergence speed. On the other hand, SSC explores the statistical correlation of depth measurements to reduce the memory overhead. In order to implement SSC with small hardware area and low power, we design mert-dToF, a memory-efficient and real-time dToF sensor for efficient execution. mert-dToF abstracts mainly operations in SSC into four basic operators and designs corresponding hardware with a fine-grained pipeline to maximize resource reuse and computational parallelism. Extensive experiments show that compared with state-of-the-art (SOTA) histogram-based dToF sensors, mert-dToF achieves ~8% accuracy improvement and 7.80× speedup (from 6.24 FPS to 48.70 FPS). The memory overhead is reduced by up to 60.91% (from 48 KB to 18.75 KB). Zhenhua Zhu 0002, Qingpeng Zhu, Jiangwei Zhang, Wenxiu Sun, Guohao Dai 0001, Fei Qiao, Huazhong Yang, Yu Wang 0002 |
DAC | 8 |
| 2023 | A Three-Step Multi-Resolution Time-to-Digital ConverterabstractThis work proposes a three-step multi-resolution time-to-digital converter (TDC) architecture based on the vernier delay line (VDL). The proposed architecture uses a delay-locked loop (DLL) to control TDC with a smooth coarse-to-fine strategy. In addition, the fine TDC uses a combination of multiple resolutions to reduce the number of delay cells and flip-flops. This architecture helps to reduce the area and power consumption and maintains high resolution. We proposed architecture performs better trade-offs between power consumption, linearity, accuracy, and measurement range. The simulation results show that the 7-bit TDC based on VDL designed in 180 nm CMOS achieves 5 ps of time resolution, 0.76/-0.8 LSB DNL and 1.02/-1.39 LSB INL at 100 MHz clock frequency while consuming 3.1 mW, which corresponds to the figure of merit (FoM) of 0.242 pJ/Conv. Jiang Yan, Yu Wang 0002, Fei Qiao, Jiangwei Zhang, Qi Wei 0001, Qingpeng Zhu, Wenxiu Sun, Ge Shi 0001 |
ISCAS | 5 |
| 2023 | A Two-Stage Search-Enhanced Evolutionary Algorithm for an Aerospace Component Production Scheduling ProblemabstractAerospace components (ACs) are an important part of the aerospace equipment. Because of process specialty, AC production scheduling has always been seen a challenging work. In this paper, an improved dual-resource-constrained flexible job shop scheduling model is constructed to formulate the problem. The characteristics of AC manufacturing process are fully considered by analyzing the relationship among processes, machines, and workers. Meanwhile, the effect of machine type, worker skill level and worker labor efficiency are also incorporated into the model. To solve the model, we design an evolutionary strategy combined with TOPSIS and used Metropolis guidelines and perturbation operators to ameliorate the search process. A two-stage search-enhanced multi-objective evolutionary algorithm is proposed, which aims to minimize the completion time, production cost and worker load imbalance. Finally, experiments are conducted based on real AC production data. Experimental results verify the proposed algorithm is effective and competitive, which provides certain application value for relevant enterprises. Zizhao Chen, Fei Qiao, Dongyuan Wang, Juan Liu 0011 |
SMC | 2 |
| 2023 | A Phased Scheduling Method with an Improved GA for Material Delivery Problem of Aircraft Pulsating Assembly LineabstractTimely delivery of materials is essential in ensuring the smooth operation of aircraft pulsating assembly lines. According to the production characteristic of aircraft pulse-like movement in pulsating assembly lines, a material delivery model is established to minimize the time window constraint penalties and delivery costs. Due to the problem characteristics of large scale, long decision cycle and uneven distribution of tasks in time, a phased scheduling method based on an improved genetic algorithm is proposed. Firstly, the decision cycle can be divided into the aircraft moving phase and the assembling phase. Secondly, the assembling phase is subdivided into several small phases to adjust the number of used AGVs and the delivery starting time of each phase. Finally, a genetic algorithm is used within each phase to decide the delivery tasks, delivery starting time and driving paths for every AGV, which has been improved for solving the large-scale problem by optimizing the generation of initial populations and adding a local search operator. The effectiveness of the proposed algorithm is verified in a practical case of an aircraft pulsating assembly line. Yiwen Fang, Fei Qiao, Juan Liu 0011 |
SMC | 2 |
| 2023 | Knowledge graph modeling method for product manufacturing process based on human-cyber-physical fusion
Chen Ding 0008, Fei Qiao, Juan Liu 0011, Dongyuan Wang |
Adv. Eng. Informatics | 2 |
| 2023 | Breaking the energy-efficiency barriers for smart sensing applications with "Sensing with Computing" architectures
Xinghua Yang, Zheyu Liu, Kechao Tang, Xunzhao Yin, Cheng Zhuo, Qi Wei 0001, Fei Qiao |
Sci. China Inf. Sci. | 7 |
| 2023 | A new boredom-aware dual-resource constrained flexible job shop scheduling problem using a two-stage multi-objective particle swarm optimization algorithm
Jiaxuan Shi, Mingzhou Chen, Yumin Ma, Fei Qiao |
Inf. Sci. | 4 |
| 2023 | A Survey of Approximate Computing: From Arithmetic Units Design to High-Level Applications
Haohua Que, Mingkai Liu, Xinghua Yang, Fei Qiao |
J. Comput. Sci. Technol. | 6 |
| 2023 | Solving a many-objective PFSP with reinforcement cumulative prospect theory in low-volume PCB manufacturing
Fei Qiao, Guangyu Zhu 0005 |
Neural Comput. Appl. | 2 |
| 2023 | Double deep Q-network-based self-adaptive scheduling approach for smart shop floor
Yumin Ma, Shengyi Li, Juan Liu 0011, Jianmin Xing, Fei Qiao |
Neural Comput. Appl. | 6 |
| 2023 | Human-Machine Interactive Learning Method Based on Active Learning for Smart Workshop Dynamic SchedulingabstractIn the field of dynamic scheduling, workers and scheduling models (SMs) play a crucial role in decision-making. Workers are able to help SM training by sample labeling, thereby enhancing the decision-making ability of SMs. However, existing supervised learning methods require a large number of labeled samples to train SMs, which limits the learning efficiency between workers and SMs. In this article, a human-machine interactive learning method based on active learning (HMILM/AL) is proposed. The method introduces active learning (AL) techniques to reduce labeling costs and improve learning efficiency. Referring to the AL framework, only a small subset of samples are selected from an unlabeled dataset and are labeled by workers, to train SMs. To further reduce labeling costs, sample selection, the key to the HMILM/AL, is improved by two strategies. First, a novel hybrid selection strategy (NHSS) is developed. By identifying and selecting more useful samples in an unlabeled dataset, the NHSS promotes efficient use of workers, and reduces labeling costs. Second, an enhanced NHSS (E-NHSS) is proposed, which considers both the difficulty of labeling samples and the usefulness of the samples. It reduces labeling costs by selecting easily labeled samples as much as possible. Finally, the proposed method is evaluated through experiments conducted in a real smart workshop. The results demonstrate that the HMILM/AL is very competitive compared with existing supervised learning methods. Moreover, both the NHSS and the E-NHSS can reduce labeling costs efficiently. Dongyuan Wang, Liuen Guan, Juan Liu 0011, Chen Ding 0008, Fei Qiao |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2022 | A 2.17μW@120fps Ultra-Low-Power Dual-Mode CMOS Image Sensor with Senputing ArchitectureabstractThis paper proposes an ultra-low-power CMOS Image Sensor (CIS) chip based on sensing-with-computing (Senputing) architecture to reduce the power bottleneck of vision system. This Senputing chip achieves BNN 1st-layer convolution in analog domain with ultra-low power consumption. It has two working modes, Normal-Sensor (NS) mode and Direct- Photocurrent-Computation (DPC) mode. The prototype measurement results under 65nm CMOS process on MNIST classification task shows that the power of feature map computation is 2.17μW with 120fps frame rates and 98.1% accuracy. The computation efficiency reaches to 11.49TOPs/W, which is 14.8× higher than state-of-art works. Han Xu 0006, Zheyu Liu, Qi Wei 0001, Fei Qiao |
ASP-DAC | 6 |
| 2022 | In-situ self-powered intelligent vision system with inference-adaptive energy scheduling for BNN-based always-on perceptionabstractThis paper proposes an in-situ self-powered BNN-based intelligent visual perception system that harvests light energy utilizing the indispensable image sensor itself. The harvested energy is allocated to the low-power BNN computation modules layer by layer, adopting a light-weighted duty-cycling-based energy scheduler. A software-hardware co-design method, which exploits the layer-wise error tolerance of BNN as well as the computing-error and energy consumption characteristics of the computation circuit, is proposed to determine the parameters of the energy scheduler, achieving high energy efficiency for self-powered BNN inference. Simulation results show that with the proposed inference-adaptive energy scheduling method, self-powered MNIST classification task can be performed at a frame rate of 4 fps if the harvesting power is 1μW, while guaranteeing at least 90% inference accuracy using binary LeNet-5 network. Maimaiti Nazhamaiti, Haijin Su, Han Xu 0006, Zheyu Liu, Fei Qiao, Qi Wei 0001, Zidong Du, Xinghua Yang |
DAC | 5 |
| 2022 | OCTOANTS: A Heterogeneous Lightweight Intelligent Multi-Robot Collaboration System with Resource-constrained IoT DevicesabstractAs the focus on highly intelligent robots continues, a problem that cannot be ignored has emerged: resource con-straints. Considering the game problem of resource limitation and the level of intelligence, we focus on lightweight intelligence. This work is a further refinement of our previous work, a heterogeneous lightweight intelligent multi-robot system. In-spired by the nature creatures “octopus” and “ants”. First, we propose a heterogeneous centralized-distributed architecture, which can make robots collaboration more flexible and non-redundant. Second, to reflect lightweight intelligence, we use the Raspberry Pi, a low computing and power consumption internet of things (IoT) device, as a processing platform and first propose a quantitative definition of the lightweight intelligent system. Then, combining the centralized-distributed architecture and the lightweight computing platform, we propose an adapted algorithm called OCTOANTS and apply it to the simultaneous localization and mapping (SLAM) field. The OCTOANTS architecture consists of one brain and eight tentacles, which can achieve complex things with proper collaboration between them. Finally, we use heterogeneous cameras and heterogeneous algorithms to form a lightweight intelligent collaborative system that can run in the real world. On the low-grade platform Raspberry Pi our heterogeneous tentacles frame rate can reach 41fps and 99.8fps respectively, power consumption is only 2W and 1.2W. At the same time, our heterogeneous system is on average 7.2% more accurate than the state-of-the-art homogeneous system and can be applied to a wider range of application scenarios, demonstrating the superiority and feasibility of our OCTOANTS. Ruiyang Quan, Siqin Qimuge, Peimin Xia, Xin Zan, Fangshi Wang, Changchuan Chen, Qi Wei 0001, Huichan Zhao, Fei Qiao |
IROS | 12 |
| 2022 | HOGEye: Neural Approximation of HOG Feature Extraction in RRAM-Based 3D-Stacked Image SensorsabstractMany computer vision tasks, ranging from recognition to multi-view registration, operate on feature representation of images rather than raw pixel intensities. However, conventional pipelines for obtaining these representations incur significant energy consumption due to pixel-wise analog-to-digital (A/D) conversions and costly storage and computations. In this paper, we propose HOGEye, an efficient near-pixel implementation for a widely-used feature extraction algorithm—Histograms of Oriented Gradients (HOG). HOGEye moves the key but computation-intensive derivative extraction (DE) and histogram generation (HG) steps into the analog domain by applying a novel neural approximation method in a resistive random-access memory (RRAM)-based 3D-stacked image sensor. The co-location of perception (sensor) and computation (DE and HG) and the alleviation of A/D conversions allow HOGEye design to achieve significant energy saving. With negligible detection rate degradation, the entire HOGEye sensor system consumes less than 48μ[email protected] for an image resolution of 256 × 256 (equivalent to 24.3pJ/pixel) while the processing part only consumes 14.1pJ/pixel, achieving more than 2.5 × energy efficiency improvement than the state-of-the-art designs. Tianrui Ma, Weidong Cao 0001, Fei Qiao, Ayan Chakrabarti, Xuan Zhang 0001 |
ISLPED | 3 |
| 2022 | A method of remaining useful life prediction of multi-source signals aero-engine based on RF-Transformer-LSTMabstractThe aeroengine remaining useful life (RUL) prediction problem of prognostics and health management (PHM) is becoming more complicated and challenging due to the large-scale, high dimension and difficult feature extraction of sensor data. This paper proposes a RUL prediction method with multi-source signals based on Random Forest (RF)-Transformer-Long Short-Term Memory (LSTM) to enhance the prediction accuracy. The method proposed in this paper mainly includes the following structures: Firstly, RF is used to select the feature of multi-sensor data. Secondly, Transformer is used to extract the feature after feature selection. Thirdly, the data after feature extraction is put into the LSTM model to obtain the aeroengine RUL. Finally, a case study is conducted to validate the superiority of the proposed RF-Transformer-LSTM method based on C-MAPSS aeroengine dataset. The results show that under the four data sets, the root mean square error(RMSE) of the method used in this paper can reach a minimum of 10.23, which is far lower than other common methods. So, the proposed RF-Transformer-LSTM method has higher accuracy than other common methods. Hanshuo Mu, Xiaodong Zhai, Debin Yin, Fei Qiao |
SMC | 4 |
| 2022 | An efficient adaptive genetic algorithm for energy saving in the hybrid flow shop scheduling with batch production at last stageabstractAbstract This article deals with energy saving in the hybrid flow shop scheduling problem with batch production at last stage, which has important application in energy‐intensive steelmaking‐continuous casting (SCC) process. We first establish a mixed integer programming model to reduce extra energy consumption, and then adopt genetic algorithm to solving the scheduling problem. Based on traditional genetic algorithm (TGA), the calculation of the fitness function as well as adaptive crossover and mutation are designed. Due to the complexity of the problem in this article, we then propose an efficient adaptive genetic algorithm (EAGA) to improve the search ability of TGA. The EAGA has new features including layered strategies and enhanced adaptive adjustment method. To evaluate the proposed model and algorithm, we conduct computational experiments under practical background and compare the EAGA with the several algorithms presented previously. The results illustrate that scheduling with our model can greatly reduce the extra energy consumption. Meanwhile, the proposed EAGA is very efficient in comparison. Fei Qiao |
Expert Syst. J. Knowl. Eng. | 2 |
| 2022 | Towards lifelong object recognition: A dataset and benchmark
Chuanlin Lan, Qi Liu 0042, Qi She, Qihan Yang, Xinyue Hao 0001, Ivan Mashkin, Ka Shun Kei, Dong Qiang, Vincenzo Lomonaco, Xuesong Shi, Yimin Zhang 0002, Fei Qiao, Rosa H. M. Chan |
Pattern Recognit. | 15 |
| 2022 | Senputing: An Ultra-Low-Power Always-On Vision Perception Chip Featuring the Deep Fusion of Sensing and ComputingabstractAlways-on intelligent visual perception applications are widely deployed in edges in the AIoT era. In order to eliminate power costs of data conversion and transmission, this paper proposes Senputing, an ultra-low-power processing-in-sensor chip that completely fuses sensing and computing together for a BNN-based hierarchical processing system. This chip could operate in two modes. In computation mode, photocurrents are directly utilized for computing without being converted into voltages, and the computation results of 1-st BNN layer are directly sent out to subsequent BNN processors for an always-on coarse classification, eliminating conversion power and storage cost of raw images. Once an interested objected is detected, this chip switches to sensor mode and sends raw images to potential full-precision processors or cloud servers for fine-grained recognition or segmentation. A$32\times 32$prototype is fabricated with 180nm CMOS process. It accomplishes MNIST dataset classification task with the accuracy of 93.76% and the power consumption of 147nW at 156fps, achieving$13.1\times $energy efficiency compared with state-of-the-art work. Han Xu 0006, Ningchao Lin, Qi Wei 0001, Runsheng Wang, Cheng Zhuo, Xunzhao Yin, Fei Qiao, Huazhong Yang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2021 | Puncturing the memory wall: Joint optimization of network compression with approximate memory for ASR applicationabstractThe automatic speech recognition (ASR) system is becoming increasingly irreplaceable in smart speech interaction applications. Nonetheless, these applications confront the memory wall when embedded in the energy and memory constrained Internet of Things devices. Therefore, it is extremely challenging but imperative to design a memory-saving and energy-saving ASR system. This paper proposes a joint-optimized scheme of network compression with approximate memory for the economical ASR system. At the algorithm level, this work presents block-based pruning and quantization with error model (BPQE), an optimized compression framework including a novel pruning technique coordinated with low-precision quantization and the approximate memory scheme. The BPQE compressed recurrent neural network (RNN) model comes with an ultra-high compression rate and finegrained structured pattern that reduce the amount of memory access immensely. At the hardware level, this work presents an ASR-adapted incremental retraining method to further obtain optimal power saving. This retraining method stimulates the utility of the approximate memory scheme, while maintaining considerable accuracy. According to the experiment results, the proposed joint-optimized scheme achieves 58.6% power saving and 40x memory saving with a phone error rate of 20%. Qin Li 0016, Peiyan Dong, Zijie Yu, Changlu Liu, Fei Qiao, Yanzhi Wang 0001, Huazhong Yang |
ASP-DAC | 5 |
| 2021 | RaP-Net: A Region-wise and Point-wise Weighting Network to Extract Robust Features for Indoor LocalizationabstractFeature extraction plays an important role in visual localization. Unreliable features on dynamic objects or repetitive regions will interfere with feature matching and challenge indoor localization greatly. To address the problem, we propose a novel network, RaP-Net, to simultaneously predict region-wise invariability and point-wise reliability, and then extract features by considering both of them. We also introduce a new dataset, named OpenLORIS-Location, to train the proposed network. The dataset contains 1553 images from 93 indoor locations. Various appearance changes between images of the same location are included and can help the model to learn the invariability in typical indoor scenes. Experimental results show that the proposed RaP-Net trained with OpenLORIS-Location dataset achieves excellent performance in the feature matching task and significantly outperforms state-of-the-arts feature algorithms in indoor localization. The RaPNet code and dataset are available at https://github.com/ivipsourcecode/RaP-Net. Dongjiang Li, Jinyu Miao, Xuesong Shi, Qiwei Long, Tianyu Cai, Hongfei Yu, Wei Yang 0029, Haosong Yue, Qi Wei 0001, Fei Qiao |
IROS | 12 |
| 2021 | A 5.9μW Ultra-Low-Power Dual-Resolution CIS Chip of Sensing-with-Computing for Always-on Intelligent Visual DevicesabstractIn the intelligent IoT edge devices, power consumption is increasing due to the deployment of high-precision algorithms, which greatly limits the working time of the devices. The power of A/D conversion and data transmission has become the bottleneck of traditional visual system. In this paper, a new sensing-with-computing (Senputing) architecture is proposed to reduce this power bottleneck by combining imaging and BNN 1st-layer feature map computation. This Senputing architecture has two working modes, Normal-Sensor mode and Direct-Photocurrent-Computation mode with different resolutions (128×128 and 32×32). An ultra-low-power CMOS image sensor (CIS) chip with Senputing architecture is proposed to verify the feasibility. Our CIS chip is simulated with 180nm CMOS technology, the power of feature map computation is 5.9μW, and the frame rate is 208fps. The computation efficiency reaches to 8.23TOPs/W, which is 10.1 x higher than previous works. Han Xu 0006, Qi Wei 0001, Fei Qiao |
ISCAS | 5 |
| 2021 | A Hybrid Metaheuristic Algorithm with Novel Decoding Methods for Flexible Flow Shop Scheduling Considering Human FatigueabstractHuman are the key production resources of enterprises, and factors such as human skills and fatigue will affect the implementation of the scheduling strategy. Aiming at the dual resource constrained flexible flow shop scheduling problem (DRC-FFSP), this paper proposes a mixed integer programming (MIP) model for the flexible flow shop to minimize makespan, with the constraints of machines and heterogeneous human who have different skills and characteristics. According to the characteristics of the model, two paradigms of a hybrid metaheuristic algorithm (HMA) are proposed, which combine genetic algorithm with two novel heuristic decoding methods, respectively. A new methodology which aims to provide an adaptable assignment heuristic algorithm is designed to allocate human during manufacturing process to meet the fatigue constraint. A case study from benchmarks demonstrates the effectiveness of the proposed model and algorithms. Furthermore, different production scales are designed to verify the superiority and stability of the two paradigms. The experimental results show that the scheduling strategy based on the hybrid metaheuristic algorithm can meet the human fatigue constraint while ensuring economic benefits. Hangming Du, Fei Qiao, Hong Lu 0013 |
SMC | 2 |
| 2021 | Reducing SRAM Reading Power With Column Data Segment and Weights Correlation Enhancement for CNN ProcessingabstractConvolutional neural network (CNN) has been widely deployed in various processors for intelligent visual signal processing. However, the large amount of activations and weights in CNN causes huge power consumption on SRAM access. Data-adaptive SRAM design is a widely studied method to reduce SRAM reading power based on the utilization of data patterns, while current designs only exploit data patterns in a coarse granularity, and have no advantages when faced with randomly distributed weight data. In this article, we propose a hardware–software co-design scheme to reduce SRAM reading power for CNN processing. First, we propose a reconfigurable data-adaptive SRAM architecture with column data segmentation (CDS-RSRAM) to utilize data patterns. Data in one column is partitioned into several segments, and finer-grained data patterns are exploited within each segment for further reading power reduction. Then, a novel training method—minimum segmented neighbor difference (miniSND)—is proposed for enhancing the correlation of weights. MiniSND improves the similarity of weights without classification accuracy degradation, thus weights could benefit from CDS-RSRAM and be read out with less power consumption. Simulation results demonstrate that the co-design scheme saves up to 66%(8b)/89%(2b) power consumption compared with 8T SRAM. Han Xu 0006, Ziru Li, Deliang Fan, Fei Qiao, Qi Wei 0001, Huazhong Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | NS-FDN: Near-Sensor Processing Architecture of Feature-Configurable Distributed Network for Beyond-Real-Time Always-on Keyword SpottingabstractAlways-on keyword spotting (KWS) that detects wake-up words has been the indispensable module in the voice interaction system. However, the ultra-low-power embedded devices put forward strict requirements on energy consumption, latency, and recognition accuracy of KWS. In this work, we propose a near-sensor processing architecture of feature-configurable distributed network (NS-FDN) for always-on KWS applications. The proposed distributed network adapts to the flexible keywords demands in the actual scene by splitting the conventional single network into distributed sub-networks. We design a channel-independent training framework to improve the recognition accuracy of distributed networks. The speech features are evaluated and the redundancy is reduced in NS-FDN, which can also configure the speech features to further reduce the computing complexity and improve processing speed. For deeper optimization, we implement a 65nm-process prototype chip with near-sensor mixed-signal processing architecture avoiding energy-consuming analog-to-digital converter. By improving the system, algorithm, and hardware designs of the KWS, our co-optimized architecture eliminates the energy consumption bottleneck long-standing in conventional KWS systems and achieves state-of-the-art system performance. The experiment results show that NS-FDN achieves 31.6% energy consumption savings, 1.6 times memory savings, 57 times speedup, and 3.4% higher recognition accuracy compared with the state of the art. Qin Li 0016, Changlu Liu, Peiyan Dong, Sheng Lin 0001, Minda Yang, Fei Qiao, Yanzhi Wang 0001, Huazhong Yang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2021 | Equipment Health Assessment Based on Improved Incremental Support Vector Data DescriptionabstractWith the rapid development of Internet-of-Things and big data, health assessment of equipment is receiving more attention in recent years. It is critical to bridge the gap between real-time production data and health status evaluation, which helps maintenance team understand the health status of equipment exactly, and then make rational maintenance plans. For this purpose, this paper proposes a framework to realize real-time equipment health assessment with health status quantitatively characterized by health degree (HD). The proposed framework begins with removing redundant features using a principal component analysis (PCA) method. Then, to represent the optimal operation status, a support vector data description (SVDD) algorithm is employed for extracting normal observations in the offline part. Thereafter, HD is introduced based on the Euclidean distance between current observation and the normal sample set. In order to achieve online updating of the normal sample set, and promote accuracy and computational efficiency of the offline part, an improved incremental SVDD algorithm based on adaptive threshold N (NISVDD) is proposed. A case study is used to demonstrate the effectiveness of the proposed framework and model using a benchmark dataset of rolling bearing. Results suggest that the proposed framework is effective, and PCA shows good potential to extract features and keep most of the original information. The proposed NISVDD model is able to trace the dynamics of equipment health status for whole run-to-failure process, and outperforms other models in both accuracy and computational efficiency. Lianlian Zhang, Fei Qiao, Xiaodong Zhai |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2020 | Utilizing Direct Photocurrent Computation and 2D Kernel Scheduling to Improve In-Sensor-Processing EfficiencyabstractDeploying intelligent visual algorithms in terminal devices for always-on sensing is an attractive trend in the IoT era. In-sensor-processing architecture is proposed to reduce power consumption on A/D conversion and data transmission, which performs pre-processing and only converting low-throughput features. However, current designs still require high energy consumption on photoelectric conversion and analog data movement. In this paper, two methods are proposed to improve the energy efficiency of in-sensor-processing architecture, including direct photocurrent computation and 2D kernel scheduling. Photocurrents are directly involved in computation to avoid data conversion; thus the indispensable imaging power is also utilized for computing. Since the location of the pixel data is fixed, data scheduling is conducted on digital weights to eliminate analog data storage and movement. We implement a prototype chip with an array of 32 × 32 units to calculate the first layer of binarized LeNet-5. The post-simulation shows that the proposed architecture reaches the energy efficiency of 11.49TOPs/W, about 14.8x higher than previous works. Han Xu 0006, Maimaiti Nazhamaiti, Yidong Liu, Fei Qiao, Qi Wei 0001, Huazhong Yang |
DAC | 4 |
| 2020 | OpenLORIS-Object: A Robotic Vision Dataset and Benchmark for Lifelong Deep LearningabstractThe recent breakthroughs in computer vision have benefited from the availability of large representative datasets (e.g. ImageNet and COCO) for training. Yet, robotic vision poses unique challenges for applying visual algorithms developed from these standard computer vision datasets due to their implicit assumption over non-varying distributions for a fixed set of tasks. Fully retraining models each time a new task becomes available is infeasible due to computational, storage and sometimes privacy issues, while naïve incremental strategies have been shown to suffer from catastrophic forgetting. It is crucial for the robots to operate continuously under open-set and detrimental conditions with adaptive visual perceptual systems, where lifelong learning is a fundamental capability. However, very few datasets and benchmarks are available to evaluate and compare emerging techniques. To fill this gap, we provide a new lifelong robotic vision dataset ("OpenLORIS-Object") collected via RGB-D cameras. The dataset embeds the challenges faced by a robot in the real-life application and provides new benchmarks for validating lifelong object recognition algorithms. Moreover, we have provided a testbed of 9 state-of-the-art lifelong learning algorithms. Each of them involves 48 tasks with 4 evaluation metrics over the OpenLORIS-Object dataset. The results demonstrate that the object recognition task in the ever-changing difficulty environments is far from being solved and the bottlenecks are at the forward/backward transfer designs. Our dataset and benchmark are publicly available at https://lifelong-robotic-vision.github.io/dataset/object. Qi She, Xinyue Hao 0001, Qihan Yang, Chuanlin Lan, Vincenzo Lomonaco, Xuesong Shi, Yimin Zhang 0002, Fei Qiao, Rosa H. M. Chan |
ICRA | 11 |
| 2020 | Are We Ready for Service Robots? The OpenLORIS-Scene Datasets for Lifelong SLAMabstractService robots should be able to operate autonomously in dynamic and daily changing environments over an extended period of time. While Simultaneous Localization And Mapping (SLAM) is one of the most fundamental problems for robotic autonomy, most existing SLAM works are evaluated with data sequences that are recorded in a short period of time. In real-world deployment, there can be out-of-sight scene changes caused by both natural factors and human activities. For example, in home scenarios, most objects may be movable, replaceable or deformable, and the visual features of the same place may be significantly different in some successive days. Such out-of-sight dynamics pose great challenges to the robustness of pose estimation, and hence a robot’s long-term deployment and operation. To differentiate the forementioned problem from the conventional works which are usually evaluated in a static setting in a single run, the term lifelong SLAM is used here to address SLAM problems in an ever-changing environment over a long period of time. To accelerate lifelong SLAM research, we release the OpenLORIS-Scene datasets. The data are collected in real-world indoor scenes, for multiple times in each place to include scene changes in real life. We also design benchmarking metrics for lifelong SLAM, with which the robustness and accuracy of pose estimation are evaluated separately. The datasets and benchmark are available online at lifelong-robotic-vision.github.io/dataset/scene. Xuesong Shi, Dongjiang Li, Pengpeng Zhao 0005, Qinbin Tian, Qiwei Long, Chunhao Zhu, Jingwei Song, Fei Qiao, Yangquan Guo, Yimin Zhang 0002, Baoxing Qin, Wei Yang 0029, Fangshi Wang, Rosa H. M. Chan, Qi She |
ICRA | 9 |
| 2020 | DXSLAM: A Robust and Efficient Visual SLAM System with Deep FeaturesabstractA robust and efficient Simultaneous Localization and Mapping (SLAM) system is essential for robot autonomy. For visual SLAM algorithms, though the theoretical framework has been well established for most aspects, feature extraction and association is still empirically designed in most cases, and can be vulnerable in complex environments. This paper shows that feature extraction with deep convolutional neural networks (CNNs) can be seamlessly incorporated into a modern SLAM framework. The proposed SLAM system utilizes a state-of-the-art CNN to detect keypoints in each image frame, and to give not only keypoint descriptors, but also a global descriptor of the whole image. These local and global features are then used by different SLAM modules, resulting in much more robustness against environmental changes and viewpoint changes compared with using hand-crafted features. We also train a visual vocabulary of local features with a Bag of Words (BoW) method. Based on the local features, global features, and the vocabulary, a highly reliable loop closure detection method is built. Experimental results show that all the proposed modules significantly outperforms the baseline, and the full system achieves much lower trajectory errors and much higher correct rates on all evaluated data. Furthermore, by optimizing the CNN with Intel OpenVINO toolkit and utilizing the Fast BoW library, the system benefits greatly from the SIMD (single-instruction-multiple-data) techniques in modern CPUs. The full system can run in real-time without any GPU or other accelerators. The code is public at https://github.com/ivipsourcecode/dxslam. Dongjiang Li, Xuesong Shi, Qiwei Long, Shenghui Liu, Wei Yang 0029, Fangshi Wang, Qi Wei 0001, Fei Qiao |
IROS | 8 |
| 2020 | NS-KWS: joint optimization of near-sensor processing architecture and low-precision GRU for always-on keyword spottingabstractKeyword spotting (KWS) is a crucial front-end module in the whole speech interaction system. The always-on KWS module detects input words, then activates the energy-consuming complex backend system when keywords are detected. The performance of the KWS determines the standby performance of the whole system and the conventional KWS module encounters the power consumption bottleneck problem of the data conversion near the microphone sensor. In this paper, we propose an energy-efficient near-sensor processing architecture for always-on KWS, which could enhance continuous perception of the whole speech interaction system. By implementing the keyword detection in the analog domain after the microphone sensor, this architecture avoids energy-consuming data converter and achieves faster speed than conventional realizations. In addition, we propose a lightweight gated recurrent unit (GRU) with negligible accuracy loss to ensure the recognition performance. We also implement and fabricate the proposed KWS system with the CMOS 0.18μm process. In the system-view evaluation results, the hardware-software co-design architecture achieves 65.6% energy consumption saving and 71 times speed up than state of the art. Qin Li 0016, Sheng Lin 0001, Changlu Liu, Yidong Liu, Fei Qiao, Yanzhi Wang 0001, Huazhong Yang |
ISLPED | 5 |
| 2020 | Human-Machine Cooperation Based Adaptive Scheduling for a Smart Shop FloorabstractWith the increasing demand of personalized products and the application of emerging technologies, substantial unexpected events appears in smart factories. Machine learning based adaptive scheduling shows significant appeal in smart shop floors, yet still has limitations in accommodating unexpected events. This paper presents a novel framework of HCPS (Human Cyber Physical System) based on the conventional CPS. A human-machine cooperative mechanism is proposed to coordinate task allocation between human and machine. Meanwhile, in order to integrate human intelligence and machine intelligence within scheduling decision making, a novel human-machine cooperative approach for adaptive scheduling is put forward. In the process of online scheduling, human operators adjust the deviation of production indicators on the basis of current condition. Subsequently, an enhanced fuzzy inference system combining with human intelligence is designed to obtain optimal dispatching rules, in which parameters are reduced by a K-means algorithm and optimized by a PSO algorithm. Finally, a case study is performed on the Minifab model. The simulation results validate the superiority of the proposed framework and approaches, and show good potential in efficiency and stability. Dongyuan Wang, Fei Qiao, Juan Liu 0011, Weichang Kong |
SMC | 2 |
| 2020 | Fine-grained access control based on Trusted Execution Environment
Yongkai Fan, Shengle Liu, Gang Tan, Fei Qiao |
Future Gener. Comput. Syst. | 4 |
| 2020 | A Novel Rescheduling Method for Dynamic Semiconductor Manufacturing SystemsabstractThis work is motivated by the need to adapt an optimal production schedule for dynamic and stochastic environments in semiconductor manufacturing systems. When unexpected events occur, such as machine breakdown and due date changes, manufacturers need to react quickly and revise the schedule accordingly. This paper presents a novel partial repair rescheduling solution that consists of a criterion and a scheduler. The former decides the segment of the original schedule to be repaired by detecting a match-up point. The impact of disturbances caused by unexpected events can be limited between the disturbing time and this match-up point. The scheduler generates a new schedule segment to replace the original and impacted one with single machine-oriented or machine-group-oriented match-up rescheduling algorithms. On practical demand, the proposed solution can be further updated due to the changing consideration of unexpected events and/or expanded performance criteria. The simulation results show that the proposed rescheduling methods have higher stability and efficiency than some well-known existing rescheduling methods. Fei Qiao, Yumin Ma, MengChu Zhou, Qidi Wu |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2020 | ASP-SIFT: Using Analog Signal Processing Architecture to Accelerate Keypoint Detection of SIFT AlgorithmabstractThe scale-invariant feature transform (SIFT) algorithm is still one of the most reliable image feature extraction methods. Despite its excellent robustness on various image transformations, SIFT's intensive computational burden has been severely preventing it from being used in real-time and energy-efficient embedded machine vision systems. To reduce processing time and energy cost while executing SIFT, an analog signal processing architecture, analog signal processing (ASP)SIFT, is proposed in this article. In ASP-SIFT, the Gaussian pyramid construction, difference-of-Gaussian (DoG) pyramid construction and keypoint locating, which are the primary steps of the keypoint detection part of the SIFT algorithm, are done directly with analog circuit networks. Thus, by completing keypoint detection in the analog domain, the total processing time is approximately equal to the settling time of the circuit network. Besides, by adopting a current-mode circuit network operating in the subthreshold region, the power dissipation would be very low. Simulation results show that the total processing speed for a typical video graphics array (VGA)-format (640 × 480) image is up to 2.3 kframes per second, which is at least 3.26× faster than the state-of-the-art digital hardware accelerators, while the system power is 94.5 mW and the energy consumption is only 40 μJ per frame. Zichen Fan, Zheyu Liu, Zheng Qu 0002, Fei Qiao, Qi Wei 0001, Shuzheng Xu, Huazhong Yang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | Concrete: A Per-layer Configurable Framework for Evaluating DNN with Approximate OperatorsabstractApproximate computing has drawn considerable attention to both academia and industry in the area of DNN hardware. Despite substantial efforts to design approximate circuits and building blocks, the resilience of DNN layers and structures remains an untapped field to explore. This paper presents an efficient framework to evaluate DNN resilience with fine-grained approximate operations, such as multipliers, adders and low-bit operators. The framework can execute large-scale approximate DNNs with relatively less time overhead. Massive experiments are conducted with the proposed framework to reveal the relationship between network structures and error tolerance. Additionally, a case study of fine-tuning the approximate DNN is presented. Zheyu Liu, Guihong Li, Fei Qiao, Qi Wei 0001, Ping Jin, Huazhong Yang |
ICASSP | 3 |
| 2019 | INA: Incremental Network Approximation Algorithm for Limited Precision Deep Neural NetworksabstractApproximate computing is a promising paradigm to deal with large computing workloads in fault-tolerant applications, providing opportunities to improve hardware efficiency of Deep Neural Networks (DNNs). However, it is still difficult to apply highly approximate arithmetics (e.g., multipliers) to DNNs due to the effect of error accumulation and the convergence problem in re-training phase. To tackle this limitation, we propose a hardware-software co-design algorithm, namely Incremental Network Approximation (INA). By addressing the convergence problem, INA promotes fault tolerance of DNNs, and yields more tradeoffs between accuracy and implementation cost. Experiments show that the approximate inference models re-trained by INA could achieve up to 80% hardware reduction in various hardware design level, while the classification accuracy degradation is less than 2%. Moreover, the experiments also exhibit the generality of INA algorithm for applying to various approximate multiplier design. Zheyu Liu, Kaige Jia, Weiqiang Liu 0001, Qi Wei 0001, Fei Qiao, Huazhong Yang |
ICCAD | 5 |
| 2019 | Energy-Aware Cascade optimization for Proportioning in the Sintering Process Using Improved Immune-Simulated Annealing AlgorithmabstractSintering is recognized as one of the most energy-intensive process in the iron and steel enterprise. Energy-aware optimization introduces energy indicators into the sintering proportioning process, so that a reduction in energy consumption could be achieved without affecting normal production on the sintering shop floor. This paper proposes a novel cascade multi-objective optimization model (CMOM) that dynamically produces optimal dosing schemes. Firstly, an improved back-propagation neural network (IBPNN) with appending momentum and adaptive variable learning rate is proposed to better mimic solid energy consumption (SEC) and other indictors of sinter. Then, a cost based proportioning model (CBPM) is established with the consideration of chemical and physical characteristics of the resulting sinter. An improved immune-simulated annealing algorithm (ISAA) is designed to fast converge toward optimal solutions. Subsequently, by achieving feasible dosing scheme with initial chemical and physical requirements, energy-aware indicator is introduced into the cascade optimization framework. The CBPM would be repeatedly optimized by adjusting the expected performance until the predicted SEC and other indicators would have been in accord with the expected intervals. A case study of an actual sintering process in steel enterprise demonstrates the effectiveness of the framework and models. Meanwhile, results show that energy consumption is reduced by 5.28% on an optimal scenario. Yumin Ma, Fei Qiao, Xiaodong Zhai, Juan Liu 0011 |
SMC | 3 |
| 2018 | Mechanical strain and temperature aware design methodology for thin-film transistor based pseudo-CMOS logic arrayabstractThin-film transistor (TFT) circuits are facing the challenges of unipolar device, process variation, and yield problems, which can be addressed by pseudo-CMOS logic array with multi-layer interconnect. However, existing design methodology does not take mechanical strain and temperature into consideration which may seriously affect the carrier mobility of TFT and thus the performance of whole logic array circuits. This paper presents a novel cell mapping algorithm including intrarow mapping step and inter-row mapping step for flexible logic array to mitigate the mobility influence. Experimental results indicate that there is more than 40% performance improvement in critical path delay at best case with the proposed algorithm. Wenyu Sun, Qinghang Zhao, Fei Qiao, Tsung-Yi Ho, Huazhong Yang, Yongpan Liu |
ASP-DAC | 4 |
| 2018 | Calibrating process variation at system level with in-situ low-precision transfer learning for analog neural network processorsabstractProcess Variation (PV) may cause accuracy loss of the analog neural network (ANN) processors, and make it hard to be scaled down, as well as feasibility degrading. This paper first analyses the impact of PV on the performance of ANN chips. Then proposes an in-situ transfer learning method at system level to reduce PV's influence with low-precision back-propagation. Simulation results show the proposed method could increase 50% tolerance of operating point drift and 70% ∼ 100% tolerance of mismatch with less than 1% accuracy loss of benchmarks. It also reduces 66.7% memories and has about 50× energy-efficiency improvement of multiplication in the learning stage, compared with the conventional full-precision (32bit float) training system. Kaige Jia, Zheyu Liu, Qi Wei 0001, Fei Qiao, Yi Yang 0039, Hua Fan 0001, Huazhong Yang |
DAC | 4 |
| 2018 | MINTIN: Maxout-Based and Input-Normalized Transformation Invariant Neural NetworkabstractConvolutional Neural Network (CNN) is a powerful model for image classification, but it is insufficient to deal with the spatial variance of the input. This paper presents a Maxout-based and input-normalized transformation invariant neural network (MINTIN), which aims at addressing the nuisance variation of images and accumulating transformation invariance. We introduce an innovative module, the Normalization, and combine it with the Maxout operator. While the former focuses on each image itself, the latter pays attention to augmented versions of input, resulting in fully-utilized information. This combination, which can be inserted into existing CNN architectures, enables the network to learn invariance to rotation and scaling. While the authors of TI-POOLING acclaimed that they reached state-of-the-art results, ours reach a maximum decrease of 0.71%, 0.23% and 0.51% in error rate on MNIST-rot-12k, half-rotated MNIST and scaling MNIST, respectively. The size of the network is also significantly reduced, leading to high computational efficiency. Jingyang Zhang, Kaige Jia, Pengshuai Yang, Fei Qiao, Qi Wei 0001, Huazhong Yang |
ICIP | 4 |
| 2018 | DS-SLAM: A Semantic Visual SLAM towards Dynamic EnvironmentsabstractSimultaneous Localization and Mapping (SLAM) is considered to be a fundamental capability for intelligent mobile robots. Over the past decades, many impressed SLAM systems have been developed and achieved good performance under certain circumstances. However, some problems are still not well solved, for example, how to tackle the moving objects in the dynamic environments, how to make the robots truly understand the surroundings and accomplish advanced tasks. In this paper, a robust semantic visual SLAM towards dynamic environments named DS-SLAM is proposed. Five threads run in parallel in DS-SLAM: tracking, semantic segmentation, local mapping, loop closing and dense semantic map creation. DS-SLAM combines semantic segmentation network with moving consistency check method to reduce the impact of dynamic objects, and thus the localization accuracy is highly improved in dynamic environments. Meanwhile, a dense semantic octo-tree map is produced, which could be employed for high-level tasks. We conduct experiments both on TUM RGB-D dataset and in real-world environment. The results demonstrate the absolute trajectory accuracy in DS-SLAM can be improved one order of magnitude compared with ORB-SLAM2. It is one of the state-of-the-art SLAM systems in high-dynamic environments. Chao Yu 0005, Zuxin Liu, Fugui Xie, Yi Yang 0039, Qi Wei 0001, Fei Qiao |
IROS | 7 |
| 2018 | Design of Approximate FFT with Bit-width Selection AlgorithmsabstractThis paper presents the approximate designs of Fast Fourier Transformation (FFT) circuit. The tradeoff between accuracy and hardware performance is achieved by using bit-width selection for each stage. The error rate can be tuned with bit-width selection. We proposed two algorithms for bit-width selection under certain error restriction. The first algorithm is targeting an approximate FFT design with low hardware cost. While the second algorithm is proposed to achieve high performance. Both of proposed algorithms allow the designer to tradeoff hardware performance and computation accuracy in each stage. The proposed two designs are implemented on FPGA. The results show that the approximate FFT design using the first algorithm can reduce hardware resource consumption up to 30.2%. The second algorithm can increases the performance of the approximate FFT design up to 24.0%, while it also saves 25.2% resource consumption. Qicong Liao, Weiqiang Liu 0001, Fei Qiao, Chenghua Wang, Fabrizio Lombardi |
ISCAS | 3 |
| 2018 | Real Manufacturing Oriented Data Process Techniques with Domain KnowledgeabstractIn the field of manufacturing industry, it is difficult to make full use of the research results for production optimization and/or management due to the low quality of real workshop data. Typical quality problems of the real workshop data include: data conflict, missing recessive data, and false error identification. The conventional data analysis methods cannot handle most such issues because they fail to consider professional insights into and domain knowledge about the data. The real production data from an actual semiconductor manufacturing workshop are adopted as the objective data in this paper. A series of data process techniques with domain knowledge are proposed to solve those data quality problems according to specific flaws of the data respectively. The work in this paper has the potential to be further extended and applied to other big data applications beyond the manufacturing industry. Weichang Kong, Fei Qiao, Qidi Wu |
SMC | 2 |
| 2017 | Evaluating Data Resilience in CNNs from an Approximate Memory PerspectiveabstractDue to the large volumes of data that need to be processed, efficient memory access and data transmission are crucial for high-performance implementations of convolutional neural networks (CNNs). Approximate memory is a promising technique to achieve efficient memory access and data transmission in CNN hardware implementations. To assess the feasibility of applying approximate memory techniques, we propose a framework for the data resilience evaluation (DRE) of CNNs and verify its effectiveness on a suite of prevalent CNNs. Simulation results show that a high degree of data resilience exists in these networks. By scaling the bit-width of the first five dominant data subsets, the data volume can be reduced by 80.38% on average with a 2.69% loss in relative prediction accuracy. For approximate memory with random errors, all the synaptic weights can be stored in the approximate part when the error rate is less than 10--4, while 3 MSBs must be protected if the error rate is fixed at 10--3. These results indicate a great potential for exploiting approximate memory techniques in CNN hardware design. Yuanchang Chen, Yizhe Zhu, Fei Qiao, Jie Han 0001, Yuansheng Liu, Huazhong Yang |
ACM Great Lakes Symposium on VLSI | 3 |
| 2017 | Region ensemble network: Improving convolutional network for hand pose estimationabstractHand pose estimation from monocular depth images is an important and challenging problem for human-computer interaction. Recently deep convolutional networks (ConvNet) with sophisticated design have been employed to address it, but the improvement over traditional methods is not so apparent. To promote the performance of directly 3D coordinate regression, we propose a tree-structured Region Ensemble Network (REN), which partitions the convolution outputs into regions and integrates the results from multiple regressors on each regions. Compared with multi-model ensemble, our model is completely end-to-end training. The experimental results demonstrate that our approach achieves the best performance among state-of-the-arts on two public datasets. Hengkai Guo, Guijin Wang, Xinghao Chen 0001, Cairong Zhang, Fei Qiao, Huazhong Yang |
ICIP | 5 |
| 2017 | From "MISSION: IMPOSSIBLE" to mission possible: Fully flexible intelligent contact lens for image classification with analog-to-information processingabstractA prototype of fully flexible intelligent contact lens, which are shown in the impressive action movie series of “MISSION: IMPOSSIBLE”, has become the Possible Mission in this work. Hereon, the system adopts analog-to-information processing method to build a specific Multi-Layer Perceptron network for image classification tasks with flexible devices and circuits, where the information is extracted from raw data of the sensing analog signal directly. Simulated with HSPICE of Level-62 TFT device model, for standard test image data set of MNIST, the classification accuracy of the presented flexible neural network circuit is up to 92.99%; meanwhile, the classification speed is as fast as 10k fps, and the energy consumption is low to only 15.16μJ. Additionally, for the imperfections of flexible devices of larger devices mismatch and process variations, the fault-tolerance of the system has been evaluated as well, which demonstrates the feasibility of the presented methods and lowers the barrier to integrated all kinds of FLEXIBLE Devices into a FULLY FLEXIBLE Systems with sensors, processing parts and even energy harvesting parts, etc., in the future wearable smart terminals. Qin Li 0016, Zheyu Liu, Fei Qiao, Xing Wu 0005, Chaolun Wang, Qi Wei 0001, Huazhong Yang |
ISCAS | 3 |
| 2017 | An 8b 0.8kS/s configurable VCO-based ADC using oxide TFTs with Inkjet printing interconnectionabstractFlexible electronic is a promising technology for flexible and large-area sensing IoT applications, where ADC is a fundamental component This paper proposes a configurable and flexible VCO-based ADC, implemented with Oxide Thin-Film Transistors(TFT) technology. A VCO with four connecting modes is designed to configure the VCO-based ADC working under different power and resolutions. An Inkjet printing interconnection technology is introduced to enable the configurability of ADC, even after all TFT transistors are fabricated. It allows the ADC to be customized for different applications and avoids fabrication failure of devices. Experimental results show that the proposed ADC achieves a 0.8kS/s sampling rate. Its power consumption ranges from 541 to 866uW with ENOB from 3 to 6b. Wenyu Sun, Qinghang Zhao, Fei Qiao, Yongpan Liu, Huazhong Yang |
ISCAS | 3 |
| 2016 | A precision-improved processing architecture of physical computing for energy-efficient SIFT feature extractionabstractA precision-improved processing architecture of physical computing for energy-efficient SIFT feature extraction algorithm has been proposed in this paper. With the novel physical computing technology of active resistor network (PC: ARN), the SIFT algorithm could be processed in analog signal domain without synchronizing clock signals, which means the complex algorithm could be completed within the setup time of the circuit. Especially for the multi-scale Gaussian convolution of SIFT algorithm, an architecture of two-layer 1-dimension PC: ARN has been adopted to compute the horizontal and vertical 1-dimension gaussian filter, in which way higher accuracy can be obtained when compared with the results processed by a 2D active circuit network. A circuit-level simulation with 65nm CMOS technology has been carried out, which shows the energy consumption of gaussian pyramid multi-scale-filtering hardware architecture is about 25.3pJ, where the size of input frame is assigned as 256×256 pixels. Additionally, the average matching ratio in different image pairs is around 80%. Moreover, integrated into the dominating CMOS image sensor with column-parallel readout technology of analog-to-digital convertor, about 20× speedup can be achieved comparing with previous implementations with FPGA, GPU, etc. Fei Qiao, Xinghua Yang, Qi Wei 0001, Huazhong Yang |
ICASSP | 2 |
| 2016 | Approximate Radix-8 Booth Multipliers for Low-Power and High-Performance OperationabstractThe Booth multiplier has been widely used for high performance signed multiplication by encoding and thereby reducing the number of partial products. A multiplier using the radix-$4$(or modified Booth) algorithm is very efficient due to the ease of partial product generation, whereas the radix-$8$Booth multiplier is slow due to the complexity of generating the odd multiples of the multiplicand. In this paper, this issue is alleviated by the application of approximate designs. An approximate$2$-bit adder is deliberately designed for calculating the sum of$1\times$and$2\times$of a binary number. This adder requires a small area, a low power and a short critical path delay. Subsequently, the$2$-bit adder is employed to implement the less significant section of a recoding adder for generating the triple multiplicand with no carry propagation. In the pursuit of a trade-off between accuracy and power consumption, two signed$16\times 16$bit approximate radix-8 Booth multipliers are designed using the approximate recoding adder with and without the truncation of a number of less significant bits in the partial products. The proposed approximate multipliers are faster and more power efficient than the accurate Booth multiplier. The multiplier with 15-bit truncation achieves the best overall performance in terms of hardware and accuracy when compared to other approximate Booth multiplier designs. Finally, the approximate multipliers are applied to the design of a low-pass FIR filter and they show better performance than other approximate Booth multipliers. Honglan Jiang, Jie Han 0001, Fei Qiao, Fabrizio Lombardi |
IEEE Trans. Computers | 3 |
| 2015 | Design methodology for approximate accumulator based on statistical error modelabstractApproximate computing technology has aroused growing interest in circuit and system design for its well-performed trade-off between output quality and performance. Numerous basic circuits and system design methodologies for approximate computing have been proposed. Considering that the existing methodologies for the evaluation of tradeoff between output quality and performance is time-consuming, this paper presents a fast design methodology for approximate accumulator based on statistical error model, in which the inexact multistage speculative adder is adopted and modeled for its advantage of compact error pattern. To validate the proposed methodology, Support Vector Machine(SVM) algorithm is analyzed and mapped to a hardware system composed of inexact and accurate computing circuits. Results show that our time for searching the optimal mapping circuits has been saved by 22.08% than functional-based simulation where the final approximate system design achieves 1.57× speedups with 8.56% accuracy degradation. Xinghua Yang, Fei Qiao, Qi Wei 0001, Huazhong Yang |
ASP-DAC | 3 |
| 2015 | Physical computing circuit with no clock to establish Gaussian pyramid of SIFT algorithmabstractPhysical computing scheme of active resistor network is proposed in this paper to set up a multi-scale Gaussian filter, which is also called Gaussian Pyramid in image signal processing. The analog output signal of each image photodiode could be directly processed by the circuit topology of an active resistor network, which has been pre-designed to meet the requirement of the complicated Gaussian Pyramid in SIFT algorithm. Since it is operated with no clock, the physical computing scheme, with only setup time of the whole processing circuits, is much faster than various current digital realizations. The circuit-level simulation with 65nm CMOS technology has been carried out, which shows the energy consumption of the Gaussian Pyramid processing circuit is around 74.79pJ to filter a frame of 256 × 256 pixel image, and the circuit setting time of the processing is about 138.03ps, considering the parasitic capacitances of each nodes of the circuits. Furthermore, the presented circuit is integrated into a smart CMOS image sensor architecture. With the same processing procedure, the new method of active resistor network could achieve 1.58X speedup when compared with its counterpart of FPGA implementation, in which the column-parallel readout technology of analog-to-digital convertor is used. Fei Qiao, Qi Wei 0001, Huazhong Yang |
ISCAS | 2 |
| 2014 | Design of multi-stage latency adders using detection and sequence-dependence between successive calculationsabstractMulti-stage latency adders based on different prediction schemes have been proved promising to enhance the circuit performance with negligible overhead. This paper presents a novel predictor exploiting both the detection and the sequence-dependence between the successive calculations. The detection of carry-kill pattern of the input data can lower the probability of the operation with multiple clock cycles and the sequence-dependence between the successive calculations is adapted to eliminate redundant cycles. The improved predictors have been inserted into Ripple Carry Adder (RCA) and a multistage latency structure has been setup. Compared with the previous predictors, the proposed one could have the same function with less prediction bits, which results in more energy-efficiency. Simulation results show that 2.41X-3.05X speedups can be achieved than the non-prediction counterpart. Furthermore, a design flow and a method for error control are proposed when applying the adder to approximate computation so that more performance improvement could be obtained after trading off certain precision. Xinghua Yang, Fei Qiao, Qi Wei 0001, Huazhong Yang |
ISCAS | 2 |
| 2013 | Adaptive Dispatching Rule for Semiconductor Wafer Fabrication FacilityabstractUncertainty in semiconductor fabrication facilities (fabs) requires scheduling methods to attain quick real-time responses. They should be well tuned to track the changes of a production environment to obtain better operational performance. This paper presents an adaptive dispatching rule (ADR) whose parameters are determined dynamically by real-time information relevant to scheduling. First, we introduce the workflow of ADR that considers both batch and non-batch processing machines to obtain improved fab-wide performance. It makes use of such information as due date of a job, workload of a machine, and occupation time of a job on a machine. Then, we use a backward propagation neural network (BPNN) and a particle swarm optimization (PSO) algorithm to find the relations between weighting parameters and real-time state information to adapt these parameters dynamically to the environment. Finally, a real fab simulation model is used to demonstrate the proposed method. The simulation results show that ADR with constant weighting parameters outperforms the conventional dispatching rule on average; ADR with changing parameters tracking real-time production information over time is more robust than ADR with constant ones; and further improvements can be obtained by optimizing the weights and threshold values of BPNN with a PSO algorithm. Li Li 0008, Zijin Sun, MengChu Zhou, Fei Qiao |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2013 | A Petri Net and Extended Genetic Algorithm Combined Scheduling Method for Wafer FabricationabstractAs one of the most complicated manufacturing processes, semiconductor manufacturing consists of four steps, wafer sort, wafer fabrication, assembly, and testing. Among them, wafer fabrication is the most costly, complex, and time consuming step. Its operation management and optimization are challenging modeling and scheduling researchers. To address its modeling issue, a hierarchical colored timed Petri net (HCTPN) is proposed, which can be used to describe various states, behavior and substructures of a wafer fabrication system. To address its scheduling issue, intelligent algorithms are introduced to the proposed HCTPN. An extended genetic algorithm (EGA) embedded scheduling strategy over HCTPN is studied to optimize the combination of scheduling policies. The combined approach can conduct more efficient search with better scheduling performance. At last, a real case is presented to illustrate the results. Based on comparing simulation results of different scheduling strategies, the HCTPN and EGA combined scheduling is proved to be valid and efficient. Fei Qiao, Yumin Ma, Li Li 0008, Hong-xia Yu |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2012 | A novel memetic algorithm based on the comprehensive learning PSOabstractA memetic algorithm MCLPSO based on the comprehensive learning PSO (CLPSO) is presented in this study. In MCLPSO, a chaotic local search operator is used and a Simulated Annealing (SA) based local search strategy is developed by combining the cognition-only PSO model with SA. The memetic scheme can enable the stagnant particles which cannot be improved by the comprehensive learning strategy to escape from the local optima and enable some elite particles to give fine-grained local search around the promising regions. The experimental result demonstrates a good performance of MCLPSO in optimizing the multimodal functions compared with some other variants of PSO including CLPSO. Li Li 0008, Fei Qiao, Qidi Wu |
IEEE Congress on Evolutionary Computation | 3 |
| 2012 | Design and implementation of motion compensator in memory reduced HDTV decoder with embedded compression engine
Hongli Gao, Fei Qiao, Huazhong Yang |
Multim. Tools Appl. | 2 |
| 2011 | System-Level Evaluation of Video Processing System Using SimpleScalar-Based Multi-core Processor SimulatorabstractMulti-core processor Simulation Platform is always a very important tool in modern multi-core processor design for the system-level design and evaluation. In this paper, a multi-core processor simulator is proposed by modifying Simple Scalar v3.0to simulate parallelized multi-core programs. Shared memory is used for the communication between different cores, which is the communication network among several different parts of the parallelized program separately. Two simulators are designed for different kinds of usage, one for functional simulation and the other for the simulation of the system with two-level cache. The mismatch of such simulator is less than 10% on average, and the presented simulator is used to evaluate the high-performance video processing systems. Zidong Du, Bingbing Xia, Fei Qiao, Huazhong Yang |
ISADS | 3 |
| 2011 | Low-Power Off-Chip Memory Design for Video Decoder Using Embedded Bus-Invert CodingabstractIn this paper, a simple, efficient, low power off-chip memory design is proposed, which fully exploits the features of DRAM memory and video application, as well as overcomes the drawbacks of algorithm complexity and system modification of embedded compression, which is a popular way to decrease power consumption of the off-chip memory. The integration of the scheme into video decoder will not involve any extra video decoding complexity. It adopts the simple bus-invert encoding scheme. Based on the fact that the power consumption of logic `0' bit is less than that of logic `1', bus-invert encoding scheme is applied to the transferring data between video decoder and off-chip memory. Meanwhile, the features of fault tolerance of human eyes and lossy processing of video decoding application are exploited to solve the extra flag-bit of encoder scheme in off-chip SDARM memory, which has the fixed bit width and is less flexible than on-chip SRAM. This scheme is integrated into MPEG-2 decoder system. The experiment results show that this scheme can archive 20%-35% reduction in power consumption of logic `1' bit, and the objective quality of image has about 1.5db PSNR improvement on average. Ni Zhou, Fei Qiao, Huazhong Yang, Hui Wang 0004 |
ISADS | 2 |
| 2008 | Implementation of low-swing differential interface circuits for high-speed on-chip asynchronous interconnection
Fei Qiao, Huazhong Yang, Hui Wang 0004 |
Sci. China Ser. F Inf. Sci. | 1 |
| 2007 | A Lot Dispatching Strategy Integrating WIP Management and Wafer Start ControlabstractCompound priority dispatching (CPD) is a new dispatching strategy for semiconductor wafer fabrication. It takes into account both work-in-progress (WIP) management and wafer start control. The compound priority of wafers is calculated based upon fab wafer start status, the current processing step, and the amount of WIP in the current, the upstream, and the downstream steps. Simulation results demonstrate that, for a given fab model, CPD can reduce the mean total queue time (MTQT) by 50% and increase the throughput rate by 20% compared with first- in-first-out (FIFO) and shortest remaining processing time (SRPT) scheduling (dispatching) strategies. Zuntong Wang, Qidi Wu, Fei Qiao |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2004 | The new method of dynamic scheduling in semiconductor fabrication lineabstractSemiconductor manufacturing is considered as one of the most complex manufacturing process. Due to its large-scale, uncertainty, reentrance and mixed processing, it needs true dynamic scheduling method urgently. On the analysis of the self-organizing activities of ant colony system, the pheromone based dynamic scheduling rule is presented. The rule is simulated and the result analysis is given by comparing with other heuristic scheduling methods, which show that this method can optimize multi objectives of the semiconductor fabrication line and has a good perspective. Jiang Hua, Li Li 0008, Fei Qiao, Qidi Wu |
ICARCV | 3 |
| 2004 | The research on dispatching rule for improving on-time delivery for semiconductor wafer fababstractSemiconductor wafer fab cries for really dynamic dispatching approach, due to its high uncertainty, re-entrance and large-scale. A dispatching rule for improving on-time delivery for semiconductor wafer fab (ODDR) is proposed, which considers the dispatching of bottleneck machines, not-bottleneck machines, batching machines and hot lots. As a result, it can distinctly improve on-time delivery without decreasing the throughput and increasing the cycle time. Finally, a simplistic model, but with essential characteristics of semiconductor wafer fab, is used to compare ODDR with FIFO, EDD and CR. It can be seen from the experimental results that ODDR is prior to FIFO, EDD and CR on throughput, cycle time and on-time delivery with better performance, especially for on-time delivery. Li Li 0008, Fei Qiao, Qidi Wu |
ICARCV | 2 |