EDBT 2026 Demo / reviewers in the wild / expert
Cong Shi 0003
dblp:07/6946-3
· DBLP profile ↗
19ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0003-0040-4411ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AM-CIM: Approximate Memory Based Near Sensor Compute-in-Memory Architecture for Keyword SpottingabstractCompute-In-Memory (CIM) has emerged as a promising solution to address the von-Neumann bottleneck, making it a key technology for intelligent computing in edge IoT devices, particularly for real-time applications like keyword spotting (KWS). However, traditional CIM architectures face challenges such as high resource consumption, especially in data conversion, which can significantly impact chip area and energy efficiency. To address these challenges, this work proposes a computational CIM architecture utilizing multilevel analog memory, named AM-CIM, tailored for near-sensor (NS) computation of real-time KWS applications. Additionally, approximate memory technology is integrated into the AM-CIM architecture, employing data resilience scheduling for analog memory which contributes to significant reductions in hardware overhead. This integration facilitates a hardware-software co-design approach. To deploy KWS tasks in AM-CIM, a gated recurrent unit (GRU) network, referred to as MAC-GRU, is implemented. By employing Mel-energy as the input feature at the near-sensor end, the system achieves a 93.13% reduction in feature extraction power consumption. Evaluation results based on TSMC 180-nm technology demonstrate that the AM-CIM architecture achieves an accuracy of 88.51% for 10-keyword classification with a power consumption of$546~\mu W$, while reducing analog memory area by 43.32%. Xiaotao Jia, Guangcai Yuan, Jianyi Yu, Cong Shi 0003, Qi Wei 0001, Youguang Zhang, Weisheng Zhao 0001, Fei Qiao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | A Real-Time 2D/3D Perception Visual Vector Processor for 1920 × 1080 High-Resolution High-Speed Intelligent Vision ChipsabstractEdge computing of reliable multimodal (2D RGB/3D RGB-Depth) data has a wide range of applications. However, many of currently reported visual processors cannot flexibly handle multimodal data, e.g., the visual streams of RGB-Depth data. The key challenge exists that these prior visual processors do not come with efficient and unified instruction set architecture (ISA) for both conventional and intelligent cognition on the 2D/3D multimodal sensory data. To fill such a gap, this paper proposes a programmable intelligent visual vector processor compatible with multimodal 2D/3D visual data processing ($1920\times 1080$-pixel resolution). The processor consists of a reconfigurable processing element (PE) array, a memory access network flexibly configurable to be fine- or coarse-grained, and a high throughput I/O interface. The vectorial PE array with neighbor PE access increases the data reuse rate and parallel computation efficiency, and can implement both convolutional neural networks (CNNs) and conventional image processing algorithms. The proposed ISA is customized and optimally tailored targeting 2D/3D image processing from RGB/Time-of-Flight(ToF) raw data to intelligent inference results. The chip is fabricated in a 55-nm CMOS process. The experimental results showed that the area efficiency, peak performance, and peak throughput of our chip attained as high as 14.41GOPS/mm2, 409.6GOPS, and 9.6Gbps at 200MHz, respectively. The measured processing speeds of this chip on ToF depth reconstruction is 87fps ($480\times 270$) or 31 fps($1920\times 1080$),on 3D object classification is 219fps ($256\times 256$), and on CNN-based 2D object tracking is 36fps ($256\times 256$). Siyuan Wei, Lei Kang 0006, Xuemin Zheng, Mingxin Zhao, Mengmeng Xu 0005, Xuanzhe Xu, Runjiang Dou, Shuangming Yu, Xu Yang 0017, Jian Liu 0021, Cong Shi 0003, Nanjian Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 13 |
| 2023 | Low-cost real-time VLSI system for high-accuracy optical flow estimation using biological motion features and random forests
Cong Shi 0003, Junxian He, Shrinivas J. Pundlik, Xichuan Zhou, Nanjian Wu, Gang Luo 0003 |
Sci. China Inf. Sci. | 1 |
| 2023 | Network Pruning for Bit-Serial AcceleratorsabstractBit-serial architectures (BSAs) are becoming increasingly popular in low-power neural network processor (NNP) designs for edge scenarios. However, the performance and energy efficiency of state-of-the-art BSA NNPs heavily depends on both the proportion and distribution of ineffectual weight bits in neural networks (NNs). To boost the performance of typical BSA accelerators, we present Bit-Pruner, a software approach to learn BSA-favored NNs without resorting to hardware modifications. Bit-Pruner not only progressively prunes but also restructures the nonzero bits in weights so that the number of nonzero bits in the model can be reduced and the corresponding computing can be load-balanced to suit the target BSA accelerators. On top of Bit-Pruner, we further propose a Pareto frontier optimization algorithm to adjust the bit-pruning rate across network layers and fulfill diverse NN processing requirements in terms of performance and accuracy for various edge scenarios. However, an aggressive Bit-Pruner can lead to nontrivial accuracy loss, especially for lightweight NNs and complex tasks. To this end, the alternating direction method of multipliers (ADMMs) is adapted to the retraining phase in Bit-Pruner to smooth the abrupt disturbance due to bit-pruning and enhance the resulting model accuracy. According to the experiments, Bit-Pruner increases the bit-sparsity up to 94.4% with negligible accuracy degradation and achieves an optimized tradeoff between NN accuracy and energy efficiency even under very-aggressive performance constraints. When pruned models are deployed onto typical BSA accelerators, the average performance is$2.1\times $and$1.6\times $higher than the baseline networks without pruning and those with classical weight pruning, respectively. Xiandong Zhao, Ying Wang 0001, Cheng Liu 0008, Cong Shi 0003, Kaijie Tu, Lei Zhang 0008 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | An 8-T Processing-in-Memory SRAM Cell-Based Pixel-Parallel Array Processor for Vision ChipsabstractVision chip is a high-speed image processing device, featuring a massively-parallel pixel-level processing element (PE) array to boost pixel processing speed. However, the collocated processing unit and fine-grained data memory unit inside each PE impose a huge requirement on memory access bandwidth as well as big area and energy consumption. To overcome this bottleneck, this paper proposes a full custom 8T SRAM-based Processing-in-Memory (PIM) architecture together with a multiplexer-based arithmetic-logic unit (mux-based ALU) to realize pixel-parallel array processor for energy-efficient vision chips. The proposed PIM architecture is constructed by embroidering each dual-port 8T SRAM cell with mux-based ALU, so as to form a PIM PE array. Each PIM PE holds a 130-bit 8T SRAM cell block embedding in-memory logic functions, of which 128-bit 8T SRAM cells serve as the PE memory, and 2-bit 8T SRAM cells act as a buffer register in the PE. A full custom physical layout of a$128\times128$prototyping PIM PE array is designed and evaluated using a 65 nm CMOS technology. The simulation results demonstrate that our proposed PIM PE architecture could operate under a 200 MHz clock frequency with a 1.0 V power supply, and reach a high energy efficiency of 512 GOPS/W and a high area efficiency of 29 GOPS/mm2. Leyi Chen, Cong Shi 0003, Junxian He, Jianyi Yu, Haibing Wang, Nanjian Wu, Min Tian 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | A Lightweight Spiking GAN Model for Memristor-centric Silicon Circuit with On-chip Reinforcement Adversarial LearningabstractAs a powerful generative model, Generative Adversarial Network (GAN) is widely studied to automatically generate high-quality new data to greatly enhances the capabilities of artificial intelligence (AI) technology. However, the unique training process of GAN comes at a very high computational complexity and high cost of memory accesses. In this work, a memristor-based spiking-GAN neuromorphic hardware system is proposed to address the challenges. Both the generator and discriminator of GAN are in the form of spiking neural network (SNN) to improve the computational performance, and the memristor synapse circuit with 1 memristor and 4 transistors (1M4T) is proposed as Computing in Memory (CIM) to avoid the cost of memory accesses. The reinforcement learning rule (i.e., reward-modulated spike-timing dependent plasticity, or R-STDP) is used to train both discriminator and generator networks, with a new backpropagation method for the reward/punishment signal. Tests on the MNIST and Fashion-MNIST datasets showed that the proposed GAN can efficiently generate data samples. The results demonstrate the great potential of this memristor-based spiking-GAN for high-speed energy-efficient data augmentations. Min Tian 0003, Haibing Wang, Jianyi Yu, Cong Shi 0003 |
ISCAS | 6 |
| 2021 | Optimizing Information Theory Based Bitwise Bottlenecks for Efficient Mixed-Precision Activation QuantizationabstractRecent researches on information theory shed new light on the continuous attempts to open the black box of neural signal encoding. Inspired by the problem of lossy signal compression for wireless communication, this paper presents a Bitwise Bottleneck approach for quantizing and encoding neural network activations. Based on the rate-distortion theory, the Bitwise Bottleneck attempts to determine the most significant bits in activation representation by assigning and approximating the sparse coefficients associated with different bits. Given the constraint of a limited average code rate, the bottleneck minimizes the distortion for optimal activation quantization in a flexible layer-by-layer manner. Experiments over ImageNet and other datasets show that, by minimizing the quantization distortion of each layer, the neural network with bottlenecks achieves the state-of-the-art accuracy with low-precision activation. Meanwhile, by reducing the code rate, the proposed method can improve the memory and computational efficiency by over six times compared with the deep neural network with standard single-precision representation. The source code is available on GitHub: https://github.com/CQUlearningsystemgroup/BitwiseBottleneck. Xichuan Zhou, Cong Shi 0003, Haijun Liu 0001 |
AAAI | 3 |
| 2021 | A Heterogeneous Spiking Neural Network for Computationally Efficient Face RecognitionabstractComputational efficiency is critical to many mobile and always-on face recognition applications. To this end, a heterogeneous spiking neural network (SNN) is proposed for face recognition. To obtain high recognition accuracy at minimal computational overheads, the heterogeneous SNN consists of an encoding subnet for sparse image feature encoding and classification subnet for feature classification. The experimental results suggest that the proposed heterogeneous algorithm can achieve high recognition accuracy on small datasets of human face samples with labeled identities at a high computational efficiency with very low neuronal activities. The proposed SNN is promising for low-cost mobile or always-on systems with strictly constrained resource and energy budgets. Xichuan Zhou, Zhenghua Zhou, Zhengqing Zhong, Jianyi Yu, Tengxiao Wang, Min Tian 0003, Cong Shi 0003 |
ISCAS | 8 |
| 2021 | CompSNN: A lightweight spiking neural network based on spatiotemporally compressive spike features
Tengxiao Wang, Cong Shi 0003, Xichuan Zhou, Yingcheng Lin, Junxian He, Ping Gan, Ping Li 0042, Ying Wang 0001, Nanjian Wu, Gang Luo 0003 |
Neurocomputing | 2 |
| 2021 | Hapke Data Augmentation for Deep Learning-Based Hyperspectral Data Analysis With Limited SamplesabstractThe emerging technology of deep neural networks has been proven to be successful for hyperspectral image analysis. However, it is still a great challenge to apply the deep learning method for quantitatively retrieving mineralogical composition, because typical deep neural networks generally require thousands of labeled samples for training, while only a few mineral samples can be acquired and examined for quantitative examination in practice. To address this challenge, this letter proposes a training data augmentation approach which incorporates the prior-knowledge of hyperspectral reflectance characteristics using the classic Hapke equations. Experiments over both laboratory and airborne hyperspectral remote sensing data show that the proposed method outperforms the widely used approaches for quantitative mineral analysis. Fangyuan Ge, Yingjun Zhao, Ming Li 0082, Cong Shi 0003, Dong Li 0007, Xichuan Zhou |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2021 | An Edge 3D CNN Accelerator for Low-Power Activity Recognitionabstract3D convolutional neural networks (CNNs) are gaining increasing popularity in the area of video-based action/activity analysis. Compared to 2D convolutions that share the filters in a 2D spatial domain, 3D convolutions further reuse filters in the temporal dimension to capture temporal-domain features in the video. How to exploit the data locality in the temporal dimension directly impacts the energy efficiency of specialized architectures for 3D CNN inference. Prior works on specialized 3D-CNN accelerators employ additional on-chip memories and multicluster architecture to reuse data among the process element (PE) arrays, which is very expensive for low-power chip implementation. Instead of harvesting in-memory data locality, we propose the architecture of systolic cube to exploit the spatial and temporal localities in 3D CNNs, which moves the reusable data in-between PEs connected via a 3D-cube network-on-chip. Furthermore, due to the existence of visual feature reappearance in the temporal domain, there exists a considerable portion of repetitive pixels and activations among the feature maps captured at adjacent time slots. To eliminate such temporal redundancy in 3D CNNs, the proposed accelerator architecture is equipped with a redundancy detection and elimination mechanism, capable of skipping the computations with the same activations and parameters when reusing the convolutional filters along the temporal dimension. In our evaluation, the experimental results show that the systolic-cube architecture contributes to a considerable energy-efficiency boost for state-of-the-art activity-recognition benchmarks and datasets. Ying Wang 0001, Yongchen Wang, Cong Shi 0003, Long Cheng 0003, Huawei Li 0001, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | BitPruner: Network Pruning for Bit-serial AcceleratorsabstractBit-serial architectures (BSAs) are becoming increasingly popular in low power neural network processor (NNP) design. However, the performance and efficiency of state-of-the-art BSA NNPs are heavily depending on the distribution of ineffectual weight-bits of the running neural network. To boost the efficiency of third-party BSA accelerators, this work presents Bit-Pruner, a software approach to learn BSA-favored neural networks without resorting to hardware modifications. The techniques proposed in this work not only progressively prune but also structure the non-zero bits in weights, so that the number of zero-bits in the model can be increased and also load-balanced to suit the architecture of the target BSA accelerators. According to our experiments on a set of representative neural networks, Bit-Pruner increases the bit-sparsity up to 94.4% with negligible accuracy degradation. When the bit-pruned models are deployed onto typical BSA accelerators, the average performance is 2.1X and 1.5X higher than the baselines running non-pruned and weight-pruned networks, respectively. Xiandong Zhao, Ying Wang 0001, Cheng Liu 0008, Cong Shi 0003, Kaijie Tu, Lei Zhang 0008 |
DAC | 4 |
| 2020 | MoNet3D: Towards Accurate Monocular 3D Object Localization in Real TimeabstractMonocular multi-object detection and localization in 3D space has been proven to be a challenging task. The MoNet3D algorithm is a novel and effective framework that can predict the 3D position of each object in a monocular image, and draw a 3D bounding box on each object. The MoNet3D method incorporates the prior knowledge of spatial geometric correlation of neighboring objects into the deep neural network training process, in order to improve the accuracy of 3D object localization. Experiments over the KITTI data set show that the accuracy of predicting the depth and horizontal coordinate of the object in 3D space can reach 96.25% and 94.74%, respectively. Meanwhile, the method can realize the real-time image processing capability of 27.85 FPS. Our code is publicly available at https://github.com/CQUlearningsystemgroup/YicongPeng Xichuan Zhou, Yicong Peng, Chunqiao Long, Fengbo Ren, Cong Shi 0003 |
ICML | 5 |
| 2019 | Systolic Cube: A Spatial 3D CNN Accelerator Architecture for Low Power Video Analysisabstract3D convolutional neural networks (CNN) are gaining popularity in action/activity analysis. Compared to 2D convolutions that share the filters in 2D spatial domain, 3D convolutions further reuse filters in the temporal dimension to capture time-domain features. Prior works on specialized 3D-CNN accelerators employ additional on-chip memories and multi-cluster architecture to reuse data among the process element (PE) arrays, which is too expensive for low-power chips. Instead of harvesting in-memory locality, we propose a 3D systolic cube architecture to exploit the spatial-and-temporal localities of 3D CNNs, which moves the reusable data in-between PEs connected via a 3D-cube Network-on-Chip. Evaluation shows that systolic-cube contributes to considerable energy-efficiency boost for activity-recognition benchmarks. Yongchen Wang, Ying Wang 0001, Huawei Li 0001, Cong Shi 0003, Xiaowei Li 0001 |
DAC | 4 |
| 2018 | A Compact VLSI System for Bio-Inspired Visual Motion EstimationabstractThis paper proposes a bio-inspired visual motion estimation algorithm based on motion energy, along with its compact very-large-scale integration (VLSI) architecture using low-cost embedded systems. The algorithm mimics motion perception functions of retina, V1, and MT neurons in a primate visual system. It involves operations of ternary edge extraction, spatiotemporal filtering, motion energy extraction, and velocity integration. Moreover, we propose the concept of confidence map to indicate the reliability of estimation results on each probing location. Our algorithm involves only additions and multiplications during runtime, which is suitable for low-cost hardware implementation. The proposed VLSI architecture employs multiple (frame, pixel, and operation) levels of pipeline and massively parallel processing arrays to boost the system performance. The array unit circuits are optimized to minimize hardware resource consumption. We have prototyped the proposed architecture on a low-cost field-programmable gate array platform (Zynq 7020) running at 53-MHz clock frequency. It achieved 30-frame/s real-time performance for velocity estimation on 160 × 120 probing locations. A comprehensive evaluation experiment showed that the estimated velocity by our prototype has relatively small errors (average endpoint error < 0.5 pixel and angular error < 10°) for most motion cases. Cong Shi 0003, Gang Luo 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | A low power global shutter pixel with extended FD voltage swing range for large format high speed CMOS image sensor
Yangfan Zhou 0001, Zhongxiang Cao, Quanliang Li, Cong Shi 0003, Runjiang Dou, Jian Liu 0021, Nanjian Wu |
Sci. China Inf. Sci. | 5 |
| 2014 | A massively parallel keypoint detection and description (MP-KDD) algorithm for high-speed vision chip
Cong Shi 0003, Jie Yang 0033, Nanjian Wu, Zhihua Wang 0001 |
Sci. China Inf. Sci. | 1 |
| 2014 | A high speed multi-level-parallel array processor for vision chips
Cong Shi 0003, Jie Yang 0033, Nanjian Wu, Zhihua Wang 0001 |
Sci. China Inf. Sci. | 1 |
| 2014 | A high speed 1000 fps CMOS image sensor with low noise global shutter pixels
Yangfan Zhou 0001, Zhongxiang Cao, Quanliang Li, Cong Shi 0003, Nanjian Wu |
Sci. China Inf. Sci. | 5 |