VLDB 2026 Research / reviewers in the wild / expert
Shan Cao 0001
dblp:70/4512-1
· DBLP profile ↗
41ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0003-3713-8671ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 7 first-author · 12 since 2021Computer networks · 9 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Late Breaking Results: A Power-Efficient RISC-V Baseband System-on-Chip for Multi-Standard Integrated Sensing and CommunicationsabstractWe present Ishtar, a power-efficient RISC-V baseband system-on-chip (SoC) tailored for multi-standard integrated sensing and communications (ISAC) in low-altitude wireless networks (LAWNs). Ishtar integrates a hierarchical scheduling scheme and a system-level power-gating architecture that dynamically controls power domains to balance performance and energy efficiency. It supports dynamic task scheduling across heterogeneous protocols using a domain-specific, graph-based representation. Implemented in 40 nm technology and running at 300 MHz, Ishtar achieves better normalized efficiency than state-of-the-art SDR SoCs, delivering real-time multi-standard sniffing under stringent power and area constraints. Limin Jiang, Yi Shi 0004, Yihao Shen, Yintao Liu 0001, Siyi Xu, Qingyu Deng, Shan Cao 0001, Zhiyuan Jiang, Sheng Zhou 0001 |
DATE | 9 |
| 2026 | Live Demonstration: A Flexible and Upgradable GNSS Receiver on Venus Architecture
Yule Jiao, Shiji Ruan, Limin Jiang, Zhiyuan Jiang, Shan Cao 0001 |
ISCAS | 5 |
| 2026 | Efficient network compression via gradient-score aware pruning
Qiuying Li, Zhixiang Chen 0003, Yu Li 0051, Zhiyuan Jiang, Shan Cao 0001 |
Neurocomputing | 5 |
| 2026 | Venusian: Rapid Wireless Baseband Validation via High-Level Programming and FPGA-Based RISC-V Accelerator Co-Design
Limin Jiang, Yi Shi 0004, Yihao Shen, Yintao Liu 0001, Shan Cao 0001, Zhiyuan Jiang, Sheng Zhou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | A Hierarchical Dataflow-Driven Heterogeneous Architecture for Wireless Baseband ProcessingabstractWireless baseband processing (WBP) is a key element of wireless communications, with a series of signal processing modules to improve data throughput and counter channel fading. Conventional hardware solutions, such as digital signal processors (DSPs) and more recently, graphic processing units (GPUs), provide various degrees of parallelism, yet they both fail to take into account the cyclical and consecutive character of WBP. Furthermore, the large amount of data in WBPs cannot be processed quickly in symmetric multiprocessors (SMPs) due to the unpredictability of memory latency. To address this issue, we propose a hierarchical dataflow-driven architecture to accelerate WBP. A pack-and-ship approach is presented under a non-uniform memory access (NUMA) architecture to allow the subordinate tiles to operate in a bundled access and execute manner. We also propose a multi-level dataflow model and the related scheduling scheme to manage and allocate the heterogeneous hardware resources. Experiment results demonstrate that our prototype achieves 2× and 2.3× speedup in terms of normalized throughput and single-tile clock cycles compared with GPU and DSP counterparts in several critical WBP benchmarks. Additionally, a link-level throughput of 288 Mbps can be achieved with a 45-core configuration. Limin Jiang, Yi Shi 0004, Yintao Liu 0001, Qingyu Deng, Siyi Xu, Yihao Shen, Fangfang Ye, Shan Cao 0001, Zhiyuan Jiang |
ASP-DAC | 8 |
| 2025 | Packetized Pipelined Pillar Feature Net Accelerator for LiDAR 3D Object DetectionabstractImplementing LiDAR-based 3D object detection algorithms in practical autonomous driving situations presents a significant challenge. In current research algorithms, the inherent sparsity and randomness of point cloud data necessitate significant memory usage and frequent data read/write operations during preprocessing. Such demands are not well-suited for terminal devices with stringent real-time requirements and constrained resources. In this paper, we present a packetized processing Pillar Feature Net accelerator for LiDAR 3D object detection. By integrating voxelization and feature extraction into a pipelined architecture, the proposed accelerator significantly reduces the storage requirements for point cloud data and enhances the speed of feature extraction and pseudo-image generation. Experimental results indicate that the proposed method improves the computational throughput from point cloud data to pseudo-image generation by 1.2 times and eliminates the need for off-chip memory access during preprocessing. Qingyu Deng, Xinyu Chen 0007, Wei Zhang 0001, Beining Zhao 0001, Yuhang Gu, Shan Cao 0001, Zhiyuan Jiang |
ISCAS | 6 |
| 2025 | MPQA: Mixed-Precision Quantization Accelerator for CNN InferenceabstractMixed-precision quantization CNNs have become crucial for edge vision algorithms. While model quantization techniques have improved inference efficiency, existing approaches have not fully addressed data coherence in coarse-grained parallelism or adaptation to various 2D data sizes. This study presents a novel CNN accelerator equipped with mixed-precision multiplier and feature linking mechanisms to enhance mixed-precision inference on edge NPUs. The accelerator incorporates reconfigurable multi-precision multiplication support in MAC units, enabling 8-bit/4-bit mixed-precision quantization. A feature linking mechanism is developed to overcome feature map size limitations, achieving scalable processing dimensions. The accelerator design, implemented on a Xilinx Zynq UltraScale+ MPSoC FPGA platform, demonstrates improved hardware performance and enhanced inference speeds for edge device CNN models such as Tiny-YOLOv3 and ResNet18. This work provides a significant advancement in deploying mixed-quantization CNNs on edge NPUs, enhancing the adaptability and performance of edge computing devices for complex neural network models. Beining Zhao 0001, Yu Li 0051, Jiahao Zuo, Wei Zhang 0001, Xinyu Chen 0007, Shan Cao 0001, Zhiyuan Jiang |
ISCAS | 6 |
| 2025 | Zoozve: A Strip-Mining-Free RISC-V Vector Extension with Arbitrary Register Grouping Compilation Support (WIP)abstractVector processing is crucial for boosting processor performance and efficiency, particularly with data-parallel tasks. The RISC-V ”V” Vector Extension (RVV) enhances algorithm efficiency by supporting vector registers of dynamic sizes and their grouping. Nevertheless, for very long vectors, the static number of RVV vector registers and its power-of-two grouping can lead to performance restrictions. To counteract this limitation, this work introduces Zoozve, a RISC-V vector instruction extension that eliminates the need for strip-mining. Zoozve allows for flexible vector register length and count configurations to boost data computation parallelism. With a data-adaptive register allocation approach, Zoozve permits any register groupings and accurately aligns vector lengths, cutting down register overhead and alleviating performance declines from strip-mining. Additionally, the paper details Zoozve’s compiler and hardware implementations using LLVM and SystemVerilog. Initial results indicate Zoozve yields a minimum 10.10× reduction in dynamic instruction count for fast Fourier transform (FFT), with a mere 5.2% increase in overall silicon area. Siyi Xu, Limin Jiang, Yintao Liu 0001, Yihao Shen, Yi Shi 0004, Shan Cao 0001, Zhiyuan Jiang |
LCTES | 6 |
| 2025 | Near-Sensor LiDAR and Visual Feature Extraction and Communication for Low-Latency Roadside Cooperative PerceptionabstractAutonomous driving technologies are swiftly evolving, characterized by two main strategies: Single-Vehicle Autonomous Driving (SVAD) and Vehicle-Infrastructure Cooperative Autonomous Driving (VICAD). SVAD depends entirely on the vehicle’s internal sensors and processing capabilities, whereas VICAD benefits from a synergistic network combining roadside infrastructure, connected vehicles, and cloud services to boost safety and efficiency. Nevertheless, VICAD encounters challenges with high-bandwidth data transmission and perception latency. To mitigate these concerns, we introduce an innovative intelligent roadside unit (I-RSU) platform integrating perception, computing, and communication into one cohesive system. The platform features dual neural processing units (NPUs) for the effective extraction of images and LiDAR features, alongside a C-V2X communication module, all realized on a Field-Programmable Gate Array (FPGA). This setup minimizes latency and expenses by enabling computation near the sensors and facilitating selective data transmission. Our system also supports multi-modal fusion, enhancing overall perception and safety. Through extensive real-world trials and simulations, our system demonstrates a substantial reduction in end-to-end latency, providing a scalable solution for VICAD scenarios. Wei Zhang 0388, Yuhang Gu, Beining Zhao 0001, Qingyu Deng, Xinyu Chen 0007, Yi Shi 0004, Limin Jiang, Shan Cao 0001, Zhiyuan Jiang, Ruiqing Mao, Sheng Zhou 0001 |
IEEE Internet Things J. | 8 |
| 2025 | A Heterogeneous CNN Compilation Framework for RISC-V CPU and NPU Integration Based on ONNX-MLIRabstractWith the continuous advancement of convolutional neural networks (CNNs), many neural network processing units (NPU) have emerged in recent years. NPUs offer advantages such as improved energy efficiency and low latency compared to traditional processors. However, NPUs often struggle to keep pace with the growing complexity of algorithmic models, which limits the implementation of certain AI applications. Central processing units (CPU) are known for their versatility, but are computationally inefficient. Combining the strengths of both architectures could present a viable solution to these challenges. Despite this potential, there is currently insufficient academic focus on CPU/NPU co-computing, and few architectures or toolchains have been developed to support such integration. In this paper, we propose a RISC-V CPU/NPU co-computing-based compilation framework to solve the problem of mapping from algorithms to heterogeneous architectures. To bridge the computational gap between heterogeneous platforms, within ONNX-MLIR, we developed three phases to support operator packing, algorithm adaption, and data coordination. To integrate heterogeneous compilation platforms, we introduce four additional functions to handle task scheduling, weight quantization, weight rearrangement, and memory management. In particular: (1) to optimize the transition overhead, an operator packing technique is introduced by extending the ONNX-MLIR framework with custom operators, which enhances NPU efficiency by reducing the number of transitions between operations. (2) At the aspect of task scheduling, a three-mode task scheduler is developed to allow users to customize task distribution for correctness verification, performance optimization, or detailed analysis on heterogeneous platforms, which demonstrates the versatility of our framework. (3) In terms of memory management, an optimized scheme for CPU/NPU architectures is proposed to ensure accurate data access. This memory management employs a dual-end approach with self-checks and a release-after-use mechanism to further improve efficiency. Experimental results demonstrate that the framework effectively generates instructions for various CNNs, significantly enhancing the efficiency and performance of heterogeneous architectures. It achieves up to a 7.06× reduction in code density and up to a 5.58× improvement in memory usage, narrowing the computational gap between CPUs and NPUs. Shan Cao 0001, Meiling Yang, Yintao Liu 0001, Yu Li 0051, Beining Zhao 0001, Xinyu Chen 0007, Zhiyuan Jiang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | A Critical-Set-Based Multi-Bit Successive Cancellation List Decoder for Polar Codes: Algorithm and ImplementationabstractWith the evolution of wireless communication systems, there is a growing demand for high reliability and low latency in channel coding, particularly in 5G and beyond wireless systems used in applications such as autonomous driving and remote medical services. For the decoding of polar codes, the multi-bit successive cancellation list (MSCL) decoding technique was recently introduced to decrease the decoding latency by decoding several short inner codes in parallel, which preserves high reliability compared to the conventional successive cancellation list (SCL) decoding. However, as parallelism increases, the complexity of the decoding path sorting also increases significantly, which makes it resource-intensive for hardware implementation. To address this issue, this paper proposes a configurable critical-set-based multi-bit successive cancellation list (CS-MSCL) decoding algorithm, which first introduces critical sets to the MSCL decoding for the optimization of path pruning. Subsequently, an enhanced CS-MSCL algorithm is introduced for large list-size MSCL decoding, which can boost the error correction performance. Then, an area-efficient decoding architecture is introduced, which supports the cyclic redundancy check (CRC) and the CS-MSCL decoding compatible with the 5G standard. The proposed decoder is implemented in SMIC 40 nm CMOS technology with a parallelism degree of 8, which has a peak area efficiency of$4.64~\mathrm {Gbps/mm^{2}}$for list size 4 and$2.01~\mathrm {Gbps/mm^{2}}$for list size 8. Compared to state-of-the-art SCL-based decoders, the normalized area efficiency is improved by 7.16% and 17.54% for list sizes 4 and 8, respectively. Shan Cao 0001, Limin Jiang, Zhiyuan Jiang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | Unlimited Vector Processing for Wireless Baseband Based on RISC-V ExtensionabstractWireless baseband processing (WBP) serves as an ideal scenario for utilizing vector processing, which excels in managing data-parallel operations due to its parallel structure. However, conventional vector architectures face certain constraints such as limited vector register sizes, reliance on power-of-two vector length (VL) multipliers, and vector permutation capabilities tied to specific architectures. To address these challenges, we have introduced an instruction set extension (ISE) based on RISC-V known as unlimited vector processing (UVP). This extension enhances both the flexibility and efficiency of vector computations. UVP employs a novel programming model that supports non-power-of-two register groupings (RGs) and hardware strip mining, thus enabling smooth handling of vectors of varying lengths while reducing the software strip-mining burden. Vector instructions are categorized into symmetric and asymmetric classes, complemented by specialized load/store strategies to optimize execution. Moreover, we present a hardware implementation of UVP featuring sophisticated hazard detection mechanisms, optimized pipelines for symmetric tasks such as fixed-point multiplication and division, and a robust permutation engine for effective asymmetric operations. Comprehensive evaluations demonstrate that UVP significantly enhances performance, achieving up to$3.0\times $and$2.1\times $speedups in matrix multiplication and fast Fourier transform (FFT) tasks, respectively, when measured against lane-based vector architectures. Our synthesized register transfer level (RTL) for a 16-lane configuration using SMIC 40-nm technology spans 0.94 mm2and achieves an area efficiency of 21.2 GOPS/mm2. Limin Jiang, Yi Shi 0004, Yihao Shen, Shan Cao 0001, Zhiyuan Jiang, Sheng Zhou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | Hybrid-Grained Pruning and Hardware Acceleration for Convolutional Neural NetworksabstractThroughout various convolutional neural network (CNN) models, the sparsity increases as the network deepens, which poses significant potential to model compression and hardware acceleration. In this paper, a dual-factor hybrid-grained pruning method is introduced to make a good balance between model compression and accuracy preservation. The pro-posed pruning method combines hardware-friendly unstructured vector-level pruning with structured filter-level pruning to explore multiple grains of sparsity in CNNs. The architecture of the corresponding hardware accelerator is then proposed based on the row-based convolution dataflow, which could fully utilize the hybrid sparsity to accelerate CNN processing. Experimental results demonstrate that the proposed method increases the compression rate by 1.08× while causing 0.21% accuracy loss compared to the state-of-the-art filter pruning method in VGG16, and 2.39% hardware resource increase compared to the accelerator without sparsity optimization. Yu Li 0051, Shan Cao 0001, Beining Zhao 0001, Wei Zhang 0001, Zhiyuan Jiang |
ISCAS | 2 |
| 2024 | Dynamically Configurable FIR Filters Based on Serial MACs and Systolic ArraysabstractFIR (Finite Impulse Response) filters are widely used in digital communication systems, digital image processing, and many other fields. A great deal of research has been done on the flexible configuration of FIR filters, particularly on the dynamic adjustment of coefficients and orders. Existing FIR filter structures can be configured to a higher-order filter for a lower-order use, leading to low hardware utilization. This paper presents a dynamically configurable architecture for FIR filters based on a novel architecture with systolic arrays and serial multiply accumulators (MACs). This design can be configured to a higher-order filter or to several independent lower-order filters, thus increasing utilization. We demonstrate a 2048-order FIR filter that can be configured to a minimum of 16 orders and a maximum of 128 channels using only 256 multipliers and adders. Bo Ruan, Limin Jiang, Shan Cao 0001, Zhiyuan Jiang |
ISCAS | 3 |
| 2023 | Parallel Computing for Energy-Efficient Baseband Processing in O-RAN: Synchronization and OFDM Implementation Based on SPMDabstractOpen radio access network (O-RAN) is considered as a viable method for reducing the cost and enhancing the energy efficiency of cellular networks, due to its native incorporation of intelligence and open interfaces. However, the processing delay of software-based wireless protocol stacks has hindered its development. This paper presents the implementation of parallel computing acceleration for an LTE baseband system based on single program multiple data (SPMD) methodology, and proposes detailed optimization strategies for the time-consuming synchronization and OFDM modulation modules in the system. Experiment results based on the implicit SPMD program compiler (ISPC) show that continuous memory access has a significant impact on the final acceleration effect. Moreover, the processing speed of software-based physical layer can be increased up to 10–30 times through SPMD. Yihao Shen, Shan Cao 0001, Zhiyuan Jiang, Sheng Zhou 0001 |
GLOBECOM | 3 |
| 2021 | A Semi-Folded Decoding Architecture for Flexible Codeword Length Configuration of Polar CodesabstractDiverse application scenarios in 5G and beyond wireless communication systems have introduced various requirements in code lengths and rates of channel codes. For the decoding of polar codes, especially the belief-propagation (BP) decoding, flexible configuration of codeword length is still not involved in current decoders. In this paper, a semi-folded decoding structure is proposed which can be reconfigured to support multiple codeword lengths. Up to 16 codes can be decoded in parallel and the utilization of processing units is no less than 87.5% for various codeword lengths. The peak throughput of 19.29 Gbps can be achieved by the proposed decoder in SMIC 55 nm CMOS technology. Shan Cao 0001, Limin Jiang, Ting Lin, Shunqing Zhang, Shugong Xu |
ISCAS | 1 |
| 2021 | IFR: Iterative Fusion Based Recognizer for Low Quality Scene Text Recognition
Zhiwei Jia, Shugong Xu, Shiyi Mu, Yue Tao, Shan Cao 0001 |
PRCV (2) | 5 |
| 2021 | Attention based convolutional recurrent neural network for environmental sound classificationabstractEnvironmental sound classification (ESC) is a challenging problem due to the complexity of sounds. The classification performance is heavily dependent on the effectiveness of representative features extracted from the environmental sounds. However, ESC often suffers from the semantically irrelevant frames and silent frames. In order to deal with this, we employ a frame-level attention model to focus on the semantically relevant frames and salient frames. Specifically, we first propose a convolutional recurrent neural network to learn spectro-temporal features and temporal correlations. Then, we extend our convolutional RNN model with a frame-level attention mechanism to learn discriminative feature representations for ESC. We investigated the classification performance when using different attention scaling function and applying different layers. Experiments were conducted on ESC-50 and ESC-10 datasets. Experimental results demonstrated the effectiveness of the proposed method and our method achieved the state-of-the-art or competitive classification accuracy with lower computational complexity. We also visualized our attention results and observed that the proposed attention mechanism was able to lead the network tofocus on the semantically relevant parts of environmental sounds. Shugong Xu, Shunqing Zhang, Tianhao Qiao, Shan Cao 0001 |
Neurocomputing | 5 |
| 2021 | Predictive Wireless Based Status Update for Communication-Agnostic SamplingabstractIn a wireless network that conveys status updates from sources (i.e., sensors) to destinations, one of the key issues studied by existing literature is how to design an optimal source sampling strategy on account of the communication constraints which are often modeled as queues. In this paper, an alternative perspective is presented—a novel status-aware communication scheme, namelyparallel communications, is proposed which allows sensors to be communication-agnostic. Specifically, the proposed scheme can determine, based on an online prediction functionality, whether a status packet is worth transmitting considering both the network condition and status prediction, such that sensors can generate status packets without communication constraints. We evaluate the proposed scheme on a Software-Defined-Radio (SDR) test platform, which is integrated with a collaborative autonomous driving simulator, i.e., Simulation-of-Urban-Mobility (SUMO), to produce realistic vehicle control models and road conditions. The results show that with online status predictions, the channel occupancy is significantly reduced, while guaranteeing low status recovery error. Then the framework is applied to two scenarios: a multi-density platooning scenario, and a flight formation control scenario. Simulation results show that the scheme achieves better performance on the network level, in terms of keeping the minimum safe distance in both vehicle platooning and flight control. Zhiyuan Jiang, Wei Zhang 0001, Zixu Cao, Shan Cao 0001, Shunqing Zhang, Shugong Xu |
IEEE Trans. Wirel. Commun. | 4 |
| 2020 | A Cross Domain Multi-modal Dataset for Robust Face Anti-spoofingabstractFace Anti-spoofing (FAS) is a challenging problem due to the complex serving scenario and diverse face presentation attack patterns. Using single modal images which are usually captured with RGB cameras is not able to deal with the former because of serious overfitting problems. The existing multi-modal FAS datasets rarely pay attention to the cross domain problems, training FAS system on these data leads to inconsistencies and low generalization capabilities in deployment since imaging principles(structured light, TOF, etc.) and pre-processing methods vary between devices. We explore the subtle fine-grained differences betweeen multi-modal cameras and proposed a cross domain multi-modal FAS dataset GREAT-FASD and several evaluation protocols for academic community. Furthermore, we incorporate the multiplicative attention and center loss to enhance the representative power of CNN via seeking out complementary information as a powerful baseline. In addition, extensive experiments have been conducted on the proposed dataset to analyze the robustness to distinguish spoof faces and bona-fide faces. Experimental results show the effectiveness of proposed method and achieve the state-of-the-art competitive results. Finally, we visualize our future distribution in hidden space and observe that the proposed method is able to lead the network to generate a large margin for face anti-spoofing task. Qiaobin Ji, Shugong Xu, Shunqing Zhang, Shan Cao 0001 |
ICPR | 5 |
| 2020 | Revealing Much While Saying Less: Predictive Wireless for Status UpdateabstractWireless communications for status update are becoming increasingly important, especially for machine-type control applications. Existing work has been mainly focused on Age of Information (AoI) optimizations. In this paper, a status-aware predictive wireless interface design, networking and implementation are presented which aim to minimize the status recovery error of a wireless networked system by leveraging online status model predictions. Two critical issues of predictive status update are addressed: practicality and usefulness. Link-level experiments on a Software-Defined-Radio (SDR) testbed are conducted and test results show that the proposed design can significantly reduce the number of wireless transmissions while maintaining a low status recovery error. A Status-aware Multi-Agent Reinforcement learning neTworking solution (SMART) is proposed to dynamically and autonomously control the transmit decisions of devices in an ad hoc network based on their individual statuses. System-level simulations of a multi dense platooning scenario are carried out on a road traffic simulator. Results show that the proposed schemes can greatly improve the platooning control performance in terms of the minimum safe distance between successive vehicles, in comparison with the AoI-optimized status-unaware and communication latency-optimized schemes-this demonstrates the usefulness of our proposed status update schemes in a real-world application. Zhiyuan Jiang, Zixu Cao, Siyu Fu, Shan Cao 0001, Shunqing Zhang, Shugong Xu |
INFOCOM | 5 |
| 2020 | A Novel Terminal Aided Synchronization Scheme for Intelligent Transportation Systems with Vehicle-to-Anything (V2X) CommunicationsabstractSynchronization, as a critical factor of modern wireless communication systems, has attracted close research attention recent years. For the vehicle-to-anything (V2X) communication in 5G new radio, synchronization faces severe challenges due to the extremely low latency and high reliability requirements. In this paper, a terminal aided synchronization scheme is proposed for the vehicle platooning in V2X communication. The shared information, such as NSLIDfrom other cooperative vehicles, are utilized to recover the original transmitted sidelink synchronization signals. The synchronization ID detection probability is therefore improved by 49.6% compared to conventional schemes. Hardware implementation on FPGA Artix-7 AC701 board is performed of the proposed synchronization scheme and the hardware latency is reduced to 67.18 μs compared to 968,654.85 μs in conventional schemes. Shunqing Zhang, Shan Cao 0001, Shugong Xu, Yi Shi 0004 |
ISCAS | 3 |
| 2020 | Hardware-Software Co-Design for Face Recognition on FPGA SoCsabstractWith the development of deep learning, face recognition is attracting more and more attention in both industry and academia. Hardware implementation of face recognition systems on heterogeneous embedded devices, however has been rarely studies. In this paper, an embedded face recognition system is designed and implemented on FPGA SoC platforms. A hardware-software partition method is first introduced by analyzing the ratio between computation and memory access of critical tasks in the system. Several acceleration methods are then exploited to optimize the hardware implementation. The face recognition system is implemented on Xilinx FPGA MPSoC ZCU102 with 97.3% recognition accuracy and 203.7 ms latency. The neural network VIPLFace, as the most time consuming part of the system, has a 74 ms latency, 71× faster after hardware-software co-design. Shan Cao 0001, Shugong Xu, Shunqing Zhang |
ISCAS | 2 |
| 2019 | Channel Estimation for WiFi Prototype Systems with Super-Resolution Image RecoveryabstractChannel estimation is crucial for modern WiFi system and becomes more and more challenging with the growth of user throughput in multiple input multiple output configuration. Plenty of literature spends great efforts in improving the estimation accuracy, while the interpolation schemes are overlooked. To deal with this challenge, we exploit the super-resolution image recovery scheme to model the non-linear interpolation mechanisms without pre-assumed channel characteristics in this paper. To make it more practical, we offline generate numerical channel coefficients according to the statistical channel models to train the neural networks, and directly apply them in some practical WiFi prototype systems. As shown in this paper, the proposed super-resolution based channel estimation scheme can outperform the conventional approaches in both LOS and NLOS scenarios, which we believe can significantly change the current channel estimation method in the near future. Qi Shi 0004, Yangyu Liu, Shunqing Zhang, Shugong Xu, Shan Cao 0001, Vincent K. N. Lau |
ICC | 5 |
| 2019 | Robust Sub-Meter Level Indoor Localization - A Logistic Regression ApproachabstractIndoor localization becomes a raising demand in our daily lives. Due to the massive deployment in the indoor environment nowadays, WiFi systems have been applied to high accurate localization recently. Although the traditional model based localization scheme can achieve sub-meter level accuracy by fusing multiple channel state information (CSI) observations, the corresponding computational overhead is significant. To address this issue, the model-free localization approach using deep learning framework has been proposed and the classification based technique is applied. In this paper, instead of using classification based mechanism, we propose to use a logistic regression based scheme under the deep learning framework, which is able to achieve sub-meter level accuracy (97.2cm medium distance error) in the standard laboratory environment and maintain reasonable online prediction overhead under the single WiFi AP settings. We hope the proposed logistic regression based scheme can shed some light on the model-free localization technique and pave the way for the practical deployment of deep learning based WiFi localization systems. Chenlu Xiang, Shunqing Zhang, Shugong Xu, Shan Cao 0001, Vincent K. N. Lau |
ICC | 5 |
| 2019 | A Pre-RTL Simulator for Neural NetworksabstractIn this paper, a pre-RTL neural network simulator (SimuNN) is proposed which is initiated as the bridge between the algorithm design and hardware implementation of neural networks. SimuNN is compatible with TensorFlow interface, and supports inference in both floating-point numbers and configurable fixed-point numbers. It can provide inference results at layer-/module-/cycle-level to serve as a golden model for RTL designs. Besides, its embedded model for hardware performance estimation enables SimuNN to provide an accurate reference of processing speed and hardware cost at the ASIC-designed user end for algorithm designers. Shan Cao 0001, Zhenyi Bao, Chengbo Xue, Shugong Xu, Shunqing Zhang |
ISCAS | 1 |
| 2019 | Passive TCP Identification for Wired and Wireless Networks: A Long-Short Term Memory ApproachabstractTransmission control protocol (TCP) congestion control is one of the key techniques to improve network performance. TCP congestion control algorithm identification (TCP identification) can be used to significantly improve network efficiency. Existing TCP identification methods can only be applied to limited number of TCP congestion control algorithms and focus on wired networks. In this paper, we proposed a machine learning based passive TCP identification method for wired and wireless networks. After comparing among three typical machine learning models, we concluded that the 4-layers Long Short Term Memory (LSTM) model achieves the best identification accuracy. Our approach achieves better than 98% accuracy in wired and wireless networks and works for newly proposed TCP congestion control algorithms. Shugong Xu, Shan Cao 0001, Shunqing Zhang, Yanzan Sun |
IWCMC | 4 |
| 2019 | Attention Based Convolutional Recurrent Neural Network for Environmental Sound Classification
Shugong Xu, Tianhao Qiao, Shunqing Zhang, Shan Cao 0001 |
PRCV (1) | 5 |
| 2019 | Energy-Efficient Subchannel and Power Allocation for HetNets Based on Convolutional Neural NetworkabstractHeterogeneous network (HetNet) has been proposed as a promising solution for handling the wireless traffic explosion in future fifth-generation (5G) system. In this paper, a joint subchannel and power allocation problem is formulated for HetNets to maximize the energy efficiency (EE). By decomposing the original problem into a classification subproblem and a regression subproblem, a convolutional neural network (CNN) based approach is developed to obtain the decisions on subchannel and power allocation with a much lower complexity than conventional iterative methods. Numerical results further demonstrate that the proposed CNN can achieve similar performance as the Exhaustive method, while needs only 6.76% of its CPU runtime. Xiaojing Chen 0001, Changhao Wu, Shunqing Zhang, Shugong Xu, Shan Cao 0001 |
VTC Spring | 6 |
| 2019 | Fingerprint-Based Localization Using Commercial LTE Signals: A Field-Trial StudyabstractWireless localization for mobile device has attracted more and more interests by increasing the demand for location based services. Fingerprint-based localization is promising, especially in non-Line-of-Sight (NLoS) or rich scattering environments, such as urban areas and indoor scenarios. In this paper, we propose a novel fingerprint-based localization technique based on deep learning framework under commercial long term evolution (LTE) systems. Specifically, we develop a software defined user equipment to collect the real time channel state information (CSI) knowledge from LTE base stations and extract the intrinsic features among CSI observations. On top of that, we propose a time domain fusion approach to assemble multiple positioning estimations. Experimental results demonstrated that the proposed localization technique can significantly improve the localization accuracy and robustness, e.g. achieves Mean Distance Error (MDE) of 0.47 meters for indoor and of 19.9 meters for outdoor scenarios, respectively. Heng Zhang 0040, Shunqing Zhang, Shugong Xu, Shan Cao 0001 |
VTC Fall | 5 |
| 2019 | Efficient MIMO Detection with Imperfect Channel Knowledge - A Deep Learning ApproachabstractMultiple-input multiple-output (MIMO) system is the key technology for long term evolution (LTE) and 5G. The information detection problem at the receiver side is in general difficult due to the imbalance of decoding complexity and decoding accuracy within conventional methods. Hence, a deep learning based efficient MIMO detection approach is proposed in this paper. In our work, we use a neural network to directly get a mapping function of received signals, channel matrix and transmitted bit streams. Then, we compare the end-to-end approach using deep learning with the conventional methods in possession of perfect channel knowledge and imperfect channel knowledge. Simulation results show that our method presents a better trade-off in the performance for accuracy versus decoding complexity. At the same time, better robustness can be achieved in condition of imperfect channel knowledge compared with conventional algorithms. Qian Chen 0006, Shunqing Zhang, Shugong Xu, Shan Cao 0001 |
WCNC | 4 |
| 2018 | Performance Evaluation for LTE-V based Vehicle-to-Vehicle Platooning CommunicationabstractWith the raising demand for autonomous driving, vehicle-to-vehicle communications becomes a key technology enabler for the future intelligent transportation system. Based on our current knowledge field, there is limited network simulator that can support end-to-end performance evaluation for LTE-V based vehicle-to-vehicle platooning systems. To address this problem, we start with an integrated platform that combines traffic generator and network simulator together, and build the V2V transmission capability according to LTE-V specification. On top of that, we simulate the end-to-end throughput and delay profiles in different layers to compare different configurations of platooning systems. Through numerical experiments, we show that the LTE-V system is unable to support the highest degree of automation under shadowing effects in the vehicle platooning scenarios, which requires ultra-reliable low-latency communication enhancement in 5G networks. Meanwhile, the throughput and delay performance for vehicle platooning changes dramatically in PDCP layers, where we believe further improvements are necessary. Tao Yu 0008, Shunqing Zhang, Shan Cao 0001, Shugong Xu |
APCC | 3 |
| 2018 | Grey Correlation Degree Analysis on Pilot Pattern Optimization for OFDM Channel EstimationabstractFor underwater acoustic communication, pilot pattern optimization is usually investigated to improve the performance of channel estimation based on compressed sensing (CS) in orthogonal frequency division multiplexing (OFDM) systems. However, there is no deterministic criteria to design a perfect pilot pattern utilizing the measurement matrix, and no mature methods to quantitatively measure the relationship between the influence indicators and estimation performance of pilot patterns. An analytical method with grey correlation degree is proposed to try to solve the problem. The influence indicators are weighted with information entropy and the grey correlation degrees of various optimization strategies are calculated. Experimental results demonstrate the proposed method is intuitive and effective, due to the order of the grey correlation degrees entirely consists with the order of the channel estimation performance on bit error rate (BER) and mean square error (MSE). Moreover, it is indicated that the ratio of large off-diagonal entries in the Gram matrix has a greater impact on the performance of channel estimation compared to the minimal mutual coherence, the ratio of small off-diagonal entries, and the ratio of middle off-diagonal entries. Rongkun Jiang, Shan Cao 0001 |
GLOBECOM | 2 |
| 2018 | Dynamic Carrier and Power Amplifier Mapping for Energy Efficient Multi-Carrier Wireless CommunicationsabstractThe rapid increasing demand of wireless transmission has incurred mobile broadband to continuously evolve through multiple frequency bands, massive antennas and other multi-stream processing schemes. Together with the improved data transmission rate, the power consumption for multi-carrier transmission and processing is proportionally increasing, which contradicts with the energy efficiency requirements of 5G wireless systems. To meet this challenge, multi carrier power amplifier (MCPA) technology, e.g., to support multiple carriers through a single power amplifier, is widely deployed in practical. With massive carriers required for 5G communication and limited number of carriers supported per MCPA, a natural question to ask is how to map those carriers into multiple MCPAs and whether we shall dynamically adjust this mapping relation. In this paper, we have theoretically formulated the dynamic carrier and MCPA mapping problem to jointly optimize the traditional separated baseband and radio frequency processing. On top of that, we have also proposed a low complexity algorithm that can achieve most of the power saving with affordable computational time, if compared with the optimal exhaustive search based algorithm. Shunqing Zhang, Chenlu Xiang, Shan Cao 0001, Shugong Xu |
ICC | 3 |
| 2018 | A Reconfigurable Pipelined Architecture for Convolutional Neural Network AccelerationabstractThe convolutional neural network (CNN) has become widely used in a variety of vision recognition applications, and the hardware acceleration of CNN is in urgent need as increasingly more computations are required in the state-of-the-art CNN networks. In this paper, we propose a pipelined architecture for CNN acceleration. The probability of both inner-layer and inter-layer pipeline for typical CNN networks is analyzed. And two types of data re-ordering methods, the filter-first (FF) flow and the image-first (IF) flow, are proposed for different kinds of layers. Then, a pipelined CNN accelerator for AlexNet is implemented, the dataflow of which can be reconfigurably selected for different layer processing. Simulation results show that the proposed pipelined architecture achieves 43% performance improvement compared with the non-pipelined ones. The AlexNet accelerator is implemented in 65nm CMOS technology working at 200MHz, with 350mW power consumption and 24GFLOPS peak performance. Chengbo Xue, Shan Cao 0001, Rongkun Jiang |
ISCAS | 2 |
| 2018 | Deep Convolutional Neural Network with Mixup for Environmental Sound Classification
Shugong Xu, Shan Cao 0001, Shunqing Zhang |
PRCV (2) | 3 |
| 2018 | Fast intra coding based on CU size decision and direction mode decision for HEVC
Xinghua Wang 0005, Shan Cao 0001 |
Multim. Tools Appl. | 4 |
| 2016 | Temperature-aware task scheduling heuristics on Network-on-ChipsabstractChip temperature becomes a critical design issue with technology scaling to nanometer-scale, especially for NoC systems with large number of cores and shrunken core size. To reduce peak temperature and balance spatial temperature distribution on NoC-based multi-cores chips, this paper proposes a temperature-aware task scheduling approach. The thermal profiles of tasks are first extracted by accurate temperature model. Then run-time task mapping heuristic is proposed considering transient core temperatures, thermal dissipation from adjacent cores, communication overheads and the thermal influence of physical position on chip. Voltage-frequency is also scaled down when timing constraint is met to reduce power consumption and core temperature. Experimental results show that the significant reduction of peak temperature and the temperature variance compared with the current approaches is achieved. Shan Cao 0001, Zoran A. Salcic, Yingtao Ding, Zhaolin Li, Shaojun Wei, Xianli Zhao |
ISCAS | 1 |
| 2015 | Schedule refinement for homogeneous multi-core processors in the presence of manufacturing-caused heterogeneityabstractMulti-core homogeneous processors have been widely used to deal with computation-intensive embedded applications. However, with the continuous down scaling of CMOS technology, within-die variations in the manufacturing process lead to a significant spread in the operating speeds of cores within homogeneous multi-core processors. Task scheduling approaches, which do not consider such heterogeneity caused by within-die variations, can lead to an overly pessimistic result in terms of performance. To realize an optimal performance according to the actual maximum clock frequencies at which cores can run, we present a heterogeneity-aware schedule refining (HASR) scheme by fully exploiting the heterogeneities of homogeneous multi-core processors in embedded domains. We analyze and show how the actual maximum frequencies of cores are used to guide the scheduling. In the scheme, representative chip operating points are selected and the corresponding optimal schedules are generated as candidate schedules. During the booting of each chip, according to the actual maximum clock frequencies of cores, one of the candidate schedules is bound to the chip to maximize the performance. A set of applications are designed to evaluate the proposed scheme. Experimental results show that the proposed scheme can improve the performance by an average value of 22.2%, compared with the baseline schedule based on the worst case timing analysis. Compared with the conventional task scheduling approach based on the actual maximum clock frequencies, the proposed scheme also improves the performance by up to 12%. Zhixiang Chen 0003, Zhaolin Li, Shan Cao 0001, Jie Zhou 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2014 | Compiler-Assisted Leakage- and Temperature- Aware Instruction-Level VLIW SchedulingabstractWith technology scaled to nanometer-scale, leakage energy consumption is accounting for a greater proportion than ever, especially for very long instruction word (VLIW) architectures with a large number of functional units (FUs). The growing energy consumption leads to an increase in chip temperature, which again brings an exponential growth in leakage current, and consequently leakage energy. However, few studies consider both leakage energy and temperature reduction during the compiling on VLIW architectures. In this paper, a leakage- and temperature-aware design flow is presented to assist the compiling of instruction-level VLIW scheduling. And two scheduling algorithms are proposed for the design flow. First, the leakage-aware rescheduling algorithm is proposed for leakage energy reduction by concentrating operations to fewer FUs and shutting more FUs down. Then, the temperature-aware workload balance algorithm is presented to reduce peak temperature by balancing the concentrated workloads among homogenous FUs. It is proved that the proposed two algorithms can reduce the leakage energy and peak temperature without performance loss. Experimental results demonstrate that the peak temperature is reduced by 15.27% and 12.84% for FU groups with three and two FUs and the leakage energy is reduced by 78.14% and 30.31% on average compared with the communication scheduling and list algorithm, respectively. Shan Cao 0001, Zhaolin Li, Shaojun Wei |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | Energy-efficient stream task scheduling scheme for embedded multimedia applications on multi-issued stream architectures
Shan Cao 0001, Zhaolin Li, Guoyue Jiang, Zhixiang Chen 0003, Shaojun Wei |
J. Syst. Archit. | 1 |